How to Aggregate Job Postings from 500+ Companies Using Public ATS APIs
Greenhouse, Lever and Ashby expose public job board APIs with no auth. Build one aggregator that pulls from all three, dedupes, and flags new roles.
The actor referenced in this article. Pay only for results delivered.
The most underrated dataset in recruiting tech is the one sitting in plain sight: public job board APIs from Greenhouse, Lever, and Ashby. Every company using these ATSes exposes a fully queryable endpoint with no API key, no rate limiting beyond reasonable use, and structured JSON output.
TL;DR: Build a job aggregator by detecting which ATS each company uses (try Greenhouse, Lever, then Ashby in order), defining a normalized Job dataclass, and running async collection across 500+ companies in parallel with aiohttp. The hardest part is slug discovery: company names in lowercase work for roughly 80% of cases. Full aggregator in about 200 lines of Python.
Most major tech companies use one of these three. Stripe, Coinbase, Notion, Vercel, Linear, Figma, Shopify, and hundreds of others expose their open roles through these APIs. This guide shows how to build a functional job aggregator in about 200 lines of Python.
Step 1: Company Slug Discovery
The hardest part of building an ATS aggregator is not the API calls: it is knowing which company uses which ATS and what their slug is.
Greenhouse slugs: Usually the company’s name in lowercase. stripe, airbnb, notion, coinbase, cloudflare. You can find a company’s Greenhouse slug by visiting jobs.greenhouse.io/{slug} and checking if it redirects to their job board.
Lever slugs: Also usually the company name. Check jobs.lever.co/{slug}.
Ashby slugs: Check jobs.ashbyhq.com/{slug}.
A practical approach is to maintain a company list and detect which ATS each uses:
import asyncio
import aiohttp
COMPANIES = [
'stripe', 'notion', 'linear', 'vercel', 'supabase',
'figma', 'loom', 'webflow', 'retool', 'airtable',
'coda', 'clickup', 'monday', 'asana', 'basecamp',
# ... add more
]
async def detect_ats(session: aiohttp.ClientSession, slug: str) -> dict:
"""Try all three ATSes and return which one has data."""
endpoints = {
'greenhouse': f'https://boards-api.greenhouse.io/v1/boards/{slug}/jobs',
'lever': f'https://api.lever.co/v0/postings/{slug}?mode=json&limit=200',
'ashby': f'https://api.ashbyhq.com/posting-api/job-board/{slug}',
}
for ats, url in endpoints.items():
try:
async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as resp:
if resp.status == 200:
data = await resp.json()
count = (
len(data.get('jobs', [])) if ats == 'greenhouse' else
len(data) if ats == 'lever' else
len(data.get('jobs', []))
)
if count > 0:
return {'slug': slug, 'ats': ats, 'count': count}
except Exception:
continue
return {'slug': slug, 'ats': None, 'count': 0}
async def build_company_map(companies: list[str]) -> list[dict]:
async with aiohttp.ClientSession() as session:
tasks = [detect_ats(session, slug) for slug in companies]
return await asyncio.gather(*tasks)
company_map = asyncio.run(build_company_map(COMPANIES))
print(f"Found: {sum(1 for c in company_map if c['ats'])} / {len(COMPANIES)} companies")
Step 2: Normalized Data Model
The three ATSes return different schemas. Define a normalized output format before writing any collection code:
from dataclasses import dataclass
from datetime import datetime
from typing import Optional
@dataclass
class Job:
id: str
ats: str # 'greenhouse' | 'lever' | 'ashby'
company_slug: str
title: str
department: Optional[str]
location: Optional[str]
is_remote: Optional[bool]
employment_type: Optional[str]
url: str
description_html: Optional[str]
description_plain: Optional[str]
published_at: Optional[datetime]
collected_at: datetime
Step 3: Per-ATS Collectors
async def collect_greenhouse(session, slug: str) -> list[Job]:
async with session.get(
f'https://boards-api.greenhouse.io/v1/boards/{slug}/jobs',
params={'content': 'true'} # Include description
) as resp:
data = await resp.json()
jobs = []
for j in data.get('jobs', []):
location = j.get('location', {}).get('name', '')
jobs.append(Job(
id=str(j['id']),
ats='greenhouse',
company_slug=slug,
title=j['title'],
department=j['departments'][0]['name'] if j.get('departments') else None,
location=location,
is_remote='remote' in location.lower() if location else None,
employment_type=None, # Greenhouse doesn't expose this
url=j['absolute_url'],
description_html=j.get('content'),
description_plain=None,
published_at=datetime.fromisoformat(j['updated_at'].replace('Z', '+00:00')) if j.get('updated_at') else None,
collected_at=datetime.utcnow(),
))
return jobs
async def collect_lever(session, slug: str) -> list[Job]:
async with session.get(
f'https://api.lever.co/v0/postings/{slug}?mode=json&limit=200'
) as resp:
data = await resp.json()
jobs = []
for j in data:
cats = j.get('categories', {})
workplace = cats.get('workplaceType', '').lower()
jobs.append(Job(
id=j['id'],
ats='lever',
company_slug=slug,
title=j['text'],
department=cats.get('department'),
location=cats.get('location'),
is_remote=workplace == 'remote',
employment_type=cats.get('commitment'),
url=j['hostedUrl'],
description_html=j.get('description'),
description_plain=j.get('descriptionPlain'),
published_at=datetime.fromtimestamp(j['createdAt'] / 1000) if j.get('createdAt') else None,
collected_at=datetime.utcnow(),
))
return jobs
async def collect_ashby(session, slug: str) -> list[Job]:
async with session.get(
f'https://api.ashbyhq.com/posting-api/job-board/{slug}'
) as resp:
data = await resp.json()
jobs = []
for j in data.get('jobs', []):
jobs.append(Job(
id=j['id'],
ats='ashby',
company_slug=slug,
title=j['title'],
department=j.get('department') or j.get('team'),
location=j.get('location'),
is_remote=j.get('isRemote', False),
employment_type=j.get('employmentType'),
url=f"https://jobs.ashbyhq.com/{slug}/{j['id']}",
description_html=j.get('jobDescription'),
description_plain=None,
published_at=datetime.fromisoformat(j['publishedDate']) if j.get('publishedDate') else None,
collected_at=datetime.utcnow(),
))
return jobs
Step 4: Parallel Collection
With 500 companies, sequential collection takes too long. Run them concurrently:
COLLECTORS = {
'greenhouse': collect_greenhouse,
'lever': collect_lever,
'ashby': collect_ashby,
}
async def collect_all(company_map: list[dict]) -> list[Job]:
async with aiohttp.ClientSession() as session:
tasks = []
for company in company_map:
if company['ats']:
collector = COLLECTORS[company['ats']]
tasks.append(collector(session, company['slug']))
results = await asyncio.gather(*tasks, return_exceptions=True)
all_jobs = []
for result in results:
if isinstance(result, list):
all_jobs.extend(result)
return all_jobs
all_jobs = asyncio.run(collect_all(company_map))
print(f"Collected {len(all_jobs)} jobs from {len([c for c in company_map if c['ats']])} companies")
Step 5: Keep Only New Postings
An aggregator that reruns daily needs to know which jobs it has already seen. Key on the ATS job ID rather than a hash of the title: the ID stays the same across runs, while titles get edited. A small SQLite table is enough:
import sqlite3
def new_jobs_only(jobs: list[Job], db_path: str = 'jobs.db') -> list[Job]:
"""Return jobs not seen in any earlier run, and record them as seen."""
conn = sqlite3.connect(db_path)
conn.execute('CREATE TABLE IF NOT EXISTS seen (key TEXT PRIMARY KEY, first_seen TEXT)')
fresh = []
for job in jobs:
key = f'{job.ats}:{job.company_slug}:{job.id}'
cur = conn.execute(
'INSERT OR IGNORE INTO seen VALUES (?, ?)',
(key, job.collected_at.isoformat()),
)
if cur.rowcount == 1:
fresh.append(job)
conn.commit()
conn.close()
return fresh
The first run treats every posting as new, so call new_jobs_only(all_jobs) once to seed the table before you switch on any alerting.
Step 6: Score and Analyze Jobs with Claude
Once the jobs are in one schema, a language model can read them for you. Two patterns are worth the few lines of code: scoring each posting against a candidate profile, and counting which skills show up across many descriptions.
Greenhouse returns content as HTML with the tags entity-encoded, so strip it to plain text before sending it anywhere:
import html
import json
import re
import anthropic
claude = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
def plain_description(job: Job) -> str:
if job.description_plain:
return job.description_plain
raw = html.unescape(job.description_html or '')
return re.sub(r'\s+', ' ', re.sub(r'<[^>]+>', ' ', raw)).strip()
def parse_json(text: str) -> dict:
"""Pull the JSON object out of a model reply, ignoring any code fences."""
start, end = text.find('{'), text.rfind('}')
return json.loads(text[start:end + 1]) if start != -1 else {}
Fit scoring
Keyword matching a resume against a job description misses a lot. Claude can weigh transferable experience, seniority hints and the unstated requirements an experienced reader would pick up:
def score_job_fit(jobs: list[Job], candidate_profile: str) -> list[tuple[Job, dict]]:
scored = []
for job in jobs:
response = claude.messages.create(
model='claude-haiku-4-5',
max_tokens=400,
messages=[{
'role': 'user',
'content': f"""Score this job against the candidate profile. Return JSON only.
CANDIDATE PROFILE:
{candidate_profile}
JOB:
Company: {job.company_slug} ({job.ats})
Title: {job.title}
Department: {job.department or 'N/A'}
Location: {job.location or 'N/A'}
Description: {plain_description(job)[:2000]}
Return:
{{"fit_score": 1-10, "match_reasons": ["..."], "gaps": ["..."],
"apply_recommendation": "strong_yes" | "yes" | "maybe" | "no",
"standout_angle": "one sentence on how this candidate should pitch themselves"}}""",
}],
)
try:
analysis = parse_json(response.content[0].text)
except json.JSONDecodeError:
analysis = {}
scored.append((job, analysis))
return sorted(scored, key=lambda pair: pair[1].get('fit_score', 0), reverse=True)
candidate = """
Senior software engineer, 6 years. Strong: Python, TypeScript, distributed systems, PostgreSQL.
Moderate: Kubernetes, AWS, React. Wants Series A to C companies and data-heavy products.
Not interested in pure frontend or crypto.
"""
Skills trends
Job descriptions show which tools employers are adopting before any survey does. Send descriptions in batches of 10 and ask for counts:
def extract_skills(jobs: list[Job], title_filter: str | None = None) -> dict[str, int]:
if title_filter:
jobs = [j for j in jobs if title_filter.lower() in j.title.lower()]
totals: dict[str, int] = {}
for i in range(0, len(jobs), 10):
batch = '\n\n---\n\n'.join(
f'Role: {j.title}\n{plain_description(j)[:500]}' for j in jobs[i:i + 10]
)
response = claude.messages.create(
model='claude-haiku-4-5',
max_tokens=600,
messages=[{
'role': 'user',
'content': f"""List the technical skills, tools, frameworks and methodologies in these job descriptions.
{batch}
Return one JSON object mapping each skill to the number of descriptions that mention it,
for example {{"Python": 7, "Kubernetes": 4}}. Normalise names ("postgres" is "PostgreSQL",
"k8s" is "Kubernetes"). Return only JSON.""",
}],
)
try:
for skill, count in parse_json(response.content[0].text).items():
totals[skill] = totals.get(skill, 0) + int(count)
except (json.JSONDecodeError, ValueError, TypeError):
continue
return dict(sorted(totals.items(), key=lambda kv: kv[1], reverse=True))
backend = extract_skills(all_jobs, title_filter='backend')
Divide each count by the number of postings in the filter. A skill in 60% or more of postings is table stakes for that role. If you store the counts each month, the skills that keep climbing are the ones to watch.
For the competitor side of this data (what a rival’s department mix says about their roadmap), see how to monitor competitor job postings.
Step 7: A Daily Alert for New Matches
Put the pieces together and run it once a day. Only new postings get scored, so the model calls stay small after the first run:
def daily_alert(candidate_profile: str, min_fit: int = 7):
jobs = asyncio.run(collect_all(company_map))
fresh = new_jobs_only(jobs)
if not fresh:
print('No new postings since the last run.')
return
for job, analysis in score_job_fit(fresh, candidate_profile):
if analysis.get('fit_score', 0) < min_fit:
continue
print(f"{analysis['fit_score']}/10 {job.title} at {job.company_slug}")
print(f" {analysis.get('standout_angle', '')}")
print(f" {job.url}")
daily_alert(candidate)
Swap the print calls for an email or Slack message and schedule it with cron. Watching a dozen target companies this way means you hear about a new role within a day of it going up.
Using the Managed Version
If you want this without the infrastructure, our ATS Jobs scraper runs this exact pipeline on Apify:
run = client.actor('themineworks/ats-jobs').call(run_input={
'companies': COMPANIES,
'maxJobsPerCompany': 100,
'includeDescription': True,
})
The output is the same normalized schema regardless of which ATS each company uses.
Frequently Asked Questions
How do you detect which ATS a company uses without manual research?
Try each ATS in order of market share: Greenhouse first (boards.greenhouse.io/{slug}), then Lever (jobs.lever.co/{slug}), then Ashby (jobs.ashbyhq.com/{slug}). The first URL that returns a 200 with job data is the correct ATS. Automate this with aiohttp and run the probes in parallel: detecting the ATS for 500 companies takes under 2 minutes. About 80% of tech companies use Greenhouse, making it the right first check.
What is the best way to handle rate limiting when aggregating hundreds of ATS companies?
Greenhouse, Lever, and Ashby are all public APIs with no authentication, so they have soft rate limits rather than hard API quotas. Use asyncio with a concurrency cap of 10-20 simultaneous requests, add a 0.5-second delay between batches per platform, and implement exponential backoff on 429 responses. A full crawl of 500 companies across all three platforms typically completes in 5-10 minutes at moderate concurrency.
Why do you need a normalized data model when aggregating multiple ATS platforms?
Each ATS returns different field names for the same data: Greenhouse uses title, Lever uses text, Ashby uses jobTitle. Location is a string in Greenhouse but a nested object in Lever. Without a normalized Job dataclass, downstream code has to branch on platform type everywhere, which breaks silently when a platform changes its schema. A normalized model means all consumer code works on a single predictable structure regardless of source.
How do you find the job board slug for a company on Greenhouse, Lever, or Ashby?
The slug is almost always the company name lowercased with spaces replaced by hyphens: stripe, linear, vercel, notion. For companies with uncommon names, check their careers page: the ATS URL is usually visible in links like boards.greenhouse.io/stripe or embedded in the page source. Automated slug detection works for roughly 80% of companies; the remaining 20% require a manual lookup or fuzzy matching against a maintained company-slug database.
What does it take to build a job aggregator covering 500+ tech companies?
The core is a 200-line Python script: ATS detection, async data collection with aiohttp, a normalized Job dataclass, and a simple SQLite store. The hard part is maintaining the company-slug list: companies get acquired, change ATSs, or go dark. Plan for 5-10% of your company list to need updating per quarter. A weekly cron job that runs the full crawl and alerts on companies with sudden zero-job counts catches most breakage before it becomes stale data.
Why score job fit with Claude instead of keyword matching?
Keyword matching gives false positives (a DevOps engineer who once listed Python matches “senior Python developer”) and false negatives (a resume that says “data pipelines” misses a posting titled “ETL engineer”). A model reads the whole description, so it can weigh transferable experience, seniority hints and the requirements a posting implies without spelling out. Haiku is enough for per-posting scoring, and because Step 5 filters to new postings only, each daily run scores a handful of jobs rather than the whole list.
Explore the scraper referenced in this article: inputs, outputs, and pricing, then run it on Apify.
Greenhouse vs Lever vs Ashby: Which ATS Has the Best Public Job API?
Greenhouse, Lever, and Ashby all expose public job board APIs with no authentication. A field-by-field comparison of what each returns, their limits, and how to pull all three into one dataset.
Building a Job Market Intelligence Dashboard with Free ATS Data
How to build a real-time hiring dashboard that tracks roles, skills demand, and company hiring velocity using public Greenhouse, Lever, and Ashby APIs.
How to Monitor Competitor Job Postings to Predict Their Strategy
Job postings are the most honest signal of a competitor's roadmap. Track their ATS boards automatically and turn hiring data into strategy.
Recruitment Automation: Building a Job Intelligence Pipeline with Free ATS Data
How to use the public Greenhouse, Lever and Ashby APIs to build automated job monitoring and salary benchmarking for recruiting teams.