The Mine Works
ClinicalTrials.gov API v2: How to Search 600,000 Studies and Track Trial Status
← All posts
tutorial June 22, 2026 · 16 min read Updated October 9, 2026

ClinicalTrials.gov API v2: How to Search 600,000 Studies and Track Trial Status

How to use the ClinicalTrials.gov v2 REST API: what changed from v1, which filters actually exist, and Python for bulk pulls and status monitoring.

Try the scraper

The actor referenced in this article. Pay only for results delivered.

View the scraper →

ClinicalTrials.gov is the US registry for clinical studies. On 8 October 2026 it held 606,387 study records, and the /stats/size endpoint returns that count live, so you can check it rather than trust it. The data covers trial phase, status, enrollment, eligibility criteria, sponsor, interventions, outcomes and locations, with sites in 226 countries and territories.

In May 2023, NLM (the National Library of Medicine, which operates the registry) announced the retirement of the v1 API. V1 is now gone. A request to /api/query/ returns a plain 404 from the web server rather than a deprecation notice, so code written against v1 fails outright instead of quietly returning less.

This post covers what changed, how the v2 parameters actually behave (including two filters that do not exist but get copied around anyway), how to build a status monitor, and the three actors we publish for teams who would rather not own the pagination and flattening code.

Try it live: ClinicalTrials.gov Scraper, 26 Fields, Sites, No API Key. Pay per result delivered. Failed and empty results are never charged.

TL;DR: The v2 base URL is https://clinicaltrials.gov/api/v2/. Pagination uses pageToken (cursor-based), not min_rnk/max_rnk offset. Free-text search uses query.cond, query.intr, query.spons, query.locn and query.term. Structured filtering uses filter.overallStatus, filter.ids and filter.geo. There is no filter.phase and no filter.studyType: phase and study type go through filter.advanced as AREA[Phase] and AREA[StudyType]. Every study comes back nested under protocolSection.

What Changed from v1

The v1 API was at https://clinicaltrials.gov/api/query/. It used numeric offset pagination (min_rnk, max_rnk), returned XML or JSON, and had a flat field structure.

V2 breaks all of these:

AspectV1V2
Base URL/api/query//api/v2/
Paginationoffset: min_rnk/max_rnkcursor: pageToken
Response shapeFlat field mapNested protocolSection
Default formatXMLJSON
Filter syntaxexpr= (Lucene-like)Typed params: query.cond, filter.overallStatus, etc.

V1 is fully retired. https://clinicaltrials.gov/api/query/study_fields?expr=cancer&fields=NCTId&fmt=json returns HTTP 404 today, so anything still pointing there is dead rather than deprecated.

The V2 Endpoint Structure

The main endpoints in v2:

EndpointWhat it returns
/studiesSearch and filter studies, paginated
/studies/{nctId}Full record for a single study
/stats/sizeTotal number of studies in the registry
/stats/fieldValues/{field}All distinct values for a field, with counts
/versionAPI version and the timestamp of the last data load

The studies endpoint is the workhorse. All searches go through it.

/version is the cheapest health check available:

{"apiVersion": "2.0.5", "dataTimestamp": "2026-10-08T09:00:05"}

The data timestamp moves once a day. Polling faster than that buys you nothing.

Query Parameters

Search query params (free-text search over specific fields):

ParamSearchesExample
query.condCondition or diseasequery.cond=pancreatic+cancer
query.intrIntervention or drugquery.intr=pembrolizumab
query.titlesOfficial and brief titlequery.titles=phase+3+checkpoint
query.outcOutcomes measuresquery.outc=overall+survival
query.sponsSponsor and collaboratorsquery.spons=Pfizer
query.leadLead sponsor onlyquery.lead=Merck
query.locnLocation: country, state or cityquery.locn=Boston
query.termFull-text across all fieldsquery.term=BRCA1

Filter params (exact and structured filtering):

ParamValuesExample
filter.overallStatusRECRUITING, NOT_YET_RECRUITING, ENROLLING_BY_INVITATION, ACTIVE_NOT_RECRUITING, COMPLETED, SUSPENDED, TERMINATED, WITHDRAWN, UNKNOWN, plus five expanded-access values (AVAILABLE, NO_LONGER_AVAILABLE, TEMPORARILY_NOT_AVAILABLE, APPROVED_FOR_MARKETING, WITHHELD)filter.overallStatus=RECRUITING
filter.idsOne or more NCT IDsfilter.ids=NCT04280705
filter.geodistance(lat,long,radius) onlyfilter.geo=distance(39.0035,-77.1013,50mi)
filter.advancedEssie expression over indexed areasfilter.advanced=AREA[StartDate]RANGE[2024-01-01,2024-12-31]

Pagination and display params:

ParamNotes
pageSizeResults per page. Default 10, max 1000. A larger value is silently capped at 1000, not rejected
pageTokenCursor for the next page, taken from the previous response
fieldsComma-separated list of fields to return (reduces payload)
sortField and direction, e.g. sort=StartDate:desc, or sort=@relevance
countTotaltrue to include totalCount in the response

Multiple values for filter.overallStatus can be separated by a comma or a pipe. Both return the same count:

filter.overallStatus=RECRUITING,ACTIVE_NOT_RECRUITING
filter.overallStatus=RECRUITING|ACTIVE_NOT_RECRUITING

Phase and Study Type Are Not Top-Level Filters

This is the one that costs people an afternoon:

GET /api/v2/studies?query.cond=diabetes&filter.phase=PHASE3
HTTP 400  `filter.phase` is unknown parameter

filter.studyType returns the same 400. Neither parameter exists. Both go through filter.advanced, which takes an Essie expression over indexed areas:

filter.advanced=AREA[Phase](PHASE2 OR PHASE3)
filter.advanced=AREA[StudyType]INTERVENTIONAL
filter.advanced=AREA[Phase]PHASE3 AND AREA[StudyType]INTERVENTIONAL
filter.advanced=AREA[LocationCountry]"United States"
filter.advanced=AREA[StartDate]RANGE[2024-01-01,2024-12-31]

Phase takes six values: NA, EARLY_PHASE1, PHASE1, PHASE2, PHASE3, PHASE4. Study type takes three: INTERVENTIONAL, OBSERVATIONAL, EXPANDED_ACCESS. /stats/fieldValues/Phase returns the live record count behind each value, which is a fast way to sanity check a filter before you page through it. (Today: 238,480 studies with no phase assigned, 49,992 in Phase 3.)

Country is the same trap one step removed. filter.geo is real, but it only accepts distance(lat,long,radius), so a country name has to go in AREA[LocationCountry], quoted. A filter.geo=country:India style guess returns a 400 and, if your code swallows HTTP errors, a silently empty result set.

The Response Schema

The v2 response wraps each study in a protocolSection object with nested sub-sections. The key sections:

{
  "studies": [
    {
      "protocolSection": {
        "identificationModule": {
          "nctId": "NCT05123456",
          "briefTitle": "...",
          "officialTitle": "..."
        },
        "statusModule": {
          "overallStatus": "RECRUITING",
          "startDateStruct": {"date": "2024-01-15", "type": "ACTUAL"},
          "completionDateStruct": {"date": "2026-06-30", "type": "ESTIMATED"},
          "studyFirstSubmitDate": "2023-10-01"
        },
        "sponsorCollaboratorsModule": {
          "leadSponsor": {"name": "Merck Sharp & Dohme LLC", "class": "INDUSTRY"},
          "collaborators": []
        },
        "descriptionModule": {
          "briefSummary": "...",
          "detailedDescription": "..."
        },
        "conditionsModule": {
          "conditions": ["Non-Small Cell Lung Cancer"],
          "keywords": ["NSCLC", "immunotherapy"]
        },
        "designModule": {
          "studyType": "INTERVENTIONAL",
          "phases": ["PHASE3"],
          "enrollmentInfo": {"count": 450, "type": "ESTIMATED"}
        },
        "armsInterventionsModule": {
          "interventions": [
            {"type": "DRUG", "name": "Pembrolizumab", "description": "..."}
          ]
        },
        "eligibilityModule": {
          "eligibilityCriteria": "Inclusion criteria:\n...\nExclusion criteria:\n...",
          "sex": "ALL",
          "minimumAge": "18 Years",
          "maximumAge": "N/A"
        },
        "contactsLocationsModule": {
          "locations": [
            {
              "facility": "Memorial Sloan Kettering Cancer Center",
              "city": "New York",
              "state": "New York",
              "country": "United States",
              "status": "RECRUITING"
            }
          ]
        }
      }
    }
  ],
  "nextPageToken": "NF0g5AEBAQBZ...",
  "totalCount": 1247
}

A single study fetched by ID (/studies/NCT04280705) puts the same protocolSection at the top level, next to resultsSection, documentSection, derivedSection and a hasResults boolean. resultsSection only appears once the sponsor has posted results, so check hasResults before reaching into it. An NCT ID that does not exist returns 404.

Python: Pulling Phase 3 Trials for a Drug in Active Recruitment

import requests
import pandas as pd
import time

BASE = "https://clinicaltrials.gov/api/v2"

def search_trials(params, max_results=None):
    """
    Search ClinicalTrials.gov v2 API with cursor pagination.
    Returns a list of study dicts (full protocolSection).
    """
    all_studies = []
    page_params = {**params, "pageSize": 1000, "countTotal": "true"}

    while True:
        response = requests.get(f"{BASE}/studies", params=page_params)
        response.raise_for_status()
        data = response.json()

        studies = data.get("studies", [])
        all_studies.extend(studies)

        total = data.get("totalCount", 0)
        print(f"Fetched {len(all_studies)} / {total}")

        if max_results and len(all_studies) >= max_results:
            break

        next_token = data.get("nextPageToken")
        if not next_token:
            break

        page_params = {**params, "pageSize": 1000, "pageToken": next_token}
        time.sleep(0.3)

    return all_studies[:max_results] if max_results else all_studies

def extract_study_summary(study):
    """Flatten a v2 study record to a single row dict."""
    ps = study.get("protocolSection", {})
    id_mod = ps.get("identificationModule", {})
    status_mod = ps.get("statusModule", {})
    sponsor_mod = ps.get("sponsorCollaboratorsModule", {})
    design_mod = ps.get("designModule", {})
    conditions_mod = ps.get("conditionsModule", {})
    contacts_mod = ps.get("contactsLocationsModule", {})

    locations = contacts_mod.get("locations", [])
    countries = list({loc.get("country", "") for loc in locations})

    return {
        "nct_id":          id_mod.get("nctId"),
        "brief_title":     id_mod.get("briefTitle"),
        "status":          status_mod.get("overallStatus"),
        "start_date":      (status_mod.get("startDateStruct") or {}).get("date"),
        "completion_date": (status_mod.get("completionDateStruct") or {}).get("date"),
        "sponsor":         (sponsor_mod.get("leadSponsor") or {}).get("name"),
        "sponsor_class":   (sponsor_mod.get("leadSponsor") or {}).get("class"),
        "phases":          "|".join(design_mod.get("phases", [])),
        "study_type":      design_mod.get("studyType"),
        "enrollment":      (design_mod.get("enrollmentInfo") or {}).get("count"),
        "conditions":      "|".join(conditions_mod.get("conditions", [])),
        "location_count":  len(locations),
        "countries":       "|".join(countries),
    }

# All Phase 3 interventional pembrolizumab trials currently recruiting.
# Phase and study type must go through filter.advanced: filter.phase
# and filter.studyType do not exist and will return HTTP 400.
pembro_trials = search_trials({
    "query.intr":           "pembrolizumab",
    "filter.overallStatus": "RECRUITING",
    "filter.advanced":      "AREA[Phase]PHASE3 AND AREA[StudyType]INTERVENTIONAL",
})

df = pd.DataFrame([extract_study_summary(s) for s in pembro_trials])
print(f"Phase 3 pembrolizumab trials recruiting: {len(df)}")
print(df[["nct_id", "brief_title", "sponsor", "enrollment", "completion_date"]].head(10).to_string(index=False))

That query returned 125 studies on 8 October 2026, which is a useful order of magnitude to check your own run against.

Python: Getting All Trials for a Sponsor

# All trials sponsored by Moderna, any status
moderna_trials = search_trials({
    "query.spons": "Moderna",
})

df_moderna = pd.DataFrame([extract_study_summary(s) for s in moderna_trials])
print(f"\nTotal Moderna trials: {len(df_moderna)}")
print(df_moderna.groupby("status")["nct_id"].count().sort_values(ascending=False))

The query.spons parameter searches both lead sponsor and collaborators. Use query.lead if you want only lead sponsor matches.

Python: Building a Status Change Monitor

This is the most operationally useful pattern: watch a set of trials and alert when their status changes.

import json
from pathlib import Path
from datetime import datetime

SNAPSHOT_FILE = "trial_snapshot.json"

def fetch_trial(nct_id):
    """Fetch a single trial by NCT ID."""
    response = requests.get(f"{BASE}/studies/{nct_id}")
    if response.status_code == 404:
        return None
    response.raise_for_status()
    study = response.json()
    return extract_study_summary(study)

def check_for_status_changes(nct_ids, snapshot_path=SNAPSHOT_FILE):
    """
    Compare current trial statuses against a saved snapshot.
    Returns a list of trials that have changed status since last check.
    """
    # Load previous snapshot
    previous = {}
    if Path(snapshot_path).exists():
        with open(snapshot_path) as f:
            previous = json.load(f)

    current = {}
    changes = []

    for nct_id in nct_ids:
        trial = fetch_trial(nct_id)
        if not trial:
            continue
        current[nct_id] = trial

        prev = previous.get(nct_id, {})
        if prev.get("status") != trial["status"]:
            changes.append({
                "nct_id":       nct_id,
                "brief_title":  trial["brief_title"],
                "old_status":   prev.get("status", "UNKNOWN"),
                "new_status":   trial["status"],
                "checked_at":   datetime.utcnow().isoformat(),
            })
        time.sleep(0.2)

    # Save updated snapshot
    with open(snapshot_path, "w") as f:
        json.dump(current, f, indent=2)

    return changes

# List of trials you care about
watched_trials = [
    "NCT05094674",
    "NCT04516746",
    "NCT03668418",
]

changes = check_for_status_changes(watched_trials)
if changes:
    for c in changes:
        print(f"STATUS CHANGE: {c['nct_id']} | {c['old_status']} -> {c['new_status']} | {c['brief_title']}")
else:
    print("No status changes detected.")

fetch_trial works because the single-study endpoint returns protocolSection at the top level, so the same flattening function handles both shapes. Run this daily or weekly depending on how fast your portfolio moves, and you have a trial status alert without a commercial tracking subscription.

Useful Field Subsets

The full study record is large (the median is just under 10 KB, the 99th percentile is 150 KB). Use the fields parameter to request only what you need:

# Request only identification and status fields (much faster pagination)
response = requests.get(f"{BASE}/studies", params={
    "query.cond": "alzheimer",
    "filter.advanced": "AREA[Phase]PHASE3",
    "fields": "NCTId,BriefTitle,OverallStatus,StartDate,CompletionDate,LeadSponsorName",
    "pageSize": 1000,
})

The fields parameter takes the ClinicalTrials field names (the full list is published at https://clinicaltrials.gov/data-api/api). The response keeps the nested protocolSection shape but drops every module you did not ask for, so payloads fall by an order of magnitude on large pulls.

Three Actors That Return Flat Records

Everything above is code you now own: token pagination, the AREA[] syntax, module flattening, retries on 429. We publish three actors that already have it. All three read the same open API, all three are pay per result, and none needs an API key.

ClinicalTrials.gov Scraper

The broadest of the three, and the right default. It takes condition, searchTerm, intervention, sponsor, location, a status array, maxResults (default 100, up to 10,000) and includeLocations, and returns 26 flat fields per study.

Track a competitor’s pipeline:

{
  "sponsor": "Pfizer",
  "status": ["RECRUITING", "ACTIVE_NOT_RECRUITING"],
  "maxResults": 500
}

Each record carries nct_id, title, official_title, overall_status, study_type, phases, conditions, interventions, lead_sponsor, sponsor_class, collaborators, enrollment, enrollment_type, start_date, primary_completion_date, completion_date, last_update_date, sex, minimum_age, maximum_age, healthy_volunteers, brief_summary, location_count, location_countries, study_url and scraped_at. Group by phases and you have the shape of a competitor’s pipeline: how much sits in Phase 1 against Phase 3, and what is about to read out.

Monitor a whole indication rather than one company:

{
  "condition": "non-small cell lung cancer",
  "intervention": "pembrolizumab",
  "status": ["RECRUITING"],
  "includeLocations": true
}

includeLocations is off by default. Turn it on and every study gains a locations array with facility, city, state, country and per-site status, which is the input a recruitment or site-selection team actually needs. Filter by location with status: ["RECRUITING"] and the output is a sourced list of active sites.

Pricing is $0.0035 per study on the Apify Free plan and $0.0025 on Gold and above. Details and FAQs on the ClinicalTrials Scraper page, listing on Apify.

ClinicalTrials Bulk Exporter

Sixteen fields, and the one that matters is eligibility_criteria: the complete inclusion and exclusion block as text. It also goes widest per run, up to 10,000 trials pulled 1,000 per API page.

ParameterExampleNotes
condition"diabetes"Disease or condition keyword
intervention"semaglutide"Drug name or intervention
phase"PHASE2,PHASE3"Comma-separated, compiled into AREA[Phase]
status"RECRUITING,ACTIVE_NOT_RECRUITING"Comma-separated status values
country"United States"Location country name, not an ISO code
maxResults500Default 500, maximum 10,000

All parameters are optional. Running with just maxResults returns whatever the registry hands back first, so in practice you always want at least a condition or an intervention.

import requests, time

run = requests.post(
    "https://api.apify.com/v2/acts/themineworks~clinicaltrials-bulk-exporter/runs",
    headers={"Authorization": f"Bearer {APIFY_TOKEN}"},
    json={
        "condition": "non-small cell lung cancer",
        "phase": "PHASE3",
        "status": "RECRUITING",
        "maxResults": 200
    }
)

run_id = run.json()["data"]["id"]

while True:
    r = requests.get(
        f"https://api.apify.com/v2/actor-runs/{run_id}",
        headers={"Authorization": f"Bearer {APIFY_TOKEN}"}
    )
    if r.json()["data"]["status"] in ("SUCCEEDED", "FAILED"):
        break
    time.sleep(5)

trials = requests.get(
    f"https://api.apify.com/v2/actor-runs/{run_id}/dataset/items",
    headers={"Authorization": f"Bearer {APIFY_TOKEN}"}
).json()

# Use .get(): Apify drops null keys, so an observational trial with no
# phase comes back without a "phase" key at all.
for t in trials:
    print(t["nct_id"], t.get("phase"), t.get("overall_status"))
    print("  ", t.get("title"))

Each record returns nct_id, title, overall_status, phase, start_date, completion_date, conditions, interventions, lead_sponsor, enrollment, eligibility_criteria, minimum_age, maximum_age, sex, url and scraped_at. phase, conditions and interventions arrive as delimited strings rather than arrays, with interventions written as TYPE:Name and separated by semicolons.

import pandas as pd

df = pd.DataFrame(trials)
df["start_date"] = pd.to_datetime(df["start_date"], errors="coerce")

# Trials by phase
print(df.groupby("phase").size())

# Active recruiting trials
recruiting = df[df["overall_status"] == "RECRUITING"]
print(recruiting[["nct_id", "title", "start_date"]].head(20))

Pricing is $0.001 per trial on the Free plan and $0.0006 on Gold and above. Details on the ClinicalTrials Bulk Exporter page, listing on Apify.

ClinicalTrials Sponsor Intel

Sponsor-first, and aimed at finished work by default. The status default is COMPLETED,TERMINATED, which is the set you want when the question is what a company has actually run rather than what it is starting.

ParameterDefaultNotes
sponsor-Sponsor or collaborator name, partial match
condition-Narrow to a specific indication
phaseallPHASE1, PHASE2, PHASE3, PHASE4, comma-separated
statusCOMPLETED,TERMINATEDComma-separated status values
crossRefFDAtrueLook up the intervention against FDA drug approvals
maxResults100Maximum 2,000

Twelve fields come back: nct_id, title, phase, overall_status, start_date, completion_date, results_first_posted, lead_sponsor, interventions, fda_approved, url and scraped_at.

results_first_posted is the quiet one. A completed trial with results posted and a completed trial with nothing posted two years on are very different signals, and that field is how you separate them.

fda_approved is a boolean set from an openFDA drugsfda lookup on the trial’s first intervention name. Treat it as a hint rather than a regulatory record: matching a trial intervention string against an FDA brand name is approximate, and the output carries no approval date.

run = requests.post(
    "https://api.apify.com/v2/acts/themineworks~clinicaltrials-sponsor-intelligence/runs",
    headers={"Authorization": f"Bearer {APIFY_TOKEN}"},
    json={
        "sponsor": "Pfizer",
        "condition": "oncology",
        "status": "COMPLETED,TERMINATED",
        "crossRefFDA": True,
        "maxResults": 100
    }
)

Poll and fetch the dataset exactly as in the bulk example above, then look at the shape of the portfolio:

import pandas as pd

df = pd.DataFrame(trials)
df["completion_date"] = pd.to_datetime(df["completion_date"], errors="coerce")

# How much of the finished book was terminated rather than completed
print(df.groupby(["phase", "overall_status"]).size().unstack(fill_value=0))

# Completed trials with no results posted
silent = df[(df["overall_status"] == "COMPLETED") & (df["results_first_posted"].isna())]
print(f"{len(silent)} of {len(df)} finished trials have no results posted")

Pricing is $0.001 per trial on the Free plan and $0.0006 on Gold and above, and the FDA lookup costs nothing extra. Details on the ClinicalTrials Sponsor Intel page, listing on Apify.

All three also charge one actor-start event of $0.005 per run. A search that matches nothing delivers no records, so it charges no per-result events.

Which of the Three to Use

What you needUse
Broad search, the most fields, per-site detailClinicalTrials.gov Scraper
Full eligibility criteria text, up to 10,000 records in one runClinicalTrials Bulk Exporter
One company’s finished trials, results posting and an FDA flagClinicalTrials Sponsor Intel

They all read the same registry, so the same NCT ID turns up in all three with different columns around it. If the answer is not obvious, start with the scraper: it has the widest filter set (condition, free text, intervention, sponsor, location, status) and the deepest record. Move to the bulk exporter when you need the inclusion and exclusion wording, for feasibility screening or for building a corpus. Move to sponsor intel when the unit of analysis is a company rather than a disease.

Use Cases

Pharma competitive intelligence and due diligence. Track a competitor’s pipeline by sponsor name, sorted by phase and expected completion date. Before licensing or acquiring a candidate, the same pull with status: COMPLETED,TERMINATED gives you the sponsor’s track record rather than their pitch.

Patient recruitment and site selection. Parse the locations array and keep the sites whose own status is RECRUITING. For a condition and a geography, that is a sourced list of active sites in a day.

Disease-area mapping and market sizing. For one indication, pull every trial by phase and status to see how crowded it is and where trials are stalling (a large ACTIVE_NOT_RECRUITING or TERMINATED block). Counting recruiting trials by country shows where the clinical activity, and the patients, actually sit.

CRO business development. Filter by sponsor class INDUSTRY, phase and status to find sponsors running studies in therapeutic areas you serve, before they have finished enrollment.

Publication and results tracking. Registered trials are increasingly required to report results. hasResults and resultsFirstPostDateStruct tell you who has and who has not, which matters for systematic reviews and for judging a sponsor’s reporting discipline.

Academic and investor analysis. Export the full COMPLETED cohort for a condition and look at completion rates, timelines and enrollment against plan. Across a portfolio, 40 completed trials with 3 approvals tells a different story from 10 with 4.

FAQ

Do I need an API key? No. The ClinicalTrials.gov v2 API is fully open: no key, no OAuth, no approval process.

How current is the data? The registry reloads once a day and /api/v2/version reports the exact load timestamp. Every actor run reads the live API, so newly posted and updated trials show up on the next load.

Why does my phase filter return a 400? Because filter.phase does not exist. Use filter.advanced=AREA[Phase]PHASE3. Same for study type with AREA[StudyType].

Can I get every study site? Yes. Set includeLocations on the scraper and each study carries facility, city, state, country and status for every location.

Which statuses can I filter by? The scraper takes nine: recruiting, not yet recruiting, enrolling by invitation, active (not recruiting), completed, suspended, terminated, withdrawn and unknown. The API itself also returns five expanded-access statuses, which you can reach by calling filter.overallStatus directly.

Can I monitor a drug or sponsor over time? Yes. Save the search as an Apify task and schedule it, and each run returns the current matching set. The Python monitor above does the same job straight against the API.

Related Actor

Explore the scraper referenced in this article: inputs, outputs, and pricing, then run it on Apify.

Apify Store

Find a ready-made scraper for your job

Apify has over 80,000 scrapers and automations, ours included. Start free with $5 of platform credit every month.