Five Scholarly Databases in One Call: OpenAlex, Crossref, arXiv, PubMed and OpenCitations
A literature scan means five sources with five query languages and five response shapes. Here is how to pull published works, canonical metadata, preprints, biomedical hits and a citation graph for one topic in a single call, and which source to trust for what.
The actor referenced in this article. Pay only for results delivered.
A proper literature scan is not one search. It is five, each in a different place, each speaking a different language and handing you back a different shape of data.
Try it live: OpenAlex Scraper, 20 Fields, Citations & OA, No API Key. Pay per result delivered. Failed and empty results are never charged.
OpenAlex for published work. Crossref for the canonical record. arXiv for what has not been published yet. PubMed if it is biomedical. OpenCitations when you need to know who actually cited what. Doing all five by hand is why “just do a quick literature review” is never quick.
Here is the same scan in one call, with real output.
One topic, five sources
// tools/call → literature_report
{ "topic": "CRISPR base editing", "from_year": 2022 }
Four sections came back at 25 records each: published works from OpenAlex, canonical metadata from Crossref, arXiv preprints, and PubMed. Crossref returned “CRISPR-Mediated Base Editing Tools: A Genome-Wide Survey” near the top, which is exactly the record you would want.
The part that saves you: knowing which source to trust
The value of having all five side by side is that you learn which one to believe for which question.
On that run, OpenAlex ranked by citation count with fairly loose matching and returned some high-citation papers that were adjacent to the topic rather than squarely on it. Crossref matched the query much more tightly. So the honest reading is: use Crossref when you want precision, and treat the OpenAlex section as a “what is big in this neighbourhood” signal rather than a targeted search. That is not a flaw to hide, it is the kind of thing you can only see when the sources sit next to each other.
The short version: Crossref is the system of record for what a work is, the metadata every publisher registers. OpenAlex is where you go to understand how works relate, who cites whom, what is influential and what is free to read. The other three each answer one narrower question.
| Question | Source |
|---|---|
| What is the canonical DOI, registered title and journal of record? | Crossref |
| What is most cited, what is open access, and where are the abstracts? | OpenAlex |
| What has been posted but not yet published? | arXiv |
| What does the biomedical literature say, with MeSH terms? | PubMed |
| Who cited this paper, and what does it cite? | OpenCitations |
Many pipelines use Crossref to confirm a paper’s identity and OpenAlex to find and enrich everything around it.
The hard half: following a claim forward
Finding papers is the easy half. The hard half is knowing whether a striking result held up, and that lives in the citation record, not in the paper.
The citation-graph tool takes a DOI and returns who cites it and what it cites:
// tools/call → citation_graph
{ "dois": ["10.1038/nature12373"], "direction": "citations" }
That returned 20 citation records, each a paper that built on the original. Walking the forward trail tells you whether a finding is doing real work in the field or is merely being cited in the same shallow way over and over. A paper will never tell you it failed to replicate. The papers that came after it will.
Going deeper on one source
The one-call report is a scan: 25 records a section. When you need a bigger, filtered corpus from one source, run the underlying scraper directly. OpenAlex is usually the one worth going deep on, since it indexes over 250 million works. A literature review pulls the most-cited recent work on a topic, with abstracts:
{
"searchTerm": "perovskite solar cells",
"fromYear": 2022,
"minCitations": 25,
"includeAbstract": true,
"maxResults": 300
}
Results come back most-cited first, so the anchor papers sit at the top. minCitations is the simplest noise filter, since it keeps the work the field has actually engaged with rather than every paper that mentions the keyword. OpenAlex stores abstracts as an inverted index, and includeAbstract rebuilds them into readable text.
Drop the citation floor and the same records map a field. Each work carries author_institutions and concepts, so counting them shows who is publishing on a technology and where, which is the raw material for R&D landscaping, partner scouting and talent searches. Counting publication_year gives a sourced answer to whether a field is heating up or cooling down.
import os
from collections import Counter
from apify_client import ApifyClient
apify = ApifyClient(os.environ["APIFY_TOKEN"])
def search_openalex(query: str, from_year: int = 2020, min_citations: int = 0,
max_results: int = 100, abstracts: bool = True) -> list[dict]:
run = apify.actor("themineworks/openalex-scholarly-works").call(run_input={
"searchTerm": query,
"fromYear": from_year,
"minCitations": min_citations,
"includeAbstract": abstracts,
"maxResults": max_results,
})
items = apify.dataset(run["defaultDatasetId"]).iterate_items()
return [w for w in items if w.get("_type") != "summary"]
works = search_openalex("perovskite solar cells", from_year=2018,
max_results=5000, abstracts=False)
by_institution = Counter(i for w in works for i in w.get("author_institutions") or [])
by_year = Counter(w["publication_year"] for w in works)
print(by_institution.most_common(10))
print(sorted(by_year.items()))
One catch on the trend line: because results are sorted by citations, a capped sample leans toward older papers that have had time to collect them. Set maxResults high enough to cover the whole result set (it goes up to 10,000), or run one query per year. Set openAccessOnly when you want a corpus you can read in full.
Citation-grounded answers over the abstracts
With abstracts in hand you can build a small retrieval layer that answers research questions and cites the papers it used. Embed the abstracts, store them with their DOIs, and require the model to reference every source.
import anthropic
import chromadb
from openai import OpenAI
openai_client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
claude = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
chroma = chromadb.PersistentClient(path="./research_chroma")
collection = chroma.get_or_create_collection("papers")
def index_papers(papers: list[dict]):
docs, ids, metas = [], [], []
for p in papers:
abstract = p.get("abstract") or ""
if len(abstract) < 100:
continue # nothing useful to embed
docs.append(f"{p.get('title')}\n\n{abstract}")
ids.append(p.get("doi") or p["openalex_id"])
metas.append({
"title": p.get("title") or "",
"year": p.get("publication_year") or 0,
"doi": p.get("doi") or "",
"citations": p.get("cited_by_count") or 0,
})
if not docs:
return
embeddings = [e.embedding for e in openai_client.embeddings.create(
model="text-embedding-3-small", input=docs).data]
collection.upsert(documents=docs, embeddings=embeddings, ids=ids, metadatas=metas)
def ask(question: str, top_k: int = 6) -> str:
q_emb = openai_client.embeddings.create(
model="text-embedding-3-small", input=[question]).data[0].embedding
hits = collection.query(query_embeddings=[q_emb], n_results=top_k,
include=["documents", "metadatas"])
context = "\n\n---\n\n".join(
f"[{m['title']} ({m['year']}), DOI {m['doi']}, {m['citations']} citations]\n{d}"
for d, m in zip(hits["documents"][0], hits["metadatas"][0])
)
resp = claude.messages.create(
model="claude-sonnet-4-6",
max_tokens=1200,
messages=[{"role": "user", "content": (
"Answer using only these paper abstracts. Cite every claim as "
"[Title (Year), DOI]. If the abstracts do not cover it, say so.\n\n"
f"QUESTION: {question}\n\nABSTRACTS:\n{context}"
)}],
)
return resp.content[0].text
index_papers(search_openalex("retrieval augmented generation",
from_year=2022, min_citations=20, max_results=100))
print(ask("What chunking strategies have been shown to improve RAG retrieval quality?"))
Abstracts are enough for literature mapping, gap analysis and “what does the field say about X” questions, because they are written to summarise the contribution. For deep methodological detail you will want the full text of the key papers. Use the abstracts to find which ones are worth fetching, then follow the oa_url where a paper is open access, and respect the publisher’s terms where it is not.
Why one endpoint beats five tabs
Each of these sources has its own scraper and its own tutorial: OpenAlex, Crossref DOI metadata, the arXiv preprint scraper, PubMed, and the OpenCitations citation graph. Running them one at a time and reconciling five formats is the slow way.
The Academic Research MCP server puts all five behind a single endpoint your AI agent can call, so Claude pulls the whole scan in one step and reasons over it. It is the cheapest server we run, because every source is a free public API with no proxy fight and no login. You pay per result, and a topic that returns nothing adds no result charges.
Honest limits
These are metadata and abstract services, not full-text retrieval. Citation counts differ between sources because their corpora differ. PubMed is biomedical only, and the report marks that section not-found for out-of-domain topics rather than padding it. And as above, OpenAlex relevance is loose on narrow queries, which is exactly why having Crossref in the same call matters.
Ready to try it? The Academic Research MCP runs on Apify with your own account. Setup for Claude, Cursor and Windsurf is in the listing.
Explore the scraper referenced in this article: inputs, outputs, and pricing, then run it on Apify.
Frequently asked questions
Is this full-text search or metadata? +
Metadata and abstracts, not full text. You get titles, authors, venues, citation counts, DOIs, abstracts and MeSH terms, plus PDF links where the source publishes them (always for arXiv, and for open-access records on OpenAlex). It is for finding and mapping the literature, not downloading every paper.
Why do OpenAlex and Crossref disagree on citation counts? +
Because their corpora differ. That is the sources disagreeing, not a bug. Use Crossref for the canonical record and precise matching, and treat OpenAlex as a broader signal of what is influential in a neighbourhood.
What does one literature scan cost? +
This is the cheapest workflow we run because every source here is a free public API with no proxy cost. You pay per result delivered, a topic that returns nothing is free, and the bundled report costs less than calling the tools one by one.
Can I check whether a finding actually replicated? +
Yes. The citation-graph tool takes a DOI and returns who cited it and what it cites, from OpenCitations. Following the forward trail tells you whether a result was built on or merely repeated, which the paper itself will never tell you.
Can this replace Scopus or Web of Science? +
For a lot of work, yes. OpenAlex coverage and citation data hold up well against the commercial indexes in most fields, and Crossref is the registry publishers themselves use. The paid tools add curated subject taxonomies and some proprietary metrics. If you need those, keep the seat. For discovery, citation analysis and grounding an AI agent, the open stack is enough.
How do I keep a research index current? +
Re-run the OpenAlex search on a schedule with a recent fromYear, use each paper's DOI as its id, and skip any id already in your vector store. New papers get embedded and added, existing ones are never re-embedded, so a weekly refresh stays cheap.
OpenAlex API: 250 Million Research Papers, Free, No Rate-Limit Workarounds Needed
OpenAlex replaced the defunct Microsoft Academic Graph with 250M+ scholarly works. The API is free, well-documented, and returns structured data including citations and author affiliations.
Build a Clinical Trial Pipeline Tracker with the ClinicalTrials.gov Scraper
Track any drug, sponsor or indication across ClinicalTrials.gov as structured JSON, with phases, sponsors, enrollment and trial sites.
Amazon Retail Price vs AliExpress Supplier Cost: The Sourcing Check in One Call
Pull an Amazon listing's retail price and reviews against real AliExpress supplier prices in one call, and read what the sourcing spread tells you.
Building a Real Estate Investor Deal-Sourcing Pipeline with Claude and Redfin
How to wire the Redfin Scraper into Claude as an MCP tool so an investor can ask for undervalued properties by price, beds, and metro in plain language.