The citation graph, without a subscription
The Academic Papers API returns scholarly-paper search and metadata as clean JSON.
12 active endpoints, on 1 and 2 credit tiers.
- POST/academic/v1/search
- POST/academic/v1/paper_detail
- POST/academic/v1/citations
- POST/academic/v1/references
- POST/academic/v1/related
- POST/academic/v1/author
- POST/academic/v1/institution
- +5 more
What Academic Papers endpoints does ReefAPI ship?
12 live read endpoints. Read-only data API: no writes, no account actions, no dashboard access on the target site.
Academic Papers API
3 of 12 endpoints, ready to run
Papers matching a query with authors, affiliations, venue, year, citation count and open-access status, filterable and sortable.
{ "ok": true, "meta": { "api": "academic", "endpoint": "search", "mode": "live", "latency_ms": 1825.9, "record_count": 10, "cache_hit": false, "completeness_pct": 100 }, "data": { "results": [ { "id": "W4323655724", "doi": "10.1016/j.lindif.2023.102274", "title": "ChatGPT for good? On opportunities and challenges of large language models for education", "abstract": null, "authors": [ { "name": "Enkelejda Kasneci", "id": "A5008809634", "orcid": "https://orcid.org/0000-0003-3146-4484", "position": "first", "is_corresponding": true, "affiliations": [ { "name": "Technical University of Munich", "id": "I62916508", "ror": "https://ror.org/02kkvpp62", "country": "DE", "type": "education" } ] }, { "name": "Kathrin Seßler", "id": "A5013346660", "orcid": "https://orcid.org/0000-0002-3380-4641", "position": "middle", "is_corresponding": false, "affiliations": [ { "name": "Technical University of Munich", "id": "I62916508", "ror": "https://ror.org/02kkvpp62", "country": "DE", "type": "education" } ] }, { "name": "Stefan Küchemann", "id": "A5050063899", "orcid": "https://orcid.org/0000-0003-2729-1592", "position": "middle", "is_corresponding": false, "affiliations": [ { "name": "Ludwig-Maximilians-Universität München", "id": "I8204097", "ror": "https://ror.org/05591te55", "country": "DE", "type": "education" } ] } ], "author_count": 23, "venue": { "name": "Learning and Individual Differences", "id": "S9267903", "issn_l": "1041-6080", "type": "journal", "publisher": "Elsevier BV", "is_oa": false, "is_in_doaj": false }, "year": 2023, "publication_date": "2023-03-09", "type": "article", "language": "en", "is_oa": true, "oa_status": "green", "oa_url": "https://epub.ub.uni-muenchen.de/125071/1/ChatGPT_for_Good_v3.pdf", "pdf_url": "https://epub.ub.uni-muenchen.de/125071/1/ChatGPT_for_Good_v3.pdf", "cited_by_count": 6232, "reference_count": 74, "fields_of_study": [ "Artificial Intelligence in Healthcare and Education", "Topic Modeling", "Online Learning and Analytics" ], "biblio": { "volume": "103", "issue": null, "first_page": "102274", "last_page": "102274" }, "is_retracted": false, "ids": { "openalex": "W4323655724", "doi": "10.1016/j.lindif.2023.102274", "pmid": null, "mag": null }, "source": "openalex", "sources_merged": [ "openalex" ] }, { "id": "W4384561707", "doi": "10.1038/s41591-023-02448-8", "title": "Large language models in medicine", "abstract": null, "authors": [ { "name": "Arun James Thirunavukarasu", "id": "A5018942914", "orcid": "https://orcid.org/0000-0001-8968-4768", "position": "first", "is_corresponding": false, "affiliations": [ { "name": "University of Cambridge", "id": "I241749", "ror": "https://ror.org/013meh722", "country": "GB", "type": "education" } ] }, { "name": "Darren Shu Jeng Ting", "id": "A5010465757", "orcid": "https://orcid.org/0000-0003-1081-1141", "position": "middle", "is_corresponding": true, "affiliations": [ { "name": "University of Nottingham", "id": "I142263535", "ror": "https://ror.org/01ee9ar58", "country": "GB", "type": "education" }, { "name": "Singapore National Eye Center", "id": "I2799299286", "ror": "https://ror.org/029nvrb94", "country": "SG", "type": "healthcare" }, { "name": "Singapore Eye Research Institute", "id": "I4210116917", "ror": "https://ror.org/02crz6e12", "country": "SG", "type": "facility" } ] }, { "name": "Kabilan Elangovan", "id": "A5073732346", "orcid": "https://orcid.org/0000-0002-7711-7368", "position": "middle", "is_corresponding": false, "affiliations": [ { "name": "Singapore National Eye Center", "id": "I2799299286", "ror": "https://ror.org/029nvrb94", "country": "SG", "type": "healthcare" }, { "name": "Singapore Eye Research Institute", "id": "I4210116917", "ror": "https://ror.org/02crz6e12", "country": "SG", "type": "facility" } ] } ], "author_count": 6, "venue": { "name": "Nature Medicine", "id": "S203256638", "issn_l": "1078-8956", "type": "journal", "publisher": "Nature Portfolio", "is_oa": false, "is_in_doaj": false }, "year": 2023, "publication_date": "2023-07-17", "type": "article", "language": "en", "is_oa": false, "oa_status": "closed", "oa_url": null, "pdf_url": null, "cited_by_count": 3728, "reference_count": 101, "fields_of_study": [ "Artificial Intelligence in Healthcare and Education", "Machine Learning in Healthcare", "Topic Modeling" ], "biblio": { "volume": "29", "issue": "8", "first_page": "1930", "last_page": "1940" }, "is_retracted": false, "ids": { "openalex": "W4384561707", "doi": "10.1038/s41591-023-02448-8", "pmid": "37460753", "mag": null }, "source": "openalex", "sources_merged": [ "openalex" ] }, { "id": "W4384071683", "doi": "10.1038/s41586-023-06291-2", "title": "Large language models encode clinical knowledge", "abstract": "Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias. In addition, we evaluate Pathways Language Model 1 (PaLM, a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM 2 on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA 3 , MedMCQA 4 , PubMedQA 5 and Measuring Massive Multitask Language Understanding (MMLU) clinical topics 6 ), including 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%. However, human evaluation reveals key gaps. To resolve this, we introduce instruction prompt tuning, a parameter-efficient approach for aligning LLMs to new domains using a few exemplars. The resulting model, Med-PaLM, performs encouragingly, but remains inferior to clinicians. We show that comprehension, knowledge recall and reasoning improve with model scale and instruction prompt tuning, suggesting the potential utility of LLMs in medicine. Our human evaluations reveal limitations of today’s models, reinforcing the importance of both evaluation frameworks and method development in creating safe, helpful LLMs for clinical applications.", "authors": [ { "name": "Karan Singhal", "id": "A5027454515", "orcid": "https://orcid.org/0000-0001-9002-7490", "position": "first", "is_corresponding": true, "affiliations": [ { "name": "Google (United States)", "id": "I1291425158", "ror": "https://ror.org/00njsd438", "country": "US", "type": "company" } ] }, { "name": "Shekoofeh Azizi", "id": "A5047463591", "orcid": "https://orcid.org/0000-0002-7447-6031", "position": "middle", "is_corresponding": true, "affiliations": [ { "name": "Google (United States)", "id": "I1291425158", "ror": "https://ror.org/00njsd438", "country": "US", "type": "company" } ] }, { "name": "Tao Tu", "id": "A5059213795", "orcid": "https://orcid.org/0000-0003-3420-7889", "position": "middle", "is_corresponding": false, "affiliations": [ { "name": "Google (United States)", "id": "I1291425158", "ror": "https://ror.org/00njsd438", "country": "US", "type": "company" } ] } ], "author_count": 32, "venue": { "name": "Nature", "id": "S137773608", "issn_l": "0028-0836", "type": "journal", "publisher": "Nature Portfolio", "is_oa": false, "is_in_doaj": false }, "year": 2023, "publication_date": "2023-07-12", "type": "article", "language": "en", "is_oa": true, "oa_status": "hybrid", "oa_url": "https://www.nature.com/articles/s41586-023-06291-2.pdf", "pdf_url": "https://www.nature.com/articles/s41586-023-06291-2.pdf", "cited_by_count": 3689, "reference_count": 91, "fields_of_study": [ "Topic Modeling", "Artificial Intelligence in Healthcare and Education", "Natural Language Processing Techniques" ], "biblio": { "volume": "620", "issue": "7972", "first_page": "172", "last_page": "180" }, "is_retracted": false, "ids": { "openalex": "W4384071683", "doi": "10.1038/s41586-023-06291-2", "pmid": "37438534", "mag": null }, "source": "openalex", "sources_merged": [ "openalex" ] } ], "count": 4336195, "next_cursor": null, "source": "openalex" } }
How the Academic Papers API works
Academic Papers is a normal ReefAPI surface — the same four rules that hold for every other engine on the key.
No OAuth app, no request signing, no per-site account. One key covers all 184 engines.
Every route is a POST with a JSON body. Parameters are validated against the published schema before anything is charged.
Credits, not seats. Failed and blocked calls are never charged, and cache hits cost nothing.
One envelope everywhere. meta carries latency_ms, record_count and the endpoint that answered.
Following an idea forwards instead of backwards
References tell you what a paper was built on. Citations tell you what was built on it, which is the direction that answers 'is this still where the field is'.
{"query": "large language models", "sort": "cited_by_count:desc", "from_year": 2023}The anchor papers. A query on this scale returned a match count in the millions, so sorting and the year filter are doing the real work.
{"id": "W3177828909"}Then forward through the graph. Paging is cursor-based, which is what makes a deep traversal possible rather than a first page.
Authors arrive with ORCIDs and institutional affiliations attached, so a co-authorship or institution analysis needs no name disambiguation of your own.
curl -X POST https://api.reefapi.com/academic/v1/search \
-H "x-api-key: $REEF_KEY" \
-H "content-type: application/json" \
-d '{"query":"covid"}'{
"ok": true,
"data": { … },
"meta": {
"api": "academic",
"endpoint": "search",
"mode": "live",
"latency_ms": …,
"record_count": …
},
"error": null
}Identifiers, sources and the values each one really returns
Three backends sit behind this engine and they do not share an id scheme, a record granularity or a notion of what counts as a paper. Everything below came from live calls, mostly on the AlphaFold paper (OpenAlex W3177828909, DOI 10.1038/s41586-021-03819-2) and a 40-row CRISPR search.
| Input or field | Measured value | What it means |
|---|---|---|
| id on paper_detail, citations, references | W3177828909, 10.1038/s41586-021-03819-2, pmid:34265844 | OpenAlex W id, DOI, prefixed PMID or arXiv id. All three of those returned the same AlphaFold record |
| a bare PubMed number | 34265844 | rejected with NOT_FOUND. The pmid: prefix is required. arXiv ids keep their arXiv: prefix and resolved to OpenAlex W4298187611 |
| ids object on the result | openalex, doi, pmid, mag | mag is the retired Microsoft Academic id, still carried where OpenAlex has one. There is no PMCID field |
| authors[].orcid | an https://orcid.org/ URL, when the author registered one | a full URL, not a bare ORCID. 28 of the paper's 34 authors had one; 33 of 34 had an OpenAlex author id |
| authors[].affiliations[].ror | https://ror.org/00971b260 | ROR URL, plus the institution's OpenAlex id, country code and type such as company or education |
| oa_status | gold, hybrid, bronze, green, closed | across 40 CRISPR hits: 13 gold, 11 closed, 9 bronze, 5 green, 2 hybrid. is_oa was true on 29 of the 40 |
| oa_url and pdf_url | https://www.nature.com/articles/s41586-021-03819-2.pdf | direct file link where one is known. 11 of the same 40 rows had no pdf_url |
| source=arxiv row ids | arXiv:2209.15001 | arXiv preprints have no DOI on this path, and venue reads arXiv |
| source=crossref row ids | 10.7717/peerj-cs.3254/fig-10 | Crossref registers component DOIs, so individual figures and tables come back as rows next to real articles |
| cursor | send * on the first page | next_cursor stays null unless you pass cursor. Without it you get one page and no way forward |
| venue profile extras | Nature: issn_l 0028-0836, h_index 1853, apc_usd 12690 | venue also reports is_oa, is_in_doaj, publisher, country_code and works_count |
Citation totals do not agree with themselves, and that is expected. The AlphaFold paper reported cited_by_count 46,735 on its own record while the citations action reported total 45,854 rows for the same work. The first is OpenAlex's denormalized counter, the second is a live count of the citing-works index. Pick one and stay with it rather than reconciling the two.
What the open record contains, and where it thins
Measured on a high-profile paper and on broad searches.
Each author comes with a stable id, an ORCID where they have one, their position in the author list and their affiliations. Name-based author matching is the classic way to get bibliometrics wrong, and it is unnecessary here.
Every paper reports whether it is open access, which flavour, and a URL to the free full text when there is one. That is the difference between a citation you can read and one you can only cite.
Search rows often return a null abstract, because the index stores abstracts in an inverted form that has to be reconstructed. The detail endpoint fills it. If you need abstracts across a result set, budget a detail call per paper rather than expecting them on the list.
A broad query matched millions of works. Sorting by citation count, restricting the year and filtering on open access are what turn that into a result set. A bare query is a way to receive the first twenty-five of four million.
Coverage is strongest where publishers deposit metadata openly and thinner for some humanities venues, older material and non-English work. It is the largest freely-queryable scholarly index rather than every paper ever written.
What people build with Academic Papers
The jobs this data is most often used for.
endpoints
credits per call
Research tools call search then citations to map the literature around a topic.
Literature-review products use references and related to trace a paper's intellectual lineage.
Academic apps use author and institution to build researcher and organization profiles.
What Academic Papers data costs
The cheapest call here is 1 credit, so $15/mo (Pro) buys 10,000 of them — $1.50 per 1,000 credits. Credits roll over and never expire, and failed or blocked calls are not charged.
Full pricing →- 1,000 free credits on signup, no card
- One key, all 184 APIs, one credit pool
- Failed and blocked calls are never charged
- Credits roll over and never expire
Call it in two lines
Sign up, get 1,000 credits and one key that works on every engine. Then this is the whole protocol.
curl -X POST https://api.reefapi.com/academic/v1/search \
-H "x-api-key: $REEF_KEY" \
-H "content-type: application/json" \
-d '{"query":"covid"}'import requests
r = requests.post(
"https://api.reefapi.com/academic/v1/search",
headers={"x-api-key": REEF_KEY},
json={
"query": "covid"
},
)
print(r.json()["data"])Have a question? We got answers.
The questions people actually ask before wiring up Academic Papers.
Get a free key →Which identifier formats can I pass as id?▾
An OpenAlex work id starting with W, a DOI starting with 10., a prefixed arXiv id, or a PubMed id written as pmid:34265844. The prefix on the last two is not optional: W3177828909, the DOI and pmid:34265844 all returned the same AlphaFold record, while the bare number 34265844 was rejected with NOT_FOUND. arXiv:2209.15001 resolved through to OpenAlex W4298187611. The response echoes every id it knows in an ids object, so you only have to do this once per paper.
Why do cited_by_count and the citations total disagree?▾
They are two different measurements of the same thing. cited_by_count is a stored counter on the work record and read 46,735. The citations action filters the works index for papers citing that id and reported total 45,854. OpenAlex refreshes the counter and the index on different schedules, so a gap of a percent or two is normal. Use cited_by_count when you want a number to display and the citations total when you are about to page through the actual citing papers, because that is the number of rows you will get.
Do I get ORCIDs, and how good is author disambiguation?▾
You get an ORCID when OpenAlex has one, as a full https://orcid.org/ URL. On a 34-author paper, 28 authors carried one and 33 carried an OpenAlex author id, leaving one author who is a name string with id null and orcid null. That is what an unresolved author looks like, and it means you cannot key on authors[].id without a null check. Author profiles also return an affiliations list with a years array per institution, which is a clustering result rather than a verified employment record, so check the years before you treat an entry as current.
Why is abstract null on so many results?▾
OpenAlex stores abstracts as an inverted index and only for records where a publisher allowed it, so a large share simply have none. In a 40-row CRISPR search, 18 rows had no abstract while all 40 had a DOI. paper_detail can close some of that gap: fill_abstract is on by default and pulls the missing text from arXiv or Crossref, which is why the AlphaFold record returned a 1,701-character abstract and reported sources_merged as openalex plus crossref. Search results do not get that treatment, so hydrate the ones you care about.
What do the oa_status values mean, and where is the file?▾
gold means the venue itself is fully open, hybrid means an open article inside a paywalled journal, green means a repository copy exists, bronze means free to read on the publisher site with no open license, and closed means no free copy is known. The measured spread over 40 rows was 13 gold, 11 closed, 9 bronze, 5 green and 2 hybrid. When a file is known, oa_url and pdf_url carry the direct link, and license is returned separately: the AlphaFold record listed a CC BY 4.0 URL. bronze in particular tends to have a readable URL and no reusable license, so read license before you redistribute anything.
Why did a Crossref search return something called Figure 10?▾
Because Crossref registers DOIs for components of an article, not only for the article, and those component records are indexed alongside everything else. A live search with source crossref returned 10.7717/peerj-cs.3254/fig-10 and 10.7717/peerj-cs.3655/fig-2 in the top rows, with venue null on both. The tell is the trailing path segment on the DOI and the missing venue. If you want papers only, use source openalex, which returned real articles with a venue object on every row.
What does source=all actually give me?▾
One deduplicated page assembled from all three backends, not a merged corpus. A call with per_page 3 and source all returned 9 rows and reported count 9, because it fetched a page from each backend and DOI-deduplicated the union. That count is the size of what you got, unlike the single-source paths where count is the corpus total, 257,382 for the same query on arXiv and 315,449 on Crossref. Each row carries sources_merged so you can see which backend it came from.
How do I page past the first screen of results?▾
Pass cursor with the literal value * on the first call, then feed the returned next_cursor back in. Without a cursor, next_cursor comes back null and there is no continuation token, which is the usual reason a scripted crawl stops after one page. per_page goes up to 200. The offset-style parameters are backend-specific and cap at 10,000: start applies to arXiv and offset to Crossref, and neither has any effect on the OpenAlex path.
What is the Academic Papers API?▾
Academic Papers API is a ReefAPI endpoint group for search scholarly papers, authors, citations and abstracts. It returns live JSON through POST requests under /academic/v1.
Is the Academic Papers API free to try?▾
Yes. ReefAPI starts with 1,000 free credits, no card required. Academic Papers calls use the same shared credit balance as every other ReefAPI engine.
Do I need an Academic Papers login or account?▾
No login to Academic Papers is needed for the API response. You call ReefAPI with your x-api-key header, and the playground can run live examples before you create a production key.
How fresh is the Academic Papers data?▾
The page example is captured from a live search call, and production requests fetch live data through ReefAPI rather than a static sample.
How many credits does the Academic Papers API use?▾
Academic Papers actions currently cost 1-2 credits per successful call. Failed or blocked calls are free, and all APIs draw from one credit pool.
Can I call Academic Papers from an AI assistant or MCP client?▾
Yes. Connect ReefAPI once through MCP and your assistant can call academic actions with the same key, credit pool and JSON envelope used by normal REST requests.
25 Media, Film & Knowledge APIs on the same key
One key, one credit pool, one response envelope. If you are pulling Academic Papers, you are one call away from the rest of the category — no second contract, no second integration.
Need something this API does not do?
Name the endpoint, the field, or a source we do not carry yet. We ship new APIs every week and you would be first to get the key. Real people read every message and reply the same day.
Try it on your own data before you pay anything
The call above is the real endpoint, not a recording. A free key gives you 1,000 credits, the other 183 APIs, and the same envelope everywhere.
Endpoints, parameters and credit costs on this page are read from the live catalog and cannot drift from what the API accepts. Field notes were captured on 2026-08-30.