Does Google Scholar Have an API? Options for Research Agents
Google does not document a public Google Scholar API. Its official help says bulk access is unavailable and asks automated software to respect Scholar's robots.txt. For a research agent, the useful replacement depends on the job: use a documented scholarly graph for paper discovery, Crossref for DOI and publisher-deposited metadata, a third-party SERP service only when Scholar-specific results are essential, and page retrieval only after you have selected an accessible source. Treat discovery, paper identity, citation metadata, and full text as separate states rather than expecting one endpoint to provide all four.
Separate the four jobs before choosing an API
“Google Scholar API” often stands in for several different requests. One developer wants titles and links for a topic. Another needs a canonical DOI. A third wants citation relationships. A fourth wants readable paper text. Those outputs come from different systems and should have different success tests.
| Job | Minimum useful record | What success means | What it does not prove |
|---|---|---|---|
| Discover candidate papers | Query, title, source URL, authors or venue, year | The agent has a shortlist to inspect | The paper identity is correct or the source is readable |
| Resolve paper identity | Stable ID such as DOI plus title, authors, year | Versions and duplicate titles can be reconciled | The metadata source has the full paper |
| Inspect citation relationships | Citing/cited work IDs and source attribution | The agent can follow a documented graph | Citation counts are identical across databases |
| Read an accessible source | Requested URL, status, title, extracted text, access time | The expected section or evidence is present | The source is the version of record or may be redistributed |
This separation prevents a common research-agent failure: a search result looks authoritative, so the agent treats its snippet, citation count, and linked version as one verified fact. They are different claims.
What Google Scholar officially provides
The current Google Scholar Search Help documents a human-facing search interface, email alerts, citation export, related articles, cited-by navigation, and links to available versions. It does not document a public search API. In its automation guidance, Google says bulk Scholar access is unavailable and tells automated software to respect robots.txt.
That distinction matters. A browser page with query parameters is not automatically a supported API contract. Its markup, ranking, access limits, and result presentation can change without the versioning and schema guarantees developers normally expect from an API. Citation export for an individual result is also not equivalent to a bulk search endpoint.
The practical answer is therefore not “find the hidden endpoint.” It is to choose a documented source whose contract matches the output your agent needs.
Compare the available routes by contract
| Route | Best fit | Documented output | Main boundary |
|---|---|---|---|
| Google Scholar web interface | Manual discovery, related papers, cited-by exploration, citation export | Human-readable result pages and source links | No documented public search API or bulk-access contract |
| OpenAlex API | Programmatic discovery across works, authors, sources, institutions, and topics | REST entities that can be searched, filtered, sorted, selected, and fetched by ID | OpenAlex corpus and ranking are not Google Scholar's |
| Semantic Scholar API | Paper, author, citation, venue, recommendation, and dataset workflows | Academic Graph, Recommendations, and Datasets APIs | Separate corpus, license, identifiers, and rate-limit contract |
| Crossref REST API | DOI lookup and publisher/member-deposited bibliographic metadata | Works, journals, members, funders, licenses, and related metadata | Excellent for identity metadata; not a universal full-text or relevance-ranking service |
| Third-party Google Scholar SERP API | Scholar-specific result behavior delivered through a vendor contract | Structured fields derived from Scholar result pages | Not an official Google API; adds vendor, policy, cost, and availability dependencies |
| AnyCrawler Scholar Search | Scholar-oriented source discovery inside the same search/crawl workflow | Normalized result candidates that can be routed to page access | Not affiliated with Google, not a promise of Google-exclusive data, and not full paper text |
Do not choose by the size of a provider's headline corpus alone. Check whether the fields you need are present, whether the identifiers are stable, how the source handles versions, what its license permits, and whether its ranking model fits the task.
Build the agent around evidence states
A research agent becomes easier to audit when every paper moves through explicit states. One compact record is enough to expose what is known and what still needs verification:
{
"query": "large language models",
"candidate_url": "https://example.org/paper-record",
"discovery_source": "scholar-search",
"paper_id": {
"doi": null,
"source_id": "candidate-id"
},
"identity_status": "unverified",
"text_status": "not_requested",
"citation_status": "not_verified",
"accessed_at": "2026-08-28T04:36:40Z"
}
The state transitions should be conditional:
- Candidate found: preserve the query, result position when available, discovery source, title, and URL.
- Identity checked: compare title, authors, year, venue, and a stable identifier across a metadata source or the publisher record.
- Accessible version selected: distinguish a publisher page, repository copy, preprint, and author-hosted version rather than silently merging them.
- Text verified: require the section or claim the downstream task needs; HTTP 200 alone is not enough.
- Citation ready: store the exact source URL, version decision, access time, and the passage or field supporting the claim.
An agent can stop early when the task needs only discovery. It should continue when the output will be cited, summarized as evidence, or compared with another paper.
Test the read step on a public paper page
Discovery should end with a source-level check. On 2026-08-28, the following request was run against AnyCrawler's public no-key crawl endpoint using the arXiv abstract page for Attention Is All You Need:
TARGET_URL='https://arxiv.org/abs/1706.03762'
curl --silent --get 'https://api.anycrawler.com/free/v1/crawl' \
--data-urlencode "url=${TARGET_URL}"
The response reported status_code: 200, preserved the requested URL, returned the title [1706.03762] Attention Is All You Need, and produced 9,887 Markdown characters. The Markdown contained the expected author name and a References & Citations section. The observed total_ms was 911 for this single request; it is a reproducible run record, not a performance benchmark.
The same response did not include final_url, canonical_url, credits_used, or markdown_tokens, so those fields are not part of this test's success claim. For authenticated production workflows, use the documented Fetch API when the source is present in HTML and Render API only when the required content depends on JavaScript.
Reject a candidate before it becomes a citation
Paper discovery fails in ways that a successful HTTP response cannot reveal. Use a diagnostic checklist before an agent cites a result:
- Ambiguous title: two works share a similar title, or a title changed between preprint and publication.
- Version collision: the result points to a preprint while the metadata describes a later version of record.
- Author mismatch: abbreviated or reordered names cause the wrong identity to be attached.
- Identifier gap: no DOI exists, or a DOI resolves to a correction, chapter, dataset, or different work type.
- Citation disagreement: two graphs report different counts because their sources, update schedules, and deduplication rules differ.
- Snippet-only evidence: the search result appears relevant, but the source page does not contain the claim.
- Access boundary: the landing page is visible, but the full text is unavailable or subject to separate access and reuse terms.
Record these outcomes as structured rejection reasons. Silently dropping failures makes a research report look cleaner while removing the evidence needed to reproduce it.
Where AnyCrawler fits in the workflow
AnyCrawler's role is routing, not replacing scholarly identity infrastructure. Start with Scholar Search when the source is unknown. Preserve the candidate result and then use Fetch or Render on selected, accessible pages. Resolve identifiers and publication metadata with a source designed for that job, and keep the discovery record separate from the verified paper record.
This is the same boundary used across the broader search, fetch, render, and screenshot decision framework: search answers where to look; page access answers what a selected source currently exposes. Neither step by itself proves that a citation is complete, current, or the version of record.
The unresolved choice: exact Scholar ranking or a stable scholarly graph?
Some applications genuinely need the ordering and presentation a researcher sees in Google Scholar. In that case, a third-party SERP provider may be closer to the requirement, but the application inherits a vendor contract and a web-interface dependency. Other applications need stable identifiers, filters, bulk-friendly entities, or reproducible graph traversal; a documented open scholarly API is usually a better fit even though its corpus and ranking will differ.
Write this choice into the product requirement. “Find relevant papers” is too vague. Specify whether the system must reproduce Scholar-like results, discover broadly across a documented corpus, verify DOI metadata, follow citations, or read an accessible source. The correct API follows from that output contract.
Frequently asked questions
Is there a free official Google Scholar API?
No public Google Scholar search API is documented by Google as of 2026-08-28. Google Scholar's official help describes the web interface, alerts, citation export, and source links, while stating that bulk access is unavailable and automated software should respect robots.txt. A third-party product may use “Google Scholar API” in its name, but that is the vendor's service contract, not an official Google API. Recheck Google's documentation before making a long-term architecture decision.
Can I scrape Google Scholar directly for a research agent?
Treat direct automation as a policy and reliability risk, not as a hidden API strategy. Google Scholar's help asks automated software to respect robots.txt and says bulk access is not available. Even when a result page loads, its markup and access behavior are not a versioned schema. Use a documented scholarly API when possible. If exact Scholar-result behavior is essential, evaluate a third-party service's terms, provenance, retention, rate limits, and failure handling before adopting it.
Which API is best for discovering papers programmatically?
There is no universal winner because each service has its own corpus, ranking, fields, identifiers, license, and update behavior. OpenAlex supports search and filtering across connected scholarly entities. Semantic Scholar exposes paper, author, citation, recommendation, and dataset services. AnyCrawler Scholar Search can supply normalized candidates inside a larger web-access workflow. Test the same representative queries, inspect missing fields and duplicates, and select the contract that matches your required output rather than headline coverage alone.
When should I use Crossref instead of a paper search API?
Use Crossref when the job is DOI resolution or verification of bibliographic metadata deposited by publishers and other Crossref members. It is useful after discovery because it can help normalize title, publication, contributor, and identifier records. It is not a replacement for every relevance-ranked discovery system, citation graph, or full-text source. A robust agent can discover candidates elsewhere, query Crossref when a DOI or bibliographic match is available, and retain both the discovery record and verified metadata.
Does AnyCrawler Scholar Search return full paper text?
No. Scholar Search is a discovery step that returns source candidates for a research workflow. A result link may point to a publisher page, repository, preprint, abstract, or another accessible record. Use Fetch when the required content is present in HTML and Render when JavaScript is necessary. Then validate the expected title, authors, section, or passage before citing it. AnyCrawler does not provide paywall bypass, guarantee full-text availability, or make a discovered version authoritative.
How should an agent verify a paper before citing it?
Preserve the original query and discovery result, then resolve the paper's identity using title, authors, year, venue, and a stable identifier when available. Distinguish preprints, accepted manuscripts, and publisher versions. Read the selected accessible source and verify the exact passage or metadata field supporting the claim. Finally, store the source URL, version decision, access time, and rejection reasons for alternatives. Citation counts or search snippets alone should never substitute for this source-level evidence record.






