To find and verify research papers with scholar search, use search results to build a shortlist, then check each paper against its repository or publisher record before citing it. Match the title, authors, identifiers and publication history; choose the version you intend to use; and read the passage that supports your claim. Keep these checks separate. A working DOI can identify an accessible record without proving that you have the right version, the complete paper or evidence for a particular conclusion. The workflow below turns those distinctions into a checklist and a small, testable collection step.
Start with a question that can reject a paper
Write down the research question before submitting a query. “BERT” produces a topic shortlist. “How does BERT pretraining use left and right context?” gives the reviewer something to look for in the source. Add a method, task or author when the first results are ambiguous, and retain the query beside each candidate.
Use the Scholar Search API's candidate discovery workflow when your application starts with a technical question and needs scholarly source links. Treat the returned title and snippet as triage material. They help decide which records to open; they do not establish the paper's publication status or verify its findings.
At this point, a useful candidate record contains the query, discovery source, returned title, candidate URL and discovery time. Record whatever author or venue context is actually present, without filling gaps from the model's memory. Keep a rejected candidate's reason too: wrong task, wrong paper, unavailable evidence or unresolved identity are different outcomes.
Match the paper before you summarize its claims
Open a primary record: the publisher's article page, a conference proceedings entry or an established repository record. Check the bibliographic fields together. A similar title can belong to another paper; a changed title can also belong to a later version of the same work. Neither case is safely resolved by a title match alone.
| Check | What to record | When to hold the candidate for review |
|---|---|---|
| Title and authors | Full title, author list and the record that supplied them | The title matches but authors or subject differ |
| Identifier | DOI, repository ID or proceedings ID, with its source | A DOI leads to a different work or has only been guessed from a citation |
| Dates | Submission, revision and publication dates as separate fields | A year has no known role or conflicts with the cited version |
| Version | Preprint revision, accepted manuscript or published record | The version read cannot be distinguished from the version cited |
| Publication status | Venue and any visible correction, withdrawal or retraction notice | A notice changes the interpretation of the intended claim |
| Evidence access | Metadata only, abstract, partial text or full text actually retrieved | The needed passage is absent, regardless of request success |
Resolve a DOI and inspect the destination's title and authors. A successful redirect is a transport observation; it does not perform that comparison. Where records disagree, preserve both values and their sources. Do not silently overwrite an earlier year with a later one, invent a missing DOI or merge two records because their abstracts sound alike.
Version labels carry information that a single published: true flag loses. A preprint, an accepted manuscript and a publisher's version of record can represent different stages of the same research. Crossref's versioning and update guidance distinguishes those stages and explains how significant corrections and retractions should be recorded. Use those distinctions to structure your evidence record, while checking the actual source for its current status. Missing relationship metadata leaves a question open.
A real mapping: BERT's preprint and conference record
The BERT arXiv record shows an initial submission on October 11, 2018 and revision v2 on May 24, 2019. Its repository identifier is 1810.04805; the page also displays the arXiv-issued DOI 10.48550/arXiv.1810.04805.
The ACL Anthology publication record lists BERT in the NAACL proceedings in June 2019, with DOI 10.18653/v1/N19-1423. Both records name Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova. The matching title and authors support associating these records, while their identifiers and dates retain different meanings.
That produces a useful mapping: one research work, an arXiv version history, and a separately identified conference publication. Store the association with the records used to establish it. Do not turn that association into a claim that the two files are identical; their complete texts were not compared for this example.
For a citation to the conference publication, use its conference year and DOI. If a claim depends on a particular preprint revision, retain that revision's identifier and the passage read. Combining the preprint's year with the conference DOI would conceal which record supplied the citation.
Collect a readable record without pretending to read the paper
The next step is small enough to test independently: retrieve a public landing page and check that expected bibliographic strings survived extraction. The following Python example uses AnyCrawler's public free crawl endpoint and the ACL BERT page. It requires Python and curl available as curl.exe on Windows; on another system, change that executable name to curl.
import hashlib
import json
import subprocess
from datetime import datetime, timezone
from urllib.parse import urlencode
source_url = "https://aclanthology.org/N19-1423/"
endpoint = "https://api.anycrawler.com/free/v1/crawl?" + urlencode({"url": source_url})
observed_at = datetime.now(timezone.utc).isoformat()
response = subprocess.run(
["curl.exe", "-sS", "-L", "--max-time", "60", "-w", "\n%{http_code}", endpoint],
capture_output=True, check=True, timeout=65,
)
body, status = response.stdout.decode("utf-8").rsplit("\n", 1)
data = json.loads(body)
result = data.get("results") or {}
markdown = result.get("markdown") or ""
expected = ["BERT", "Jacob Devlin", "10.18653/v1/N19-1423"]
record = {
"requested_url": source_url,
"final_url": data.get("final_url"),
"observed_at": observed_at,
"retrieval_timestamp": data.get("timestamp"),
"http_status": int(status),
"source_status": data.get("status_code"),
"title": result.get("title"),
"content_sha256": hashlib.sha256(markdown.encode("utf-8")).hexdigest(),
"credits_used": data.get("credits_used"),
"expected_strings": {s: s in markdown for s in expected},
"identity_status": "needs_review",
"full_text_status": "not_verified",
}
print(json.dumps(record, indent=2, ensure_ascii=False))
The publication check returned HTTP 200 and source status 200, with all the expected strings present. The response did not report final_url or credits_used; the example leaves both as null. The collection request retrieved the landing page's metadata and abstract context. It did not retrieve and review the complete paper or execute an authenticated Scholar Search request.
These string checks are a smoke test for extraction, not an identity-matching algorithm. The DOI could appear in a reference list on an unrelated page. Inspect the title, authors and surrounding context before changing identity_status. Also check the returned statuses and ok field in a production collector; handle transport errors, non-JSON responses and missing fields as explicit failures rather than an empty successful record. The local timeout values in the example are implementation choices, not service guarantees.
Keep observed_at separate from the response's retrieval timestamp because a free response may be cached. A Markdown hash helps identify the retrieved text, but cannot establish authorship, publication status or scientific validity. Store the raw response alongside the derived record when your application's retention rules allow it.
To test your own candidate, open the free public-page crawler and submit an accessible paper landing page. Inspect the extracted title, DOI and abstract before adapting the script. For a production workflow, choose Fetch when the required content is in the returned HTML and Render when JavaScript must execute to expose it. Neither route establishes access to restricted text. A screenshot records visible state; a login or interactive browser task needs a separately authorized workflow.
Promote a citation only after reading the relevant passage
Once identity is matched, select an accessible representation of the intended version. Follow a publisher or repository full-text link when it is available, using a reader or extraction path appropriate to that format. A landing-page crawl does not establish that a linked PDF was retrieved or parsed.
Save the claim you plan to make, the supporting passage, its section or page location, the version read and its source URL. Check the passage in context: the abstract's broad statement may not describe the population, dataset or conditions needed for your question. If the text is unavailable, the correct result can be “metadata verified; claim not checked.” That is a useful research outcome.
Route failures according to the missing evidence:
- Conflicting identity: retain the competing records and require review before merging them. A second source should help resolve the disagreement, not merely repeat one record.
- Abstract only: limit any description to what the abstract supports and keep full-text-dependent claims pending.
- Access restriction: seek an authorized repository or author-provided version, identify it explicitly, and stop if the required evidence remains inaccessible.
- Correction or withdrawal: preserve the notice and reassess the affected claim. An old copy remaining readable does not remove the notice's significance.
Citation count can help prioritize a reading queue, but it is not a check of a particular claim or method. A highly cited paper still needs the same identity, version and passage checks as a less familiar one.
When should an agent keep an earlier version?
A system that always replaces earlier records with the latest publication can lose the evidence behind an older decision. A research history may need the preprint available at the time, while a current literature review may prefer a later publication after checking its updates. These tasks need different selection rules.
Choose that rule before automating citation promotion. Should the application preserve the exact version originally read, recommend a newer version for review, or maintain both with a documented relationship? The answer depends on whether the user is reconstructing past knowledge or assessing the current record. Preserve the original evidence either way, and make a version change visible to the reviewer.
Frequently asked questions
Does a valid DOI prove that a paper is trustworthy?
A DOI helps identify a registered object and locate a record. It does not establish that the object is the paper you intended, that the methods are sound or that a particular sentence is supported. Check the destination's bibliographic details, choose the relevant version and inspect the evidence for your claim. Keep identity verification separate from assessment of the research itself.
Why do the preprint and published paper have different years?
The dates can describe different events. A manuscript may first appear in a repository and later appear in conference proceedings or a journal. Store submission, revision and publication dates with their labels instead of collapsing them into one year. When constructing a citation, use the date belonging to the version cited, and retain the original metadata if the relationship still needs review.
Can I cite a preprint when a published version exists?
Choose according to the question and the evidence actually read. An earlier revision may matter when documenting the history of an idea or reproducing an earlier claim. A current review may instead need the later publication and any updates. Identify the preprint version explicitly, record its relationship to the publication where verified, and avoid implying that the two texts have been checked for equivalence.
Does Scholar Search return the full research paper?
AnyCrawler Scholar Search supplies discovery candidates and source context. Open the selected source and determine what is accessible before claiming to have read the paper. A successful retrieval might contain only a landing page or abstract; a linked PDF is another resource. Record the access state and use an appropriate, authorized reading path for the material needed to support your specific claim.
Should a research agent rank papers by citation count?
Citation count can be one signal for deciding what to inspect, but it cannot replace task relevance or source review. A paper with many citations may still be the wrong version or fail to support the claim under consideration. Keep the discovery ranking separate from citation approval, and require the same bibliographic and passage checks for every paper promoted into an answer.
What should I save with each verified citation?
Save the research query, discovery source, matched title and authors, identifiers, selected version, bibliographic dates, source URL and retrieval time. Add the exact supporting passage with its section or page location, plus the claim it supports. Preserve access limitations and unresolved disagreements. When a field was not returned, leave it unknown rather than constructing a plausible value that later readers could mistake for evidence.






