Every frontier lab now ships a "deep research" button, and every one of them shares the same dirty secret: the citations are the weakest link. Ask for sources and you get plausible-looking papers that don't exist, links that 404, and quotes that appear nowhere in the cited text. The synthesis got good. The evidence didn't.

This tutorial fixes that from the bottom up. You will build the evidence engine of a deep-research agent in about 360 lines of Python: a pipeline that decomposes a question into sub-queries, searches three keyless sources in parallel, ranks passages with BM25, cross-verifies every paper DOI against an independent index, and emits a report where every claim is backed by a real URL and a quoted passage. No API keys, no paid services, no embeddings to train. Every API call and parameter below was verified live while writing this — including two failure modes I hit and fixed, which you'll learn to avoid.

Run it once and you'll have a CLI that turns "How do diffusion models generate images?" into a cited markdown report in under a minute. Plug any LLM on top of it later and you have a research agent whose sources you can actually click.

What you'll need

  • Python 3.10+ and pip install requests. That's the only dependency — no torch, no API keys, no accounts.
  • Internet access to three public APIs: the Wikipedia Action API, OpenAlex, and Crossref. All are keyless; OpenAlex and Crossref ask only that you identify yourself with a mailto parameter, which joins their polite pool.
  • ~15 minutes and a terminal. The optional synthesis step at the end needs a local model (Ollama) or any LLM API you already use — the evidence engine itself needs neither.

One honest scoping note before we start: this tutorial builds the retrieval and verification half of deep research — the part commercial products hide. The final prose synthesis is a short, strictly-grounded prompt you run on your own model in Step 6. Nothing here pretends a rule-based decomposer is as smart as an LLM planner; what it is, is deterministic, free, and fully inspectable.

Step 1 — Decompose the question into sub-queries

A single search query is a single point of failure. Deep-research systems split the question into several angles — the original phrasing, a keyword core, a definitional angle (great for encyclopedias), and a research angle (great for papers). Our decomposer does this deterministically, which means it never invents a sub-query and behaves identically on every run:

STOPWORDS = set(
    "a an the and or of to in on for with is are was were be been what how why "
    "when where which who whom does do did can could should would it its this that".split()
)

def content_words(text):
    words = re.findall(r"[a-zA-Z][a-zA-Z0-9+\-#\.]{1,}", text.lower())
    return [w.strip(".-") for w in words if w.strip(".-") not in STOPWORDS]

def decompose(question):
    cw = content_words(question)
    core = " ".join(cw)
    subs = [question.strip()]
    if core and core != question.strip().lower():
        subs.append(core)
    if cw:
        subs.append("what is " + " ".join(cw[:6]))
    if cw:
        subs.append(" ".join(cw[:6]) + " research")
    seen, out = set(), []
    for s in subs:
        if s not in seen:
            seen.add(s)
            out.append(s)
    return out

For "How do diffusion models generate images?" this yields four sub-queries: the original, "diffusion models generate images", "what is diffusion models generate images", and "diffusion models generate images research". Crude? Yes. Effective? Measurably — each angle hits a different source's strengths, and the ranking step later discards the misses.

Step 2 — Search three keyless sources in parallel

Most "build a research agent" tutorials stop at one search API, usually one that needs a key. We use three that don't, each with a distinct job:

  • Wikipedia Action API — background knowledge and definitions. Two calls: list=search to find pages, then prop=extracts|info with explaintext=1 for plain-text content and inprop=url for the canonical URL.
  • OpenAlex — scholarly papers: titles, abstracts, publication years, citation counts, and landing-page URLs. The abstract arrives as an inverted index ({word: [positions]}), so we rebuild the text by sorting on position.
  • Crossref — not a retrieval source but an independent verifier. Every paper DOI we keep gets checked against Crossref's metadata index; a DOI that doesn't resolve there doesn't make the report.

Politeness matters with free APIs: set a descriptive User-Agent (Wikimedia requires one) and pass mailto to OpenAlex and Crossref. Wrap every request in retries with backoff — transient failures are normal:

UA = "DeepResearchTutorial/1.0 (educational example; contact: [email protected])"
TIMEOUT = 25

def http_get(url, params=None, headers=None, retries=3):
    h = {"User-Agent": UA}
    if headers:
        h.update(headers)
    last = None
    for attempt in range(retries):
        try:
            r = requests.get(url, params=params, headers=h, timeout=TIMEOUT)
            if r.status_code == 200:
                return r.json()
            last = f"HTTP {r.status_code}"
        except Exception as e:
            last = repr(e)
        time.sleep(1.5 * (attempt + 1))
    raise RuntimeError(f"GET {url} failed after {retries} tries: {last}")

The search functions themselves are thin wrappers around verified endpoints:

def search_wikipedia(query, limit=5):
    data = http_get(
        "https://en.wikipedia.org/w/api.php",
        params={"action": "query", "list": "search", "srsearch": query,
                "srlimit": limit, "format": "json"},
    )
    hits = data["query"]["search"]
    if not hits:
        return []
    titles = [h["title"] for h in hits]
    data = http_get(
        "https://en.wikipedia.org/w/api.php",
        params={"action": "query", "prop": "extracts|info", "inprop": "url",
                "explaintext": "1", "exsectionformat": "plain",
                "exlimit": "max", "titles": "|".join(titles), "format": "json"},
    )
    out = []
    for page in data["query"]["pages"].values():
        if "missing" in page or not page.get("extract"):
            continue
        out.append({"source": "wikipedia", "title": page["title"],
                    "url": page["fullurl"], "text": page["extract"],
                    "year": None, "cited_by": None})
    return out

def de_invert(inverted):
    """OpenAlex stores abstracts as {word: [positions]}; rebuild the text."""
    if not inverted:
        return ""
    pairs = [(pos, w) for w, poss in inverted.items() for pos in poss]
    pairs.sort()
    return " ".join(w for _, w in pairs)

def search_openalex(query, limit=8, since_year=None):
    # OpenAlex's search parser rejects punctuation like "?" — strip it.
    clean = re.sub(r"[^\w\s\-]", " ", query)
    params = {"search": clean, "per-page": limit,
              "select": "id,doi,title,publication_year,abstract_inverted_index,"
                        "cited_by_count,authorships,primary_location",
              "mailto": "[email protected]"}
    if since_year:
        params["filter"] = f"from_publication_date:{since_year}-01-01"
    data = http_get("https://api.openalex.org/works", params=params)
    out = []
    for w in data.get("results", []):
        loc = w.get("primary_location") or {}
        url = loc.get("landing_page_url") or w.get("doi") or w.get("id")
        authors = [a.get("author", {}).get("display_name", "")
                   for a in (w.get("authorships") or [])[:3]]
        out.append({"source": "openalex",
                    "title": w.get("title") or "Untitled",
                    "url": url,
                    "text": (w.get("title") or "") + "\n" + de_invert(w.get("abstract_inverted_index")),
                    "year": w.get("publication_year"),
                    "cited_by": w.get("cited_by_count"),
                    "authors": ", ".join(a for a in authors if a)})
    return out

Two real failure modes surfaced while verifying this, and both are handled above. First, OpenAlex returns HTTP 400 on queries containing ? — its search parser chokes on punctuation — so the full original question must be sanitized before it goes to OpenAlex (Wikipedia handles it fine). Second, arXiv's API rate-limited this machine with HTTP 429 from a shared IP; OpenAlex indexes the same papers through a far friendlier API, so it's the primary paper source here rather than a fallback.

All searches run concurrently in a thread pool — one future per (sub-query × source):

with cf.ThreadPoolExecutor(max_workers=6) as ex:
    futs = []
    for sq in subqueries:
        futs.append(ex.submit(search_wikipedia, sq, n_wiki))
        futs.append(ex.submit(search_openalex, sq, n_papers, since_year))
    for f in cf.as_completed(futs):
        try:
            docs.extend(f.result())
        except Exception as e:
            print(f"  [warn] a search failed: {e}", file=sys.stderr)
Diagram: one research question fans out to Wikipedia, OpenAlex, and Crossref, whose results funnel into a passage collection
Diagram generated for AI Frontier Post: the question fans out to Wikipedia, OpenAlex, and Crossref, and every result funnels into a single passage pool.

Step 3 — Chunk documents into passages

Ranking whole pages is too coarse: a 50,000-character Wikipedia article contains one paragraph you need and forty-nine you don't. Split on paragraph boundaries and greedily pack up to ~900 characters per passage, dropping stubs under 120 characters:

def chunk(text, max_chars=900):
    paras = [p.strip() for p in re.split(r"\n{2,}|\n", text) if p.strip()]
    chunks, cur = [], ""
    for p in paras:
        if len(cur) + len(p) + 1 > max_chars and cur:
            chunks.append(cur)
            cur = p
        else:
            cur = (cur + " " + p).strip() if cur else p
    if cur:
        chunks.append(cur)
    return chunks

A typical run produces 100–150 passages. That volume is exactly why the next step matters.

Step 4 — Rank passages with BM25 (no embeddings needed)

Here is the contrarian core of this tutorial: you don't need embeddings to rank research passages. Okapi BM25 — the algorithm behind Lucene, Elasticsearch, and decades of information retrieval — is deterministic, runs in milliseconds on CPU, needs no model download, and is brutally effective on keyword-rich research questions. It scores each passage by term frequency (dampened by k1), inverse document frequency across your passage set, and length normalization (b):

def tokenize(text):
    return [w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOPWORDS]

def bm25_scores(corpus_tokens, query_tokens, k1=1.5, b=0.75):
    N = len(corpus_tokens)
    df = Counter()
    for toks in corpus_tokens:
        for t in set(toks):
            df[t] += 1
    idf = {t: math.log((N - df[t] + 0.5) / (df[t] + 0.5) + 1.0) for t in df}
    avgdl = sum(len(t) for t in corpus_tokens) / max(N, 1)
    scores = []
    for toks in corpus_tokens:
        tf = Counter(toks)
        dl = len(toks)
        s = 0.0
        for t in query_tokens:
            if t not in tf:
                continue
            num = tf[t] * (k1 + 1)
            den = tf[t] + k1 * (1 - b + b * dl / avgdl)
            s += idf.get(t, 0.0) * num / den
        scores.append(s)
    return scores

Sanity-check it on three toy documents before trusting it on 140 real ones:

python3 - <<'EOF'
from research import tokenize, bm25_scores
corpus = [
    "diffusion models generate images by iterative denoising",
    "transformers use self attention for language modeling",
    "stable diffusion is a latent diffusion model for text to image",
]
scores = bm25_scores([tokenize(d) for d in corpus],
                     tokenize("how do diffusion models generate images"))
print([round(s, 2) for s in scores])
EOF
# [3.34, 0.0, 0.66] — the two diffusion passages win, the transformer text scores zero

Passages scoring zero are dropped outright. The rest are sorted best-first — that ordering is what the report shows.

Step 5 — Cross-verify every DOI, then build the cited report

This is the step commercial deep-research tools skip, and it's the whole point. For each paper in the top results, resolve its DOI against Crossref's independent metadata index. A DOI that 404s there never reaches your report — which is exactly how hallucinated citations get caught before a reader does:

def verify_doi(doi_url):
    if not doi_url or "doi.org" not in doi_url:
        return False
    doi = doi_url.split("doi.org/")[-1]
    try:
        r = requests.get(
            f"https://api.crossref.org/works/{urllib.parse.quote(doi, safe='')}",
            params={"mailto": "[email protected]"},
            headers={"User-Agent": UA}, timeout=TIMEOUT)
        return r.status_code == 200
    except Exception:
        return False

The report renderer then quotes each top passage verbatim, capped at 500 characters, with its title, source, year, citation count, and verification badge — plus a full source list of clickable URLs:

def build_report(question, ranked, top_k=12):
    top = ranked[:top_k]
    lines = [f"# Research: {question}", "",
             "## Key findings (each backed by a cited source)", "",
             "> Evidence-only draft: every bullet below quotes a real, "
             "retrieved passage. Run the synthesis step to turn this "
             "into prose with your own LLM.", ""]
    for i, (score, p) in enumerate(top, 1):
        quote = p["passage"][:500].replace("\n", " ").strip()
        meta = p["source"]
        if p.get("year"):
            meta += f", {p['year']}"
        if p.get("cited_by"):
            meta += f", cited by {p['cited_by']}"
        badge = " ✓ cross-verified" if p.get("verified") else ""
        lines.append(f"**[{i}] {p['title']}** ({meta}{badge})")
        lines.append(f"> {quote}…")
        lines.append("")
    lines.append("## Sources")
    lines.append("")
    for i, (score, p) in enumerate(top, 1):
        lines.append(f"[{i}] [{p['title']}]({p['url']}) — {p['source']}")
    return "\n".join(lines)

Wire it together with a CLI and run it for real:

python research.py "How do diffusion models generate images?" \
    --papers 5 --top 6 --since 2023 --out report.md
Q: How do diffusion models generate images?
Sub-queries (4):
  - How do diffusion models generate images?
  - diffusion models generate images
  - what is diffusion models generate images
  - diffusion models generate images research
Searching Wikipedia + OpenAlex + Crossref …
Ranked 140 passages.
Cross-verifying paper DOIs in Crossref …
Wrote report.md

The report opens with the Wikipedia Diffusion model article explaining denoising backbones, followed by cross-verified papers like SynthBuster (2023, 93 citations) and InstructPix2Pix (2023, 1,395 citations) — each with a DOI link that was independently confirmed to resolve. 140 passages went in; 6 cited findings came out.

Diagram: passages flow through BM25 ranking, DOI verification, and emerge as a final report with numbered citations
Diagram generated for AI Frontier Post: the citation verification stage — passages are ranked by BM25, paper DOIs are checked against Crossref, and the final report carries numbered, clickable citations.

Step 6 — Synthesize with your own LLM (the grounded way)

The evidence engine is LLM-free by design, but a report of quoted passages isn't a research brief yet. The rule for this step is absolute: the model only ever sees your verified evidence, never the bare question. This prompt template enforces it:

You are writing a research brief. You may ONLY use the evidence below.
Every factual claim must cite its source as [N]. If the evidence does not
support an answer, say so — do not fill gaps from training data.

QUESTION: {question}

EVIDENCE:
{numbered passages: [N] title, url, quoted passage}

Write:
1) a 3-sentence direct answer,
2) key findings as bullets, each with [N] citations,
3) open questions the evidence does not resolve.

Run it anywhere: paste it into any chat UI with the report attached, or automate it against a local model. Ollama exposes an OpenAI-compatible endpoint at http://localhost:11434/v1/chat/completions (any non-empty API key is accepted), so the same code works against Ollama, vLLM, or LM Studio:

from openai import OpenAI  # pip install openai

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
resp = client.chat.completions.create(
    model="llama3.2",  # pull it first: ollama pull llama3.2
    messages=[{"role": "user", "content": prompt}],
    temperature=0.2,   # low temperature: this is synthesis, not brainstorming
)
print(resp.choices[0].message.content)

Keep temperature low and the evidence attached. The model's job is now writing, not knowing — which is precisely why its citations stop being fiction.

Which approach should you use?

This keyless stack is the right default, but not the only option. Choose deliberately:

  • Keyless stack (this tutorial: Wikipedia + OpenAlex + Crossref) — free forever, fully reproducible, excellent for scientific and technical questions. Weak on breaking news, paywalled content, and niche forums.
  • Paid search APIs (Tavily, Brave Search) — better for fresh web content and long-tail queries; you pay per search. Worth it when the question is "what happened this week," not "how does this work."
  • Managed deep-research products — zero code, but you surrender control of source selection and can't audit the retrieval step. Fine for casual use; unacceptable when citations are the deliverable.
  • BM25 vs. embeddings for ranking — BM25 (this tutorial) is free, deterministic, and strong on keyword-heavy research questions. Embeddings win when queries and passages paraphrase each other heavily. Production systems use both: see our guides to late-interaction retrieval and cross-encoder reranking for the upgrade path.
  • Wikipedia vs. papers — definitions, background, and established mechanisms: Wikipedia. Recent findings, numbers, and contested claims: papers, weighted by citation count and DOI-verified.

The takeaway

Deep research is two problems wearing one trench coat: finding evidence and writing about it. The industry obsessed over the writing half and left the finding half to vibes — which is why citations hallucinate. Build the evidence engine first: decompose deterministically, search sources you don't have to pay for, rank with an algorithm from the 1970s that still works, verify every DOI against an independent index, and only then let a model turn it into prose. Do that, and your research agent's sources survive the one test that matters: a reader clicking them.

The complete script is ~360 lines in a single file — every function above, plus the CLI. Extend it next with recency weighting, a second verification pass that fetches each URL and checks the quote actually appears on the page, or an embeddings-based reranker on top of the BM25 shortlist.

API references