AI Frontier Post
A person at a desk chatting with two friendly robot coworkers on a large screen showing live work
Always-on AI coworkers, each with its own screen of work. Illustration generated for AI Frontier Post.
A private AI research agent drafting a cited report on a dark desktop
A research agent that lives on your machine and cites its sources. Illustration generated for AI Frontier Post.

Deep-research agents usually live on someone else’s computer: your question goes to a cloud, your data rides along. Local Deep Research — an MIT-licensed, open-source project with roughly 9,200 GitHub stars — flips that. The whole pipeline runs on your own machine: your models through Ollama, your search through a self-hosted SearXNG, your notes in an encrypted local database, and every answer returned with citations. The headline benchmark the project publishes is loud: roughly 95% on SimpleQA (n=500) with Qwen3.6-27B running locally on a single RTX 3090 — though the README itself flags the caveats (small samples, LLM-grader noise, contamination risk). Today we’ll build the working setup: the app running, search connected, a cited research report, and a private knowledge base it can search.

What you’ll need

Local Deep Research is a self-hosted app, not a hosted product: you run it, you point it at your own models and search, and the docs say exactly which pieces are still rough. Nothing below pretends otherwise.

Step 1 — pull the three pieces

On Linux, the README’s recipe is three containers: Ollama for the models, SearXNG for search, then the app itself. Model weights are multi-gigabyte, so start the pull first:

docker run -d -p 11434:11434 --name ollama ollama/ollama
docker exec ollama ollama pull gpt-oss:20b

docker run -d -p 8080:8080 --name searxng searxng/searxng

gpt-oss:20b is the README’s starter choice. If you have the VRAM, the project’s own benchmark table suggests Qwen3.6-27B or Qwen3.5-9B as the models that made the headline numbers — the same community-maintained leaderboard on Hugging Face (local-deep-research/ldr-benchmarks) tracks which Ollama and llama.cpp models actually work well for research before you download multi-GB weights.

Step 2 — start Local Deep Research

docker run -d --network host \
  --name local-deep-research \
  --volume "deep-research:/data" \
  -e LDR_DATA_DIR=/data \
  -e LDR_SEARCH_ENGINE_WEB_SEARXNG_DEFAULT_PARAMS_INSTANCE_URL=http://localhost:8080 \
  localdeepresearch/local-deep-research

Open http://localhost:5000 after about 30 seconds. Two details worth knowing from the docs: the long environment variable pins SearXNG’s address and marks it operator-approved — private/localhost engine URLs are blocked by default since v1.10.3 — and --network host only works on native Linux. On Docker Desktop (Mac/Windows/WSL2), host networking silently fails; the README offers an all-platform Docker Compose shortcut instead:

curl -O https://raw.githubusercontent.com/LearningCircuit/local-deep-research/main/docker-compose.yml && docker compose up -d

There’s also a pure pip route — pip install local-deep-research then python -m local_deep_research.web.app — but you still need Ollama and SearXNG running alongside it.

Step 3 — run your first cited research

The app offers four research modes: Quick Summary (answers in 30 seconds to 3 minutes), Detailed Research (comprehensive analysis with structured findings), Report Generation (professional reports with sections and a table of contents), and Document Analysis (search your private documents — that’s Step 4). Start with a Quick Summary on a question you can check: an obscure technical fact with a verifiable answer, not an opinion. The output comes back with citations on its claims, and you can follow each one to the source page it was scraped from.

Then run the same question as a Detailed Research and compare. The difference is the point: Quick Summary compresses the search results into an answer; Detailed Research runs the agentic loop — search, read, refine the query, search again — before it writes. If you’re evaluating the tool, this comparison tells you what the agentic loop actually buys you.

Three self-hosted services — model, search, and research app — connected on one machine
Model, search, and app: the whole research stack in three containers. Illustration generated for AI Frontier Post.

Step 4 — build your own knowledge base

This is where a local research tool earns its keep over a cloud tab. Feed it your documents — papers, reports, notes — and it builds a searchable knowledge base out of them, so Document Analysis answers from your files with the same citation habit. Your notes never leave the machine; they sit in a SQLCipher-encrypted local database.

A quietly useful addition in recent versions: a journal-quality system that scores source reputation using OpenAlex and DOAJ data (both CC0) plus a predatory-journal list, with a quality dashboard. When your research leans on papers, it’s the difference between “ten sources” and “ten sources worth trusting.”

Step 5 — drive it from code

The web UI is the demo; the API is the tool. The README ships a Python client:

from local_deep_research.api import LDRClient, quick_query

# Option 1: Simplest - one line research
summary = quick_query("username", "password", "What is quantum computing?")
print(summary)

# Option 2: Client for multiple operations
client = LDRClient()
client.login("username", "password")
result = client.quick_research("What are the latest advances in quantum computing?")
print(result["summary"])

There’s also an authenticated HTTP API (see docs/api-quickstart.md in the repo) and an MCP server that exposes the research tools to Claude Desktop and Claude Code — the natural next step if your coding agent should be able to do cited research instead of guessing. For the truly hands-off version, research subscriptions run daily or weekly digests on topics you subscribe to, filtered and summarized by the AI.

A locked private library of documents an AI can search
Your documents, your vault: the private knowledge base. Illustration generated for AI Frontier Post.

What you built

A complete deep-research stack that answers to nobody but you: local models served through Ollama, self-hosted web search through SearXNG, an app that plans searches, reads results, and writes cited reports in four modes — plus a private, encrypted, searchable knowledge base over your own documents, a Python and HTTP API for automation, and an MCP bridge into Claude. Research history is saved, exportable as PDF or Markdown, and every claim in it points at a source you can open.

Honest limitations

The bet Local Deep Research makes is simple: the research agents getting the most attention right now run in someone else’s cloud, on someone else’s models, reading someone else’s search index. Roughly 9,200 developers starred the version that runs on your machine and shows its work. That’s worth building.