Small Models Win: Why Fine-Tuned Small Models Outperform General Models started as a suspicion and became a thesis I could defend with citations. Across medicine, law, finance, code, math, and speech, the pattern repeats: a small model, fine-tuned on the right data, matches or beats a general model many times its size — at a fraction of the cost.

A few of the results that made it into the book:

  • A 13B model (Orca-2) scoring within a point of GPT-3.5 on reasoning benchmarks.
  • A 7B medical model (Meerkat) outscoring a 70B rival on MedQA: 74.3% vs 70.2%.
  • A 7B math specialist hitting 86.81% on GSM8K, ahead of GPT-3.5 and LLaMA-2-70B.
  • An API bill comparison with a 33× gap: ~$3,200/month vs ~$96/month for the same workload.

What's inside#

  • The evidence, domain by domain — medicine, law, finance, code, math, reasoning, language, retrieval, speech.
  • The mechanism — how fine-tuning actually works, and why data beats scale.
  • The economics — the full invoice math, break-even analysis, and when self-hosting wins.
  • The honesty — two full chapters on when fine-tuning fails and where scale still wins. The thesis is conditional, not a slogan.
  • The playbook — NVIDIA's published recipe for replacing giant models with small ones, plus a migration checklist.

Every figure is sourced from published papers and official reports, with per-chapter sources. Prices are September 2026 snapshots and labeled as such.

The book is free#

Small Models Win is free — the complete book, all 10 chapters plus the glossary, in both formats: PDF and ePub. The details page has the chapter breakdown, and the free sample PDF (preface plus chapters 1–2) is there if you want a taste first.