AI Frontier Post
AI News

Kisoku 1.6B: one person trained a laptop-size model from scratch — and it ties Llama 3.2 1B on 18× less data

0ARCH released Kisoku 1.6B, a fully open 1.6B-parameter language model one person trained from scratch on a free Google TPU grant. It splits ten benchmarks 5–5 with Meta's Llama 3.2 1B — on roughly 18 times less training data — and runs offline on a laptop.

0ARCH released Kisoku 1.6B on October 10, 2026: a free, open language model that one person trained from scratch. It has 1.6 billion parameters, trained on about 524 billion tokens — and the headline claim is striking. Run head-to-head against Meta's Llama 3.2 1B on the same ten benchmarks, through the same harness on the same machine, each model wins five. Llama 3.2 1B reportedly trained on around 9 trillion tokens, so Kisoku gets its draw on roughly 18 times less data.

The release is about as open as a release gets. The weights, the training code, the technical report, and every benchmark answer, sample by sample, are all published under the Apache 2.0 license. Nothing is gated behind an application form or a hosted-only endpoint.

The scorecard

Kisoku takes the reasoning and math tests; Llama takes knowledge and commonsense. The split, per 0ARCH's published numbers:

  • GSM8K (grade-school math): Kisoku 15.0 vs 5.8
  • MMLU (57 subjects): Kisoku 34.2 vs 31.3
  • ARC-Easy / ARC-Challenge (science): Kisoku 64.4 vs 61.8; 40.1 vs 36.9
  • BBH (hard multi-step reasoning): Kisoku 28.6 vs 28.3
  • PIQA, WinoGrande, HellaSwag, HumanEval, TriviaQA: Llama 3.2 1B leads on all five — including a wide 40.7 vs 23.1 gap on TriviaQA recall.

On long context (RULER, 13 tasks averaged), Kisoku scores 76.1 / 58.8 / 56.7 / 43.2 at 4K / 32K / 64K / 128K tokens against Llama's 73.5 / 56.7 / 49.2 / 43.1. It is worth noting the honest framing on the release page: Qwen2.5 1.5B and SmolLM2 1.7B beat both models on most of these tests. This is a result about data efficiency, not about the best small model on earth.

Laptop screen with a terminal and a notebook of benchmark numbers
Illustration: AI Frontier Post. Kisoku is built to run entirely on your own machine — nothing you type leaves the laptop.

Built on a free grant

The most interesting line in the spec sheet may be the hardware. Kisoku trained on a TPU v4-32 from Google's TPU Research Cloud — a free research grant. The model trained to 64K tokens of context (about 128K with YaRN), runs in Ollama with a 1.0 to 3.2 GB download, and hits about 150 words a second on a MacBook Pro with an M5 Max chip. Once the weights are downloaded, everything runs offline; the model has no internet access and no telemetry back to a vendor.

Ready-to-run GGUF builds for llama.cpp and Ollama are published alongside the base weights: Q8_0 (1.7 GB, recommended), Q4_K_M (1.0 GB), and F16 (3.2 GB) versions of both the base model and a chat fine-tune. The base model is the finished release; the chat model is still labeled a preview, and the release page is candid about its rough edges: it can lose the thread in long conversations, states wrong answers as confidently as right ones, and occasionally refuses things it could just do.

Abstract comparison of five benchmark pairs
Illustration: AI Frontier Post. Five benchmark wins each — a 524-billion-token model splitting the series with a 9-trillion-token one.

Why this matters

The sub-billion-to-few-billion parameter tier is where AI gets genuinely personal: models that run locally, offline, with no subscription and no data leaving the device. Kisoku's numbers won't dethrone the current kings of that tier — the page itself points you to Qwen2.5 and SmolLM2 for that — but it answers a different question. How far can careful training on modest compute go? Apparently: within shouting distance of a model trained on 18 times more data, built by one person, on someone else's free TPUs.

And unlike most model releases this year, you can check the homework. Every eval answer is public, sample by sample — so the 5–5 split is auditable, not just claimed. That's the standard small-model releases should be held to.