A 744-billion-parameter model does not fit on your machine. That sentence is true and useless — because a Mixture-of-Experts model never uses all its parameters at once. Colibri (JustVugg/colibri, Apache-2.0, 38,100+ GitHub stars) exploits exactly that: GLM-5.2's 744B parameters activate only ~40B per token, and of those, only about 11 GB change from token to token. So Colibri keeps the dense part — attention, shared experts, embeddings, ~17B params — resident in RAM at int4 (~9.9 GB), and streams the 19,456 routed experts from disk (~370 GB) on demand. A JIT, but for weights: measured routing heat drives a per-layer LRU cache, a learned pinned hot-store, and a router lookahead that predicts the next layer's experts with 71.6% accuracy.

Nine model families run today — GLM-5.2/5.3 (744B), GLM-5.3-Flash (321B, vision), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash (284B), DeepSeek V4.1 Flash (552B, vision), Qwen3.8-Flash-Next (125B), Qwen3.6 (35B-A3B), and OLMoE (7B) — through the same coli chat / coli serve / coli web front end. This tutorial builds the engine from source (I did it in 6 seconds), verifies the launcher, picks a model that fits your disk, and runs it — including Brio mode, the feature that scores closed questions without generating a single token.

What you'll need#

  • gcc (or clang) with OpenMP — or skip compiling and grab a prebuilt release below
  • Python 3 — only for the coli launcher and the one-time model converter; the engine itself is pure C with zero dependencies
  • Disk and RAM for the model you choose: the accessible path is OLMoE (~7 GB disk, 8 GB RAM); the full experience is GLM-5.2 (~372 GB disk, 16 GB RAM minimum)
  • No GPU. None of the models require one — a GPU only ever makes it faster

Step 1 — Get the engine: build it yourself#

Colibri's engine is one C file per model family (c/colibri.c, c/olmoe.c, …) plus small headers. No BLAS, no Python at runtime, no GPU required. Clone and build the OLMoE engine — the smallest one, and the one I built myself:

git clone https://github.com/JustVugg/colibri && cd colibri
make -C c olmoe

On a 2-core VM with gcc 13.3 that finished in about six seconds with a single compile line:

gcc -O3 -march=native -fopenmp -pthread -Wall -Wextra olmoe.c -o olmoe -lm -fopenmp -pthread

The result is a 197 KB binary. Run it bare and it tells you it is not the program you run directly — the launcher is:

colibri: this is the OLMoE engine, and it was started without a model.
The engine is not the program you run directly -- the launcher is:

    ./coli chat  --model <model directory>    interactive chat
    ./coli serve --model <model directory>    OpenAI-compatible API
    ./coli web   --model <model directory>    API plus the dashboard
    ./coli doctor --model <model directory>   check a model is usable

Prefer not to compile? Download a prebuilt release (Linux, macOS, Windows — no compiler needed) and unpack it:

mkdir colibri && tar xzf colibri-v1.8.0-linux-x86_64.tar.gz -C colibri && cd colibri
python3 coli info                         # engine ready ✓

From a source checkout, ./setup.sh checks gcc/OpenMP, builds, and self-tests, and pip install -e . puts coli on your PATH.

Step 2 — Verify the launcher#

cd colibri/c && python3 coli info

Output on my build:

    (\      colibri v1.12.1
     )·>     tiny engine, immense model
    / \      9 model families · MoE experts streamed from disk · CPU
             info

  ──────────────────────────────────────────────────────────
no model directory given.
  pass --model <dir>, or set COLI_MODEL=<dir>

The launcher reads the model's config.json and picks the matching engine binary itself — the command line never changes between models.

A three-tier memory pyramid: VRAM at top, RAM in middle, NVMe SSD at bottom, data streams feeding a central AI chip
AI-generated illustration for AI Frontier Post: Colibri's one memory hierarchy — VRAM, RAM, and storage as placement tiers for the same weights.

Step 3 — Get a model (the honest part)#

This is the step where you choose how deep to go. Two paths:

  • The starter path — OLMoE (7B). Convert with the bundled tool into an ~7 GB int8 container, then run it with 8 GB of RAM:
    python3 c/tools/convert_olmoe_merged.py   # OLMoE → ~7 GB int8 container
  • The full path — GLM-5.2 (744B). A pre-converted int4 group-scaled (gs64) container is on Hugging Face — about 372 GB, so it wants a big fast disk:
    https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp
    Or convert from the FP8 source yourself, shard by shard, without ever needing the full 756 GB at once:
    ./coli convert --model /nvme/glm52_i4

Two gotchas the README documents honestly, so you don't learn them the hard way. Use the gs64 container, not the older per-row int4 mirrors — those measure ~9 points worse on quality and caused the original think-mode loops. And the MTP speculation head must be int8, not int4: an int4 head collapses to 0–4% draft acceptance. Check with ls -l <model>/out-mtp-*.

The full requirements table (all from the README): GLM-5.3 419 GB / 16 GB+ RAM; GLM-5.3-Flash 195 GB; Inkling 469 GB; Kimi K3 ~1.6 TB (no conversion needed — its MXFP4 experts stream straight from the original checkpoint); DeepSeek V4 Flash 167 GB; Qwen3.6 ~20 GB. GPU optional everywhere.

Step 4 — Run it#

COLI_MODEL=/nvme/glm52_i4 ./coli chat     # RAM budget, cache and MTP auto-detected
COLI_MODEL=/nvme/glm52_i4 ./coli plan     # inspect the planned VRAM/RAM/disk placement
COLI_MODEL=/nvme/glm52_i4 ./coli doctor   # read-only readiness check
COLI_MODEL=/nvme/glm52_i4 ./coli tune     # measure and save this machine's fastest safe profile
./coli web  --model /nvme/glm52_i4        # API + dashboard, opens a browser
./coli serve --model /nvme/glm52_i4       # API + dashboard, headless

coli plan is the one to run first: it shows exactly where each part of the model will live — dense in RAM, hot experts pinned, the rest streamed — for your machine. coli doctor preflights the tensors, shards, index, and any disk mirror before you spend an hour on a bad download.

The engine keeps a learning cache as you use it: .coli_usage records which experts your workload actually routes to and pins the hottest ones automatically — Colibri literally gets faster the more you run it — and .coli_kv persists compressed conversation state across restarts, so conversations reopen warm with zero re-prefill.

A person at a desk with a laptop and external drive, chatting with a giant AI brain hologram
AI-generated illustration for AI Frontier Post: a frontier model running where you already are — no hyperscaler required.

Step 5 — Try Brio mode: answers without generating#

Most real questions are choices, not paragraphs: which queue, which verdict, which of four values a field may take. Brio mode hands the engine the options and reads the probability of each one instead of generating — completion_tokens is 0, no answer can fall outside your list, and every answer carries an entropy, so "the model is not sure" becomes a number you can put a threshold on:

./coli chat --model /nvme/qwen36_i4_gs64
> /brio merge | request changes | close
> 340 lines, 8 files, no tests. CI is green but nothing covers that path.

Or one JSON request against the running server:

curl -s http://127.0.0.1:8000/v1/brio -H 'Content-Type: application/json' -d '{
  "model": "qwen36",
  "state": "340 lines, 8 files, no tests. CI is green but nothing covers that path.",
  "question": "What should the reviewer do?",
  "options": ["merge", "request changes", "close"]}'

Measured on Qwen3.6 against generating the same answer on the same CPU box: 2.4× faster on a four-field schema, 5.7× on four questions about one document read once. It runs on all nine families and is opt-in per request — chat stays byte-identical for everyone who doesn't ask.

What you built#

A self-hosted inference stack that treats VRAM, RAM, and NVMe as one placement hierarchy and streams routed experts from disk through a learned cache: a 197 KB engine binary, the coli launcher, a chat TUI, an OpenAI-compatible API, a live dashboard (chat dock, Brio page, Brain page with the measured expert atlas, profiling), and Brio mode for closed-question scoring. The same engine spans a 25 GB laptop (everything streams from disk — slow but correct) to a 6× RTX 5090 host (full expert residency, disk drops out of the decode path).

Honest limitations#

  • Speed is set by your disk. Measured decode: 5.8–6.8 tok/s on 6× RTX 5090 with full residency; ~1.8 tok/s on a 128 GB CPU-only desktop; ~1 tok/s on a single laptop GPU; 0.05–0.1 tok/s cold on a 25 GB box. The project states it plainly: no SLA on speed.
  • The hard guarantee is on semantics, not speed. Placement only ever decides speed — the router's decisions and weight precision are identical whether an expert answered from VRAM or disk, with the forward pass validated against a transformers oracle.
  • The downloads are enormous. 372 GB for GLM-5.2, 1.6 TB for Kimi K3. Start with OLMoE (~7 GB) to learn the workflow.
  • It's a research platform, not a product. Backends range from experimental (Metal, Vulkan) to battle-tested; speculation and caching are measurable policies, not promises — measure on your hardware with ./coli tune.