Run a 744B-parameter model on your own machine: a hands-on tutorial for Colibri
What if a 744-billion-parameter model ran on hardware you already own? Colibri's tiny engine keeps the dense weights in RAM and streams the routed experts from disk — frontier MoE inference in pure C, no GPU required, no API, no hyperscaler.

A 744-billion-parameter model does not fit on your machine. That sentence is true and useless — because a Mixture-of-Experts model never uses all its parameters at once. Colibri (JustVugg/colibri, Apache-2.0, 38,100+ GitHub stars) exploits exactly that: GLM-5.2's 744B parameters activate only ~40B per token, and of those, only about 11 GB change from token to token. So Colibri keeps the dense part — attention, shared experts, embeddings, ~17B params — resident in RAM at int4 (~9.9 GB), and streams the 19,456 routed experts from disk (~370 GB) on demand. A JIT, but for weights: measured routing heat drives a per-layer LRU cache, a learned pinned hot-store, and a router lookahead that predicts the next layer's experts with 71.6% accuracy.
Nine model families run today — GLM-5.2/5.3 (744B), GLM-5.3-Flash (321B, vision), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash (284B), DeepSeek V4.1 Flash (552B, vision), Qwen3.8-Flash-Next (125B), Qwen3.6 (35B-A3B), and OLMoE (7B) — through the same coli chat / coli serve / coli web front end. This tutorial builds the engine from source (I did it in 6 seconds), verifies the launcher, picks a model that fits your disk, and runs it — including Brio mode, the feature that scores closed questions without generating a single token.
What you'll need#
gcc(or clang) with OpenMP — or skip compiling and grab a prebuilt release below- Python 3 — only for the
colilauncher and the one-time model converter; the engine itself is pure C with zero dependencies - Disk and RAM for the model you choose: the accessible path is OLMoE (~7 GB disk, 8 GB RAM); the full experience is GLM-5.2 (~372 GB disk, 16 GB RAM minimum)
- No GPU. None of the models require one — a GPU only ever makes it faster
Step 1 — Get the engine: build it yourself#
Colibri's engine is one C file per model family (c/colibri.c, c/olmoe.c, …) plus small headers. No BLAS, no Python at runtime, no GPU required. Clone and build the OLMoE engine — the smallest one, and the one I built myself:
git clone https://github.com/JustVugg/colibri && cd colibri
make -C c olmoe
On a 2-core VM with gcc 13.3 that finished in about six seconds with a single compile line:
gcc -O3 -march=native -fopenmp -pthread -Wall -Wextra olmoe.c -o olmoe -lm -fopenmp -pthread
The result is a 197 KB binary. Run it bare and it tells you it is not the program you run directly — the launcher is:
colibri: this is the OLMoE engine, and it was started without a model.
The engine is not the program you run directly -- the launcher is:
./coli chat --model <model directory> interactive chat
./coli serve --model <model directory> OpenAI-compatible API
./coli web --model <model directory> API plus the dashboard
./coli doctor --model <model directory> check a model is usable
Prefer not to compile? Download a prebuilt release (Linux, macOS, Windows — no compiler needed) and unpack it:
mkdir colibri && tar xzf colibri-v1.8.0-linux-x86_64.tar.gz -C colibri && cd colibri
python3 coli info # engine ready ✓
From a source checkout, ./setup.sh checks gcc/OpenMP, builds, and self-tests, and pip install -e . puts coli on your PATH.
Step 2 — Verify the launcher#
cd colibri/c && python3 coli info
Output on my build:
(\ colibri v1.12.1
)·> tiny engine, immense model
/ \ 9 model families · MoE experts streamed from disk · CPU
info
──────────────────────────────────────────────────────────
no model directory given.
pass --model <dir>, or set COLI_MODEL=<dir>
The launcher reads the model's config.json and picks the matching engine binary itself — the command line never changes between models.

Step 3 — Get a model (the honest part)#
This is the step where you choose how deep to go. Two paths:
- The starter path — OLMoE (7B). Convert with the bundled tool into an ~7 GB int8 container, then run it with 8 GB of RAM:
python3 c/tools/convert_olmoe_merged.py # OLMoE → ~7 GB int8 container - The full path — GLM-5.2 (744B). A pre-converted int4 group-scaled (gs64) container is on Hugging Face — about 372 GB, so it wants a big fast disk:
Or convert from the FP8 source yourself, shard by shard, without ever needing the full 756 GB at once:https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp./coli convert --model /nvme/glm52_i4
Two gotchas the README documents honestly, so you don't learn them the hard way. Use the gs64 container, not the older per-row int4 mirrors — those measure ~9 points worse on quality and caused the original think-mode loops. And the MTP speculation head must be int8, not int4: an int4 head collapses to 0–4% draft acceptance. Check with ls -l <model>/out-mtp-*.
The full requirements table (all from the README): GLM-5.3 419 GB / 16 GB+ RAM; GLM-5.3-Flash 195 GB; Inkling 469 GB; Kimi K3 ~1.6 TB (no conversion needed — its MXFP4 experts stream straight from the original checkpoint); DeepSeek V4 Flash 167 GB; Qwen3.6 ~20 GB. GPU optional everywhere.
Step 4 — Run it#
COLI_MODEL=/nvme/glm52_i4 ./coli chat # RAM budget, cache and MTP auto-detected
COLI_MODEL=/nvme/glm52_i4 ./coli plan # inspect the planned VRAM/RAM/disk placement
COLI_MODEL=/nvme/glm52_i4 ./coli doctor # read-only readiness check
COLI_MODEL=/nvme/glm52_i4 ./coli tune # measure and save this machine's fastest safe profile
./coli web --model /nvme/glm52_i4 # API + dashboard, opens a browser
./coli serve --model /nvme/glm52_i4 # API + dashboard, headless
coli plan is the one to run first: it shows exactly where each part of the model will live — dense in RAM, hot experts pinned, the rest streamed — for your machine. coli doctor preflights the tensors, shards, index, and any disk mirror before you spend an hour on a bad download.
The engine keeps a learning cache as you use it: .coli_usage records which experts your workload actually routes to and pins the hottest ones automatically — Colibri literally gets faster the more you run it — and .coli_kv persists compressed conversation state across restarts, so conversations reopen warm with zero re-prefill.

Step 5 — Try Brio mode: answers without generating#
Most real questions are choices, not paragraphs: which queue, which verdict, which of four values a field may take. Brio mode hands the engine the options and reads the probability of each one instead of generating — completion_tokens is 0, no answer can fall outside your list, and every answer carries an entropy, so "the model is not sure" becomes a number you can put a threshold on:
./coli chat --model /nvme/qwen36_i4_gs64
> /brio merge | request changes | close
> 340 lines, 8 files, no tests. CI is green but nothing covers that path.
Or one JSON request against the running server:
curl -s http://127.0.0.1:8000/v1/brio -H 'Content-Type: application/json' -d '{
"model": "qwen36",
"state": "340 lines, 8 files, no tests. CI is green but nothing covers that path.",
"question": "What should the reviewer do?",
"options": ["merge", "request changes", "close"]}'
Measured on Qwen3.6 against generating the same answer on the same CPU box: 2.4× faster on a four-field schema, 5.7× on four questions about one document read once. It runs on all nine families and is opt-in per request — chat stays byte-identical for everyone who doesn't ask.
What you built#
A self-hosted inference stack that treats VRAM, RAM, and NVMe as one placement hierarchy and streams routed experts from disk through a learned cache: a 197 KB engine binary, the coli launcher, a chat TUI, an OpenAI-compatible API, a live dashboard (chat dock, Brio page, Brain page with the measured expert atlas, profiling), and Brio mode for closed-question scoring. The same engine spans a 25 GB laptop (everything streams from disk — slow but correct) to a 6× RTX 5090 host (full expert residency, disk drops out of the decode path).
Honest limitations#
- Speed is set by your disk. Measured decode: 5.8–6.8 tok/s on 6× RTX 5090 with full residency; ~1.8 tok/s on a 128 GB CPU-only desktop; ~1 tok/s on a single laptop GPU; 0.05–0.1 tok/s cold on a 25 GB box. The project states it plainly: no SLA on speed.
- The hard guarantee is on semantics, not speed. Placement only ever decides speed — the router's decisions and weight precision are identical whether an expert answered from VRAM or disk, with the forward pass validated against a transformers oracle.
- The downloads are enormous. 372 GB for GLM-5.2, 1.6 TB for Kimi K3. Start with OLMoE (~7 GB) to learn the workflow.
- It's a research platform, not a product. Backends range from experimental (Metal, Vulkan) to battle-tested; speculation and caching are measurable policies, not promises — measure on your hardware with
./coli tune.