The insight behind Glyd is a bit-level fact about your models: in a bf16 weight, the sign and mantissa are noise and the exponent carries under 3 of its 8 bits. So Glyd codes the exponent, holds weights and the KV cache compressed in GPU memory, and decodes them on the GPU where they are used — inside the matrix product and inside attention. The weights come back bit for bit. Nothing is quantized, pruned, or approximated, which is why the project's claims survive the one test that matters: MMLU answers match bf16's on 98.0–100% of questions, and perplexity stays within 0.06%.

What you'll need

  • For serving models: Linux (x86_64 or aarch64) with an NVIDIA GPU, Ampere or newer, and a recent driver — plus PyTorch for CUDA 12 or 13. The GPU kernels are the part under the Business Source License 1.1: free for personal, educational, research, and other non-commercial use; commercial production use needs a license.
  • For the compression CLI: any machine — macOS or Linux via Homebrew, Linux/macOS via pip, or release binaries for Linux x86_64, Linux aarch64, and macOS arm64.
  • A Hugging Face model to pack (Qwen3-8B in the examples; Qwen3, Qwen2, Llama, Mistral, and Granite families are the ones glyd pack understands).
  • Licenses that matter: the codec — crate, CLI, C ABI, Python and Go bindings — is BSD-3-Clause or GPL-2.0, the same licenses as zstd, so anything that may ship zstd may ship Glyd. The store (glyd-store) and the GPU weights code are BUSL-1.1.

1. Install Glyd

Pick the install that matches what you want to do:

pip install "glyd[gpu]"      # models on the GPU: Linux x86_64 / aarch64, a CUDA GPU (Ampere or later), PyTorch for CUDA 12 or 13
brew install surya-koritala/glyd/glyd   # macOS / Linux: the glyd and glyd-store CLIs, glyd.h
pip install glyd             # Python bindings: Linux x86_64 / aarch64, macOS arm64
cargo install --git https://github.com/surya-koritala/Glyd glyd glyd-store glyd-gpu  # from source, Rust 1.80+

2. Load a packed model in Python

This is the two-line payoff. Glyd loads the model packed on the GPU as it downloads — Qwen3-8B arrives as 11.2 GB of weights, not 16.4:

import glyd
model = glyd.from_pretrained("Qwen/Qwen3-8B")               # packed on the GPU as it loads
model = glyd.from_pretrained("Qwen/Qwen3-8B", exact=True)   # logits bit for bit bf16's

One behavior to know about: generate() runs compiled on PyTorch 2.13.0 or later (a static cache, CUDA graphs); pass compile=False — or set GLYD_COMPILE=0 — and it runs eager. If you are on an older PyTorch, expect the eager path, not an error.

3. Pack a model on the CPU with the CLI

Packing does not need Python or a GPU — the glyd pack subcommand writes the packed bytes to disk, which is the artifact you would check into a model registry or copy to a machine without the Python stack:

glyd pack Qwen/Qwen3-8B qwen3-8b-glyd                    # saved packed on the CPU, no Python or GPU
glyd pack Qwen/Qwen3-8B qwen3-8b-glyd12 --layout mma12   # the 12-bit layout: an A10, A100 or H100 loads it as saved

The mma12 layout is the one tuned for the tensor-core path: on an H100, Qwen3-32B sits in 49.2 GB instead of 65.5 GB at 26.4 ms of GPU time a token, with the packed matrix products 1.1–1.2× faster than cuBLAS's. The default (tiered) layout takes the more standard route and is what the savings table below is measured with.

Diagram of memory savings: a long bar of bf16 weight blocks shrinks into a bar one third shorter of packed blocks, with arrows showing the flow
AI-generated diagram for AI Frontier Post — bf16 weights compress to roughly two thirds of their size, bit for bit

4. Serve it through vLLM

The fastest way into the serving path is the project's installer, which bundles Glyd with vLLM 0.30 and PyTorch as one tool (Linux, NVIDIA GPU, driver 580 or newer — on a Mac, or with no GPU, it installs the compression CLI instead):

curl -LsSf https://getglyd.com/install.sh | sh   # Glyd, vLLM 0.30 and PyTorch as one tool
glyd run Qwen/Qwen3-8B                           # downloads the model, starts it packed, and opens a chat

glyd run checks the machine, works out the memory share, the context length, and the chat parsers from the GPU it finds (a 16 GB card included), and chats in the terminal and at http://localhost:8000. Two more commands round out the serving story: glyd serve MODEL leaves the model up as an OpenAI API, and glyd doctor reports what the machine has and which models fit.

What the packed serving buys you, measured on Lambda Cloud's 4× RTX A6000 (48 GB) cards on 2026-09-26 — nvidia-smi taken during the runs:

The same runsbf16Glyd
Qwen3-32B: weights65.5 GB44.5 GB
Qwen3-32B: 48 GB GPUs it takes21
Qwen3-32B: tokens/s at 1 / 8 / 32 sequences9.5 / 74 / 27611.8 / 95 / 290
Qwen2.5-72B: weights145.4 GB97.8 GB
Qwen2.5-72B: 48 GB GPUs it takes43
KV cache, Qwen2.5-7B, 16K tokens947 MB651 MB

At Lambda Cloud's on-demand prices ($1.09 an RTX A6000-hour), that is Qwen3-32B served around the clock for $9,548 a year instead of $19,097 — a 50% cut. And quality holds where it counts: Qwen3.8 27B moves from 15.1946 to 15.1941 perplexity and 79.7% to 80.0% MMLU; Llama 3.1 8B Instruct goes 19.5910 to 19.5909 perplexity, 71.7% to 72.0% MMLU.

5. The same engine, as a file compressor

Glyd's model work sits on top of a general-purpose compression codec that competes with zstd, and you can use it on your own data without touching a GPU. The level you want first is --max, the project's designated zstd -3 slot — fewer bytes, 3–7× faster reads on a server:

glyd --max  events.json -o events.glyd            # the zstd -3 slot: fewer bytes, 3-7x faster reads
glyd --max -r access.log -o access.glyd           # record mode: logs, dumps, CSV, JSON lines as columns
glyd --ultra -r dump.sql -o dump.glyd             # fewest bytes from a parse; slow to write
glyd --cold -r dump.sql -o dump.glyd              # fewest bytes of all; 1 MB/s per core each way
glyd --base dump-mon.sql dump-tue.sql -o tue.glyd # base mode: Tuesday's dump against Monday's
glyd -d events.glyd -o events.json                # the level and mode are in the stream
glyd -b bigfile                                   # benchmark every level on your data

Two modes deserve a closer look. Record mode (-r) detects the shape of logs, SQL dumps, CSV, and JSON lines, turns each field into a typed stream (integer and date-time deltas, dictionaries with recency ranks, text), and compresses those — which is how a bf16 checkpoint becomes 33% smaller, a fine-tune against its base 44% smaller. Base mode (--base) compresses a version against the last one, like a delta patch for nightly dumps; decoding needs the same base, so keep it around. Not sure which level fits? glyd -b benchmarks every level against your data and shows you the tradeoff directly.

Diagram of the Glyd pipeline: large glowing tensor cubes enter zigzag compression waves and exit as smaller packed cubes into a softly glowing server rack
AI-generated diagram for AI Frontier Post — weights enter the encoder, packed weights leave, the GPU never sees the difference

6. The honesty check: what it can't do

The project publishes a "known gaps" section that is unusually specific, and it is the part worth reading before you commit. The short version:

  • Reads are the tax. Record-mode reads cost 2–2.7× zstd's CPU rebuilding the columns, which makes plain zstd -3 the cheaper choice when reads are CPU-billed and frequent. On Sapphire Rapids, one-core reads run 0.79–0.97× zstd's — on Graviton3 and Ryzen they beat it.
  • High concurrency is not its best shape. At 64 sequences, Glyd generates 0.82–0.86× bf16's tokens/s on the A6000s (1.04× on a 4080 SUPER), though it still costs less per token. Prompt processing takes 1.2–1.6× bf16's time on the A6000s.
  • --ultra is a specialist. It runs 0.1–3.5% larger than the better of zstd -19 and -22 on plain text, executables, source, and OS trees — and up to 15% larger than xz -9e or brotli -11 there, though it reads at 500–1,600 MB/s against their 30–125. Its territory is records and containers, where it ties or wins.
  • --cold is symmetric. Reads cost what writes cost — 1–1.3 MB/s per core — so it is for data read a few times in its life: archives, compliance holds. Not a tier that serves reads.
  • The ceiling is real. No lossless code takes more than about 34% off bf16 weights or KV cache (measured: 10.5–10.6 bits a value). Glyd's 10.80 bits is 32.5% off — close to the wall. Models already shipped in FP8 have about 18% to take, in NVFP4 about 7%.
  • The license has two faces. The codec is BSD-3-Clause or GPL-2.0 — ship it anywhere you could ship zstd. The GPU weights and the store are BUSL-1.1: fine for personal, research, and non-commercial use, but commercial production needs a license from the author. Know which one you are building on.

What you built

A packed model directory that loads on fewer GPUs than it was trained for, an OpenAI-compatible endpoint serving it, and a compression toolkit you can point at any log, dump, or checkpoint. The realistic production shape is this: pack once with glyd pack, serve packed through glyd serve or the vLLM path, and spend the saved memory on KV cache — vLLM gives the freed VRAM to the cache, which is why saturated requests-per-second rise (1.33× on an L4, 1.65× on an A100 40 GB for Qwen3-14B).

Where to go next inside the project: gpu/vllm/README.md in the repo covers the manual vllm serve setup for when the one-command path does not fit your deployment, and ROADMAP.md shows where the project is heading. And since the codec is the same zstd license as zstd, the CLI half of this tutorial is production-safe for anything you can already do with zstd.

The takeaway

Model compression in 2026 mostly means quantization — trading accuracy for memory and hoping nobody notices. Glyd's bet is the opposite direction: the bits are redundant, not the model, so take the redundant bits and change nothing else. The measured result is 33% less GPU memory with perplexity and MMLU numbers that are, for all practical purposes, the same model. The caveats are priced in plainly: the GPU code is source-available, not fully open source; reads cost more CPU than zstd in the modes that save the most; and there is a hard ceiling on how much more any lossless code can ever take. But if your constraint is GPUs — and in 2026, whose isn't — two lines of Python that turn 16.4 GB into 11.2 GB, bit for bit, deserve an afternoon of your time.


Every number in this article was measured by the Glyd project, not sketched: Qwen3-32B and Qwen2.5-72B runs on Lambda Cloud 4× RTX A6000 (48 GB), 2026-09-26, with raw logs in the repo's benchmarks/gpu/ directory; quality numbers from the popular-models harness on an H100 SXM and an A10, 2026-09-26/27. Commands verified against the README at release v0.14.3. Star count (~240) measured at time of writing; the project debuted on September 19, 2026.