Merge models, don't just train them: hands-on model merging with MergeKit
Some of the most capable open-weight models on the Hugging Face Hub were never trained — they were merged. Here is the full workflow, measured on a CPU: task arithmetic, SLERP, and how to pick a method.

Some of the most capable open-weight models on the Hugging Face Hub were never trained. They were merged — built by doing arithmetic directly on the weights of two or more existing models, with no gradient descent involved. It sounds like alchemy, and in 2023 it mostly was: forum recipes, vibes, and a lot of broken checkpoints. Then Arcee AI's Charles Goddard turned the folk practice into a toolkit, MergeKit (7,300+ stars on GitHub, with an EMNLP 2024 Industry Track paper to its name), and the community merged thousands of models with it — some of which topped the open LLM leaderboards of their era. This tutorial takes you from zero to a real, measured merge on your own CPU: no GPU, no training, no API keys.
1. Why merging works (the 90-second theory)
Two models that share an architecture and a common ancestor — say, a base model and its instruction-tuned sibling — live in the same broad basin of the loss landscape. Their weights are nearby points, not random ones. Blending those points usually produces a working model, because the path between them stays in low-loss territory. Three ideas built on this observation cover nearly every merge you'll ever run:
- Linear averaging ("model soups"). Take the element-wise weighted average of the weights. Wortsman et al. (2022) showed that averaging fine-tuned checkpoints of the same run improves accuracy with zero extra inference cost. Simple, but the midpoint can sag when the endpoints are far apart.
- SLERP. Interpolate along the great-circle arc between the two weight vectors instead of the straight chord. The arc preserves weight magnitudes, and magnitude matters: a transformer's activations are calibrated to the scale of its weights, so naive averaging can quietly degrade quality even when the direction is right.
- Task arithmetic. Compute a task vector — the difference between a fine-tuned model and its base — and treat it as a portable bundle of the skill that fine-tuning added. Add fractions of task vectors back onto a base to assemble new capability mixes. TIES, DARE, and DELLA all extend this idea by sparsifying the vectors first, so unrelated skills interfere less (Ilharco et al., 2022; Yadav et al., 2023; Yu et al., 2023).
MergeKit implements all of these and more — linear, SLERP, NuSLERP, multi-SLERP, Karcher mean, task arithmetic, TIES, DARE, DELLA, breadcrumbs, SCE, model stock, passthrough "frankenmerging" — behind one YAML config format and one command, mergekit-yaml. It streams tensors through the arithmetic out-of-core, so a merge runs on CPU or with as little as 8 GB of VRAM.
2. What you'll need
- Python 3.10+ and
pip. Everything below ran on a 2-core CPU-only Linux VM — no GPU anywhere. - About 3 GB of free disk: two half-billion-parameter models plus the merges.
- Two shape-compatible models. Our pair:
Qwen/Qwen2.5-0.5B(the base) andQwen/Qwen2.5-0.5B-Instruct(its instruction-tuned sibling). Same architecture, same tokenizer, common ancestor — the safest possible recipe. - MergeKit itself:
pip install mergekit. This run used mergekit 0.1.4 (LGPL-3.0). One real-world snag, documented honestly in the gotchas: 0.1.4's CLI can crash on startup with a pydanticPydanticUserErroraboutConfiguredModuleArchitecturenot being fully defined. The fix is a three-line wrapper that importstorchand callsmodel_rebuild()on mergekit's models before invoking the CLI entry point — we ran every merge below through it.
The golden rule of merging, worth stating before you touch anything: same architecture, same vocabulary, shared ancestor. Merging a base with its own fine-tune is the safest recipe; merging two fine-tunes of the same base is the classic use case. Cross-family merges need extra machinery and fail silently more often than you'd like — fluent text, subtly wrong behavior. Always evaluate the merged model on the tasks you care about, because weight-space arithmetic gives no guarantees.
3. Install MergeKit and learn the config format
pip install mergekit
mergekit-yaml --help # mergekit-yaml CONFIG_FILE OUT_PATH [--cuda] [--allow-crimes] ...
Every merge is a YAML document with four load-bearing fields:
merge_method— one oflinear,slerp,task_arithmetic,ties,dare_ties, and friends.models— the inputs, each with optional per-modelparameters(weights, densities).base_model— required for the task-vector methods and SLERP; the anchor everything is measured against.parameters— method knobs:tfor SLERP (0 yields the base, 1 the other model),weight/lambdafor task arithmetic,densityfor the sparsifying methods.
Two more fields control precision and tokenizers: dtype casts inputs before merging (omit it and matching bfloat16 weights stay bfloat16; linear and SLERP compute in float32 internally and cast back), and tokenizer: {source: union} decides the output vocabulary. Our pair shares a tokenizer, so union looks like a formality — but it wasn't, quite: see the gotchas.
4. Sanity check: rebuild the instruct model with task arithmetic
Before doing anything interesting, verify the machinery. Task arithmetic says: merged = base + λ · Σ weightᵢ · (modelᵢ − base). With a single model, weight 1.0, and λ = 1.0, that formula reduces to the instruct model itself. If MergeKit's implementation is correct, the output should be near-identical to Qwen2.5-0.5B-Instruct. Save this as taskarith.yml:
merge_method: task_arithmetic
base_model: Qwen/Qwen2.5-0.5B
models:
- model: Qwen/Qwen2.5-0.5B-Instruct
parameters:
weight: 1.0
parameters:
lambda: 1.0
dtype: bfloat16
tokenizer:
source: union
mergekit-yaml taskarith.yml ./merged-taskarith # 108 seconds on our 2-core CPU box
Then we compared the merged weights against the real instruct model, tensor by tensor, straight from the safetensors files:
| Tensor | ‖instruct‖ | ‖merged − instruct‖ | Relative error |
|---|---|---|---|
layers.0.attn.q_proj.weight | 60.68 | 0.0032 | 5 × 10⁻⁵ |
layers.12.mlp.down_proj.weight | 38.76 | 0.0023 | 6 × 10⁻⁵ |
embed_tokens.weight (overlap) | 175.00 | 0.0161 | 9 × 10⁻⁵ |
The distances are floating-point dust from the float32-then-cast-back round trip, exactly as the docs describe. The machinery works. This identity merge is the cheapest possible validation that your install, your config schema, and your mental model all agree — run it first whenever a merge behaves strangely later. If the identity merge doesn't reproduce, the problem is your setup, not your method.

5. The real merge: SLERP halfway between base and instruct
Now the interesting one. SLERP with t: 0.5 walks the great-circle arc halfway from the base model to the instruct model:
merge_method: slerp
base_model: Qwen/Qwen2.5-0.5B
models:
- model: Qwen/Qwen2.5-0.5B
- model: Qwen/Qwen2.5-0.5B-Instruct
parameters:
t: 0.5
dtype: bfloat16
tokenizer:
source: union
mergekit-yaml slerp.yml ./merged-slerp # 77 seconds on CPU
Note the base_model line: SLERP takes exactly two models, and one must be designated the base. The output directory contains the merged model.safetensors plus a generated README.md and the mergekit_config.yml that produced it — free provenance for your model card.
Does the arc actually beat the chord here? We measured both. L2 norms for two representative matrices:
| Tensor | ‖base‖ | ‖instruct‖ | ‖SLERP t=0.5‖ | ‖linear 0.5/0.5‖ |
|---|---|---|---|---|
layers.0.attn.q_proj.weight | 60.68 | 59.80 | 60.29 | 60.20 |
layers.12.mlp.down_proj.weight | 38.76 | 38.04 | 38.42 | 38.37 |
Honest reading: for this pair the difference is small — base and instruct are close relatives, so the chord barely cuts the corner and SLERP's norm preservation buys you ~0.15%. The arc earns its keep when the endpoints are far apart (two different specialists, say), where naive averaging visibly shrinks every matrix and the generations go dull. Also note the SLERP midpoint sits at nearly equal distance from both parents (2.16 vs 2.19 on q_proj) — a true geometric midpoint, which is exactly what t=0.5 promises.

6. Measure the merge: perplexity and real generations
Weight-space prettiness means nothing without behavior. We scored four checkpoints — base, instruct, our task-arithmetic rebuild, and the SLERP midpoint — on two probes. First, perplexity on a fixed 647-token held-out passage about model merging (lower is better):
| Model | Perplexity (647 tokens) |
|---|---|
| Qwen2.5-0.5B (base) | 25.542 |
| Qwen2.5-0.5B-Instruct | 25.762 |
| Task-arithmetic rebuild | 25.762 |
| SLERP t=0.5 | 25.421 |
Two things to read here. The rebuild lands exactly on the instruct model's perplexity — the identity check holds at the behavior level too, not just in weight space. And the SLERP midpoint is the best of the four: a genuine interpolation that landed in a slightly better spot than either parent, the "free lunch" the merging literature keeps reporting.
Second, greedy generations (temperature 0, 40 new tokens) on three plain-continuation prompts, identical for every model — no chat template, so nobody gets a formatting advantage:
| Prompt | Base | Instruct | Rebuild | SLERP |
|---|---|---|---|---|
The capital of France is | "Paris. It is the largest city in Europe and the second largest in the world…" | "Paris. It was founded in 789 AD by Charlemagne…" | "Paris. It is the largest city in Europe and the third largest city in the world…" | "Paris. It is the largest city in Europe and the third largest in the world…" |
Summarize in one sentence: The Eiffel Tower… | Continues with tower facts (ignores the instruction) | "…Summary: The Eiffel Tower, a massive iron lattice tower, stands as…" | Continues with tower facts | "It was built by Gustave Eiffel, a French engineer, and is located in Paris…" |
Once upon a time, in a land of talking robots, | "…a young robot named R2-D2… always ready to help others." | "…a young robot named R2-D2… love of exploring the galaxy." | "…a young robot named R2-D2… playful nature…" | "…a young inventor named Alex. Alex had a unique idea for a robot…" |
What this actually shows, stated carefully:
- The instruct model's
Summary:marker on prompt 2 is instruction-tuning leaking through even in plain continuation — the one place its training clearly shows. The merged models sit between base and instruct in style, which is what an interpolation should do. - The rebuild and the true instruct model have identical perplexity yet diverge on some greedy generations (compare "second largest" vs "third largest"). Tiny weight dust flips close token races — generation is chaotic under small perturbations, so never treat one greedy sample as proof of identity. Measure distributions, not anecdotes.
- The SLERP midpoint invents its own story character ("Alex") on prompt 3 instead of copying either parent's R2-D2. Interpolated weights can produce genuinely intermediate behavior, not just a coin flip between parents.
- And the necessary caveat: these are 0.5B models. Every one of them states false facts with total confidence (Paris was not founded in 789 AD by Charlemagne). This tutorial's models are plumbing for demonstrating the mechanics — repeat the recipe at 7B+ and the same pattern shows up with far more fluent text.
7. Which approach should you use?
| Situation | Method | Why |
|---|---|---|
| Averaging checkpoints of one fine-tuning run | linear | Model-soup territory; endpoints are close, the chord is fine, and it's the cheapest option. Weights default to normalizing to sum to 1. |
| Blending two distinct models (base ↔ instruct, chat ↔ code) | slerp with t to taste | Preserves weight magnitudes along the arc; start at 0.5 and walk t in 0.1 steps. |
| Combining skills from several fine-tunes of one base | ties / dare_ties | Task vectors + sparsification + sign consensus tame interference; tune density (fraction of weights kept) per model. |
| Dialing one behavior up or down | task_arithmetic | Scale a single task vector with weight/lambda; the identity check in step 4 is your debugger. |
| Stitching layers from different models | passthrough slices | "Frankenmerging": early layers from one model, late layers from another. Powerful, fragile — validate hard. |
Two parameters deserve a final word. density (TIES/DARE) is the fraction of each task vector retained after pruning — 0.5 is a sane start; lower values merge more models with less interference but wash out subtle skills. t (SLERP) is not a linear blend knob: because the arc preserves norms, t=0.5 is a true geometric midpoint, and small steps near the endpoints move behavior faster than steps in the middle.
8. Gotchas that bite everyone once
- The tokenizer union can resize your embeddings. Our models' configs claimed a 151,936-token vocabulary, but the actual tokenizers only define 151,665 tokens. MergeKit aligned the embedding matrices to the union vocabulary, so the merged model has 151,665-row embeddings. The 271 dropped rows are token IDs the tokenizer never emits — harmless here, and the merged config + tokenizer are self-consistent — but if you diff merged weights against a parent, compare overlapping rows or you'll get a shape error. This is the
tokenizer:block doing its documented job. - mergekit 0.1.4's CLI can crash before it starts. We hit
PydanticUserError: ConfiguredModuleArchitecture is not fully definedon a fresh install. Workaround: importtorchfirst, callmodel_rebuild()on mergekit's pydantic models, then invoke the CLI. If a future release fixes the import order, delete the wrapper. - Chat templates don't merge themselves. MergeKit copies a template (
chat_template: "auto"picks the most common among inputs; ours wrote achat_template.jinjainto the output). If your parents use different formats, decide explicitly which one the child speaks. - bf16 in, bf16 out — unless you say otherwise.
dtypecasts inputs before merging;out_dtypecasts after. Forgetting this is how you accidentally ship a float32 model twice the size you expected. - License hygiene. A merge is a derivative work of every parent. If a parent's license restricts redistribution, your merge inherits the restriction. Check before uploading.
- Evaluate, always. Merging gives no behavioral guarantees. Perplexity on a held-out passage plus a handful of greedy generations — the two probes in this tutorial — take minutes on CPU and catch most silent failures.
The takeaway
Model merging turns the open-weight ecosystem into composable parts: take an instruction-tuned model and a base, or two specialists, and blend them with a config file instead of a training run. The workflow that actually works is boring in the best way — verify the identity merge first, pick SLERP for distant pairs and linear for close ones, reach for TIES/DARE when combining several fine-tunes, and measure everything on real generations. On a CPU-only machine, the whole loop from pip install mergekit to a measured SLERP merge took about twenty minutes of compute. The next time you catch yourself thinking "I wish there were a model halfway between these two" — there can be, by lunchtime.