Why this is worth your time

Your LLM says it is 90% sure. That number is not a probability — it is a ranking with a costume on. Reorder the options in a multiple-choice prompt and raw logits flip their answer 23% of the time. And when every "0.9" gets routed to a human because nobody trusts it, your "AI automation" is a demo, not a deployment.

The metric that actually matters is auto-decidable traffic: the share of decisions you can safely automate at a stated error rate. On the standard BANKING77 benchmark, raw logits let you automate 7.7% of traffic at ≤5% error. The tool in this tutorial takes that to 52% — with zero to a few hundred labels, no fine-tuning, and nothing generated. It is AnyJev, from a team at Nokia's applied-research lab, and it collected ~995 stars in about ten days for good reason: it turns any open LLM into a Jev-style decision model.

What that means in practice: you ask a typed question — a choice, a yes/no, a score — and get back a decision with a probability you can threshold. No parsing "Answer: B" out of prose. No hoping the model's self-reported confidence means something. One prefill of the next-token distribution, read properly.

Diagram of AnyJev's four levels: raw uncalibrated readout, L0 debiasing over option rotations, L1 temperature scaling, L2 closed-form head on a mid-layer hidden state
AI-generated diagram for AI Frontier Post — raw → L0 → L1 → L2

What you'll need

  • Python 3.10 or newer. The base package depends on numpy alone; the Hugging Face extras add torch and transformers.
  • No GPU required to follow along. The packaged demo runs on a synthetic model with numpy only, in seconds. The full serve path (truncating and serving a real model) needs a GPU with vLLM — that section is verified against the project's own docs rather than run live here.
  • A clone of the repo for the demo: it is Apache-2.0, by Jiamu Zhang, Tianze Yang, Yucheng Shi and Liang Wu (Nokia Sunnyvale, plus Tencent Hunyuan) — and, as the README states, not affiliated with TypeSafe AI or Jev.

Step 1 — Install and run the zero-hardware demo

AnyJev's library installs like anything else. The demo ships inside the repo, so clone it:

pip install anyjev
git clone https://github.com/nokia-applied-research/AnyJev
cd AnyJev
python -m demo.jev_mode --backend fake

The fake backend runs the entire Jev-mode demonstration on a synthetic model — numpy only, no download, under a second. It walks four sections: [1/4] the levels on held-out states, L0 zero-label vs the L2 head; [2/4] the same question reworded and re-listed with no new labels; [3/4] a new head fit from a few labels; [4/4] an export-then-reload round trip. The demo is honest about its own numbers: the synthetic model's numbers are planted, and it says so on screen. The real numbers are in docs/jev_mode.md, measured at runtime.

What you are seeing in section [1/4] is the whole pitch. Everything else in this tutorial is the machinery behind it.

Step 2 — Understand the four levels

Every decision AnyJev returns carries a level. Here is the contract, with the measured numbers (Qwen3-8B, BANKING77 20-way, 300 test items):

  • raw — a restricted softmax over the label tokens at the answer position. This is what every open Jev clone does: max_tokens=1 plus logprobs. What you get: a ranking. Two biases are baked in: the model prefers some labels regardless of input ("prior bias") and some positions in the option list ("position bias" — flip the order and the argmax can change). The project says it plainly: do not ship raw.
  • L0 — zero labels. Two training-free corrections: cyclic-shift marginalization (for a K-option choice, read all K rotations so every option sits at every position once, combined via geometric mean in log space — if position bias is additive in logit space, this removes it exactly) and prior correction (estimate the model's label prior without labels, divide it out; default is batch calibration at strength 0.75, warming up over 8 items per question). Measured: order-flip rate 0.230 → 0.073, accuracy 0.747 → 0.803, calibration error (ECE) 0.240 → 0.184.
  • L1 — 100–500 labels. Temperature scaling on top of L0. It does not change the ranking; it fixes the uncertainty: ECE 0.184 → 0.095.
  • L2 — 100–300 labels. A closed-form head per question, solved in seconds on the hidden state partway down the model — no gradients, the model's weights untouched. It costs less than one plain forward pass, and it is the row that moves the automation number: auto-decidable at ≤5% error goes 7.7% → 52.0%.

Downstream code can enforce quality with require="L1", which refuses to act on a weaker-level decision. That is the mechanism that turns "a probability" into "a threshold I can automate against".

Step 3 — Truncate the model (keep ~two thirds)

Here is the counterintuitive part. A decision does not need the whole model: the head's accuracy is flat from about two thirds of the depth, because the upper blocks are busy converting the answer into token space rather than deciding anything. So AnyJev ships a truncator that writes the blocks it doesn't need away into an ordinary smaller checkpoint — vLLM, transformers, quantizers and GGUF converters all take it unchanged:

pip install "anyjev[hf]"
python -m anyjev.truncate Qwen/Qwen2.5-7B-Instruct 18 ./qwen-b18

On the project's own banking20 measurement, 18 of 28 blocks is 1.2–1.4× faster and two points more accurate (0.850 vs 0.830 at full depth), because a middle block is a better feature space for a linear head than the final one. Depth, the docs note, is usually a gain, not a trade.

One flag the project explicitly steers you away from: --quantization fp8 is available and not recommended — it buys single-question latency and costs accuracy.

Step 4 — Serve it with vLLM's embed server

L2 reads a hidden state, so the server is an embed server whose pooler hands the state back untouched. One command:

vllm serve ./qwen-b18 --task embed \
  --override-pooler-config '{"pooling_type":"LAST","normalize":false,"softmax":false}'

A deployment at L2 is, in the project's words, "a pooling server plus a few kilobytes of head": no logits, no parsing, no patched engine, nothing generated. And the head is portable across engines — one fit through transformers and served by vLLM answers 99.0% identically, with mean |Δp| of 0.0011 on BANKING77-20.

Step 5 — Ask a typed question (L0, zero labels)

With the server up, the API is three objects — a backend, a typed question, a decider:

from anyjev import Decider, Question
from anyjev.backends.vllm import VLLMBackend

d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"))
route = Question.choice("Which team should handle this?",
                        ["billing", "technical", "sales", "other"], name="route")

d.decide(ticket, [route])["route"].distribution  # {"billing": 0.81, "technical": 0.07, ...}

Nothing is generated and nothing is parsed: the distribution is read from one prefill of the next-token distribution, position bias averaged out over the option rotations, the label prior divided away. Questions come in three types — choice, noul (yes/no), and score (ordinal bins, never permuted).

Diagram of the AnyJev deployment pipeline: a model truncated at two thirds of its depth, served by a vLLM embed server, with a small closed-form head reading hidden states into calibrated decisions
AI-generated diagram for AI Frontier Post — truncate, serve, decide

Step 6 — Add the rotation budget for 2.2× throughput

L0's safety costs prefills: one per option rotation. For an 18-option choice that is 18 reads. The rotation budget — recommended for any K-option choice — reads rotations one at a time and stops when the leader is far enough ahead, with the threshold calibrated so the answer matches the full cycle's a certified fraction of the time:

d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), adaptive_shifts=True)
d.calibrate_adaptive(route, unlabelled_tickets, target=0.01)  # a few hundred states, no labels
d.decide_batch(tickets, route)      # diagnostics: shifts_used, stop_threshold

The measurement: 7.2 rotations instead of 18 at a certified 1% disagreement rate — 2.2× the decisions per second on vLLM, 2.3–2.7× on the transformers backend, accuracy unchanged. It is opt-in in version 0.2.0 only because the docs tables were measured before it existed. Note that calibrate_adaptive needs no labels at all: the budget is calibrated against the project's own full-strength readout, not against human judgments.

Step 7 — Fit an L2 head from 100–300 labels

With a hundred or so labeled decisions for your question, you get the top level — one closed-form solve:

d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), level="L2")
d.fit_head(route, states, labels, layers=[-1])  # 100-300 labels, seconds, no gradients

On LocalLLaMA/typed-decisions (20 questions × 300 labels, 2,000 held-out decisions), L2 over L0 zero-label: Qwen3-1.7B 0.494 → 0.730 at block 18/28, costing 0.70× of one forward; Qwen3-4B 0.564 → 0.786; Qwen3-8B 0.647 → 0.771; the 30B-A3B 0.630 → 0.799. Pooled calibration error: 0.03–0.05. A 1.7B at 64% of its depth reaches the number Jev publishes; a 4B ties the fine-tuned 421M Laya (0.768). One hundred labels already put the 8B head at 0.740.

If you are running one of the five Qwen3 models the project ships heads for, you can skip the fitting entirely: anyjev-heads/ carries 23 heads per model in a single 1.8–4.4 MB file.

Step 8 — Measure the whole thing on your box in one command

Don't trust their numbers — the project agrees. One command truncates, serves, fits a head, measures accuracy, ECE and milliseconds-per-decision on held-out states, tears the server down, and repeats at full depth so there is something to compare against:

python -m anyjev.pipeline Qwen/Qwen2.5-7B-Instruct --labels-from banking20

Timings are a median over --repeats passes with the spread printed next to them — the docs are explicit that on a shared machine a single pass can report the same configuration as both faster and slower than baseline. That level of measurement honesty is why the numbers in this tutorial are worth quoting.

Step 9 — Go live incrementally with observe()

The production pattern the project is designed around: day 0 runs at L0 with no labels, and the head arrives when the loop has collected enough:

d.observe(q, state, label)  # solves the head at 30 labels, re-solves at 60, 120, ...

And the head maintains itself afterwards: only its feature mean and scale move, re-estimated from unlabelled traffic, so it follows its question across rewordings and option orders on its own. A reworded question drops the Qwen3-8B head from 0.77 to 0.65–0.70, and 30 unlabelled requests bring it back to 0.74–0.75 — against 0.77 for a full relabelled refit. New labels are only needed for a new question.

What you built

A calibrated typed-decision endpoint over your own open model: L0 debiasing with zero labels from the first request, an optional rotation budget for 2× throughput, and a closed-form L2 head once you have ~100 labels — one that auto-decides half your traffic at ≤5% risk instead of 7.7%. The share of traffic you can automate, not accuracy, is the metric this project optimizes, and it's the right one.

Honest limitations

The README carries its own limitations section, and it deserves a read — but the highlights:

  • L2 is per question and per model. A head fit on one question does not transfer to another, and only Qwen3 heads ship. Heads need hidden states, which transformers and a vLLM embed server provide; other engines don't yet (SGLang included).
  • "Accuracy" on typed-decisions is agreement with a teacher LLM — the gold is the mean of three samples of one model, and a fresh sample of that teacher agrees with it only 0.735 of the time. Calibrated disagreement with the teacher is not necessarily error.
  • Calibration cannot fix a model that cannot answer. On maze edges and Minesweeper, no readout beats the trivial baseline.
  • L0's prior correction is not a free win everywhere — the batch prior hurts when one label dominates the true marginal. The shift marginalization is the safe half; the docs quantify where each half helps.
  • At most 26 options in the letter readout (a span readout is on the roadmap, not in the code), and the coverage-at-5%-risk figures are high-variance estimates at n=300. Every decision here is scored in isolation, not inside a real agent loop.
  • Reproduce the numbers before trusting them. The project regenerates every number from committed JSON (bash scripts/regen_docs.sh), and a second run from a clean checkout reproduced every zero-label number bit for bit. python -m anyjev.pipeline exists so you can run that check on your own hardware.

The takeaway

Raw logits are a ranking wearing a probability's clothes. AnyJev does the unglamorous work that converts them into something you can automate against: average out position bias over cyclic shifts (L0, zero labels), divide out the label prior (L0), fix the uncertainty with temperature scaling (L1, a few hundred labels), then fit a closed-form linear head on a mid-layer hidden state that costs less than a forward pass (L2). Two-thirds of the model is plenty; a 1.7B at 18 blocks does the job of a flagship's confidence score.

If you route tickets, triage alerts, score leads, or run any LLM-as-a-judge loop where a wrong auto-decision has a cost, this is the calibration layer you were missing. The repository is nokia-applied-research/AnyJev; run the fake-backend demo first, then anyjev.pipeline on your own model before you believe anyone's numbers — including the ones in this article.