AI Frontier Post
AnyJev banner: turn any LLM into a Jev-style decision model — typed decisions, real probabilities
AnyJev’s pitch: typed decisions with real probabilities, read from the model’s next-token distribution instead of generated text. Banner from the project’s repo (Apache-2.0).

Every routing agent, triage bot, and moderation queue eventually does the same fragile dance: prompt the model to “reply in JSON,” regex the label out of the text, and retry when the output drifts. AnyJev takes the problem out of the generated text entirely. You hand the model a state and a typed question — a choice, a yes/no, or a score — and the answer is read as a probability distribution over the option labels, straight from the model’s next-token distribution. Nothing is generated, nothing is parsed.

AnyJev (nokia-applied-research/AnyJev, ~1,100 GitHub stars, Apache-2.0 license) comes from Nokia’s applied research team, and its own numbers make the case. On Qwen3-8B over 300 BANKING77 20-way intent items: raw label logits flip their answer 23.0% of the time when you reverse the option order; AnyJev’s L0 readout drops that to 7.3%, accuracy rises from 74.7% to 80.3% — and the share of decisions confident enough to automate at ≤5% error jumps from 7.7% to 46.3%, with zero labels. Here is the full walkthrough, verified against the project’s own README and docs.

What you’ll need

1. Install the package

The core library is one pip install:

pip install anyjev

Two extras cover the serving paths, and you install exactly one depending on your setup — anyjev[client] for talking to a running vllm serve, or anyjev[vllm] to run vLLM in the same process:

pip install "anyjev[client]"
# or:
pip install "anyjev[vllm]"

2. Make your first typed decision in one forward pass

The fastest route uses Tacit, AnyJev’s self-distilled checkpoints (1.7B to 9B on Hugging Face). They were trained to give, in one forward pass, the answers the base model reaches when it reasons — no human labels, no teacher model involved:

from anyjev import Tacit

tacit = Tacit.from_pretrained("morriszjm/Tacit-9B")          # transformers, one GPU
d = tacit.decide(state="Customer: my package was due last Monday and it still has not arrived.",
                 question="What does the customer want?",
                 options=["track_order", "cancel_order", "refund", "change_address"])
d["answer"], d["probs"], d["route"]    # an option, {option: probability}, "one_forward" or "cot"

kind defaults to "choice"; "yes_no" and "score" (options are ordered levels, lowest first) cover the other shapes. Batching is built in: decide_batch([...]) takes a list of such dicts. The route field tells you how the answer was read — "one_forward" or "cot" — so downstream code can check before it acts.

3. Go training-free on any open model

No Tacit checkpoint for your domain? The Decider reads the same typed decisions from any open causal LM, with no training and no labels. The default level is L0: the options are read in each of their rotations so every option sits at every position, the rotations are combined in log space, and the model’s label prior is divided out — all estimated without a single label. That is what kills the 23% order-flip rate:

from anyjev import Decider, Question
from anyjev.backends.hf import HFBackend          # or anyjev.backends.vllm.VLLMBackend(url, model)

d = Decider(HFBackend("Qwen/Qwen3-8B"))           # level="L0" by default
route = Question.choice("Which team should handle this?",
                        ["billing", "technical", "sales", "other"], name="route")
r = d.decide("My card was charged twice for one order.", [route])["route"]
r.distribution, r.level                           # {option: probability}, "L0"

Enforcement comes free: decide(..., require="L0") raises LevelError when a decision lands below the asked level, so a degraded readout can’t silently become an automated action. Every decision also records diagnostics — answer_mass (the probability the model put on the label tokens at all; a low value means the prompt did not land), order_flip_raw, permutations, and the prior method used.

4. Send only the hard calls to reasoning — with a cap

Reasoning helps the close calls and wastes budget on the easy ones. The adaptive mode sends a decision back through with thinking turned on only when the log-probability gap between the top two options falls below tau. The answer is still read as a label distribution, so it cannot fail to parse — and a hard cap bounds the cost:

tacit = Tacit.from_pretrained("morriszjm/Tacit-9B", adaptive=True, tau=0.5, max_cot_share=0.2, cot_window=1000)

Every escalation is reported (route == "cot", with the first pass kept), and tacit.stats counts them. The training-free Decider has its own budget version: Decider(adaptive_shifts=True, canonical_order=True) reads rotations until the leader is far enough ahead, and calibrate_adaptive sets that threshold against the full-cycle answer — no labels — at a stated disagreement rate. The project reports that at a certified 1% disagreement it read 7.2 rotations instead of 18: 2.2× the decisions per second on vLLM, accuracy unchanged.

AnyJev demo: reversing the option order flips the raw logit answer, while the L0 readout stays put
Position bias, killed: reverse the option order and the raw readout flips — L0 gives the same answer both ways. Demo frames from the project’s README (Apache-2.0).

5. Serve it as a decision endpoint

When one cap has to serve many clients, AnyJev ships a gateway. Point a stock vllm serve at a Tacit checkpoint, then run the server module in front of it:

pip install "anyjev[client]"
vllm serve morriszjm/Tacit-9B --host 127.0.0.1 --port 8000
python -m anyjev.serve --model morriszjm/Tacit-9B --upstream http://127.0.0.1:8000 --adaptive --port 8100
curl -s 127.0.0.1:8100/v1/decide -H 'Content-Type: application/json' \
  -d '{"state": "Customer: my package has not arrived.", "question": "What does the customer want?", "options": ["track_order", "refund"]}'

POST /v1/decide takes one decision or {"items": [...]}; GET /v1/stats reports the escalation share across clients. Note the security posture stated in the README: the gateway listens on 127.0.0.1 and has no authentication of its own — keep it on loopback or put your own auth in front.

AnyJev's order-flip demo, final frame: L0 holds the same answer while the raw readout changes
Every number in the demo is a model output, not a mock-up — the project’s own words. Demo frames from the project’s README (Apache-2.0).

What you built

A typed-decision service with real probabilities. Instead of prompting for JSON and parsing the reply, you ask choice, yes_no, or score questions and get back an option, a probability distribution, and a readout level — plus the diagnostics (answer_mass, rotation counts, prior method) to know when to trust it. Escalation is capped and reported, so the “let it think harder” path can’t eat your budget silently.

Honest limitations

Related articles

Tutorials

Give your coding agents an org chart: hands-on with Paperclip, the 98K-star agent orchestrator

October 6, 2026 · 8 min read
Tutorials

Make your coding agent answer with pages, not walls of text: a hands-on guide to the Answer me with HTML skill

October 6, 2026 · 8 min read
Tutorials

Run your own voice studio on your own machine: hands-on with VoiceStudio

October 6, 2026 · 8 min read