nokia-applied-research/AnyJev
Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)
📦 AnyJev – Turn any LLM into a calibrated decision model
AnyJev (Any Jev) is a Python library that lets you ask an open‑source LLM a typed question (e.g. a multiple‑choice, yes/no, or numeric rating) and receive a probability‑based decision directly from the model’s next‑token logits. It does not generate text, does not require fine‑tuning, and can work with zero labelled examples. With a few hundred labels you can upgrade to a higher‑accuracy “L2” head that runs in essentially the cost of a single forward pass.
🎯 What problem does it solve?
- Raw logits from LLMs are unstable – flipping the order of answer options often flips the model’s prediction, and the confidence scores are poorly calibrated.
- Because the scores cannot be trusted, most pipelines fall back to human review, wasting resources.
- AnyJev provides three levels of calibration:
- L0 – zero‑label correction that removes option‑order bias.
- L1 – temperature scaling using 100‑500 labels per question.
- L2 – a closed‑form linear head (shrunken LDA / ridge) trained on 100‑300 labels, giving calibrated probabilities at the cost of a single truncated forward.
🚀 Quick start (from the README)
pip install "anyjev[hf]"
from anyjev import Decider, Question
from anyjev.backends.hf import HFBackend
d = Decider(HFBackend("Qwen/Qwen3-8B"))
# define typed questions
route = Question.choice(
"Which team should handle this?",
["billing", "technical", "sales", "other"],
name="route"
)
risky = Question.noul("Is this tool call destructive?", name="risky")
done = Question.score("How complete is the task?", bins=5, name="done")
# run a decision
state = {"conversation": [...], "tool_call": {...}}
result = d.decide(state, [route, risky, done])
print(result["route"].distribution) # {'billing': 0.81, 'technical': 0.07, ...}
print(result["risky"].p_true) # 0.12
print(result["done"].value) # 0.35
print(result.level) # "L0"
Add labels later to move to L1/L2:
# collect 100‑500 labelled examples for a question
d.calibrate(risky, states, labels) # L1 – temperature scaling
d.fit_head(route, states, labels) # L2 – closed‑form head
d.save_artifacts("qwen3-8b.json") # store the head (≈100 KB)
Now d.decide(..., level="auto") will automatically use L2 when a head exists, otherwise fall back to L1 or L0.
🧠 How it works (high‑level)
| Level | Labels needed | What it does | What it doesn’t do |
|---|---|---|---|
| raw | none | Returns the soft‑max over the token logits for each option (the naïve baseline). | Any bias correction or calibration. |
| L0 | none | Rotates the option list K times, averages the logits, and divides out the estimated label prior. Removes order‑flip bias. | Does not make the model’s uncertainty statistically calibrated. |
| L1 | 100‑500 per question | Fits a temperature scaling on top of L0. | Does not change the ranking of options. |
| L2 | 100‑300 per question | Solves a closed‑form linear head (shrunken LDA / ridge) on the hidden state at ~⅔ of the model depth. One prompt per input, then a cheap matrix multiply at serve time. | Cannot be transferred to a different question or model without re‑fitting. |
L2 heads are tiny artifacts (≈100 KB) containing a [hidden, K] matrix, bias, standardisation vector and temperature. They are computed in seconds on a CPU and then used for inference with a single truncated forward (e.g., stop at block 24 of a 36‑block model).
📊 Reported results (Qwen3‑8B on the BANKING77 20‑way benchmark)
| Metric | Raw logits | AnyJev L0 (zero labels) | AnyJev L1 (100‑500 labels) | AnyJev L2 (100‑300 labels) |
|---|---|---|---|---|
| Answer‑order flip rate | 0.230 | 0.073 | 0.077 | – |
| Accuracy | 0.747 | 0.803 | 0.807 | – |
| Expected Calibration Error | 0.240 | 0.184 | 0.095 | – |
| Auto‑decidable @ ≤5 % error (fraction of traffic you can safely automate) | 7.7 % | 46.3 % | 52.0 % | – |
Latency: L2 costs ≈0.68× a full forward on Qwen3‑8B (one prompt stopped early). L0 needs K prefills (≈0.25 s per decision for K=20 on an H100).
🛣️ Roadmap (as of the README)
- ✅ Implement
choice,noul,scorequestions with a single prefill. - ✅ L0 (zero‑label) and L1 artifacts; level enforcement.
- ✅ L2 closed‑form heads, automatic label‑free adaptation,
observeloop. - ✅ Pre‑built heads for five Qwen3 models and a demo CLI.
- ⏳ Speed optimisations to make every decision cheaper.
- ⏳ Support for serving via vLLM / SGLang (using the residual stream at a fixed block).
- ⏳ Full agent‑loop evaluation (replace the LLM inside a real agent).
- ⏳ Public hub for heads, an interactive Space, and a technical report.
- ⏳ Extend to more base models (Llama, Gemma, Mistral, DeepSeek) and larger option sets (>26).
⚠️ Limitations (explicitly listed)
- Accuracy is measured against a teacher LLM, not a human gold standard.
- L2 heads are per‑question and per‑model; they do not transfer across questions or to other models (currently only Qwen3 heads are shipped).
- Calibration cannot compensate for a model that simply cannot answer a task (e.g., maze navigation, Minesweeper).
- L0 may hurt accuracy when one label dominates the prior.
- The library currently supports at most 26 options per choice (span readout is planned).
- Coverage estimates at 5 % risk are high‑variance with n = 300 test items.
- All decisions are evaluated in isolation, not inside an end‑to‑end agent loop.
📦 Installation & licensing
- Install with optional Hugging‑Face backend:
pip install "anyjev[hf]". - The package is released on PyPI (
anyjev), supports Python 3.8+. - License: Apache‑2.0.
📚 Further reading & citation
- Full method description:
docs/method_v3.md. - Benchmarks:
docs/results_bench.md,docs/results_exit.md. - Demo scripts:
demo/jev_mode,demo/games(2048, Minesweeper). - Cite as:
@software{anyjev2026,
title = {AnyJev: Turn any LLM into a Jev-style decision model},
author = {Zhang, Jiamu and Yang, Tianze and Shi, Yucheng and Wu, Liang},
year = {2026},
url = {https://github.com/nokia-applied-research/AnyJev}
}
Bottom line: AnyJev provides a practical, zero‑label‑required way to turn the raw, noisy logits of any open‑source LLM into trustworthy, probability‑calibrated decisions, and a lightweight closed‑form head (L2) that brings near‑state‑of‑the‑art accuracy with minimal inference cost.
Related
- Project
- Project
- Project
- Project