nokia-applied-research/AnyJev

Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)

📦 AnyJev – Turn any LLM into a calibrated decision model

AnyJev (Any Jev) is a Python library that lets you ask an open‑source LLM a typed question (e.g. a multiple‑choice, yes/no, or numeric rating) and receive a probability‑based decision directly from the model’s next‑token logits. It does not generate text, does not require fine‑tuning, and can work with zero labelled examples. With a few hundred labels you can upgrade to a higher‑accuracy “L2” head that runs in essentially the cost of a single forward pass.


🎯 What problem does it solve?

  • Raw logits from LLMs are unstable – flipping the order of answer options often flips the model’s prediction, and the confidence scores are poorly calibrated.
  • Because the scores cannot be trusted, most pipelines fall back to human review, wasting resources.
  • AnyJev provides three levels of calibration:
    • L0 – zero‑label correction that removes option‑order bias.
    • L1 – temperature scaling using 100‑500 labels per question.
    • L2 – a closed‑form linear head (shrunken LDA / ridge) trained on 100‑300 labels, giving calibrated probabilities at the cost of a single truncated forward.

🚀 Quick start (from the README)

pip install "anyjev[hf]"
from anyjev import Decider, Question
from anyjev.backends.hf import HFBackend

d = Decider(HFBackend("Qwen/Qwen3-8B"))

# define typed questions
route = Question.choice(
    "Which team should handle this?",
    ["billing", "technical", "sales", "other"],
    name="route"
)
risky = Question.noul("Is this tool call destructive?", name="risky")
done  = Question.score("How complete is the task?", bins=5, name="done")

# run a decision
state = {"conversation": [...], "tool_call": {...}}
result = d.decide(state, [route, risky, done])
print(result["route"].distribution)   # {'billing': 0.81, 'technical': 0.07, ...}
print(result["risky"].p_true)          # 0.12
print(result["done"].value)            # 0.35
print(result.level)                     # "L0"

Add labels later to move to L1/L2:

# collect 100‑500 labelled examples for a question
d.calibrate(risky, states, labels)      # L1 – temperature scaling
d.fit_head(route, states, labels)       # L2 – closed‑form head

d.save_artifacts("qwen3-8b.json")      # store the head (≈100 KB)

Now d.decide(..., level="auto") will automatically use L2 when a head exists, otherwise fall back to L1 or L0.


🧠 How it works (high‑level)

Level Labels needed What it does What it doesn’t do
raw none Returns the soft‑max over the token logits for each option (the naïve baseline). Any bias correction or calibration.
L0 none Rotates the option list K times, averages the logits, and divides out the estimated label prior. Removes order‑flip bias. Does not make the model’s uncertainty statistically calibrated.
L1 100‑500 per question Fits a temperature scaling on top of L0. Does not change the ranking of options.
L2 100‑300 per question Solves a closed‑form linear head (shrunken LDA / ridge) on the hidden state at ~⅔ of the model depth. One prompt per input, then a cheap matrix multiply at serve time. Cannot be transferred to a different question or model without re‑fitting.

L2 heads are tiny artifacts (≈100 KB) containing a [hidden, K] matrix, bias, standardisation vector and temperature. They are computed in seconds on a CPU and then used for inference with a single truncated forward (e.g., stop at block 24 of a 36‑block model).


📊 Reported results (Qwen3‑8B on the BANKING77 20‑way benchmark)

Metric Raw logits AnyJev L0 (zero labels) AnyJev L1 (100‑500 labels) AnyJev L2 (100‑300 labels)
Answer‑order flip rate 0.230 0.073 0.077 –
Accuracy 0.747 0.803 0.807 –
Expected Calibration Error 0.240 0.184 0.095 –
Auto‑decidable @ ≤5 % error (fraction of traffic you can safely automate) 7.7 % 46.3 % 52.0 % –

Latency: L2 costs ≈0.68× a full forward on Qwen3‑8B (one prompt stopped early). L0 needs K prefills (≈0.25 s per decision for K=20 on an H100).


🛣️ Roadmap (as of the README)

  • ✅ Implement choice, noul, score questions with a single prefill.
  • ✅ L0 (zero‑label) and L1 artifacts; level enforcement.
  • ✅ L2 closed‑form heads, automatic label‑free adaptation, observe loop.
  • ✅ Pre‑built heads for five Qwen3 models and a demo CLI.
  • ⏳ Speed optimisations to make every decision cheaper.
  • ⏳ Support for serving via vLLM / SGLang (using the residual stream at a fixed block).
  • ⏳ Full agent‑loop evaluation (replace the LLM inside a real agent).
  • ⏳ Public hub for heads, an interactive Space, and a technical report.
  • ⏳ Extend to more base models (Llama, Gemma, Mistral, DeepSeek) and larger option sets (>26).

⚠️ Limitations (explicitly listed)

  • Accuracy is measured against a teacher LLM, not a human gold standard.
  • L2 heads are per‑question and per‑model; they do not transfer across questions or to other models (currently only Qwen3 heads are shipped).
  • Calibration cannot compensate for a model that simply cannot answer a task (e.g., maze navigation, Minesweeper).
  • L0 may hurt accuracy when one label dominates the prior.
  • The library currently supports at most 26 options per choice (span readout is planned).
  • Coverage estimates at 5 % risk are high‑variance with n = 300 test items.
  • All decisions are evaluated in isolation, not inside an end‑to‑end agent loop.

📦 Installation & licensing

  • Install with optional Hugging‑Face backend: pip install "anyjev[hf]".
  • The package is released on PyPI (anyjev), supports Python 3.8+.
  • License: Apache‑2.0.

📚 Further reading & citation

  • Full method description: docs/method_v3.md.
  • Benchmarks: docs/results_bench.md, docs/results_exit.md.
  • Demo scripts: demo/jev_mode, demo/games (2048, Minesweeper).
  • Cite as:
@software{anyjev2026,
  title  = {AnyJev: Turn any LLM into a Jev-style decision model},
  author = {Zhang, Jiamu and Yang, Tianze and Shi, Yucheng and Wu, Liang},
  year   = {2026},
  url    = {https://github.com/nokia-applied-research/AnyJev}
}

Bottom line: AnyJev provides a practical, zero‑label‑required way to turn the raw, noisy logits of any open‑source LLM into trustworthy, probability‑calibrated decisions, and a lightweight closed‑form head (L2) that brings near‑state‑of‑the‑art accuracy with minimal inference cost.

Related

  • Project
  • Project
  • Project
  • Project