Jev in 25 Lines of Python – How a Minimal Local Classifier Works
Quick takeaway
A 25‑line Python script can mimic the core functionality of Jev – a fast, local decision model – by loading a GGUF LLM, feeding it a prompt with labeled options, and converting the model’s final token logits into probabilities.
What the script does
It loads a GGUF model, formats a prompt with labeled choices, runs a forward pass, extracts the last‑token logits for the choice tokens, and normalizes them into probabilities.
# /// script
# requires-python = ">=3.12"
# dependencies = ["huggingface-hub", "llama-cpp-python", "numpy"]
# ///
import numpy
from llama_cpp import Llama
model = Llama.from_pretrained(
repo_id="Qwen/Qwen3-0.6B-GGUF",
filename="Qwen3-0.6B-Q8_0.gguf",
n_ctx=512,
logits_all=True,
verbose=False,
)
labels = ["A", "B", "C"]
choices = ["Legitimate", "Spam", "Phishing"]
email = "Payroll asks for your password on a non-company sign-in page."
options = "\n".join(
f"{l}. {c}" for l, c in zip(labels, choices, strict=True)
)
prompt = f"""<|im_start|>system
Choose one option.<|im_end|>
<|im_start|>user
Email: {email}\n\n{options}<|im_end|>
<|im_start|>assistant
\n\n"""
model.eval(tokens=model.tokenize(text=prompt.encode(), add_bos=False, special=True))
logits = model.scores[model.n_tokens - 1]
token_ids = [model.tokenize(text=l.encode(), add_bos=False)[0] for l in labels]
choice_logits = numpy.asarray([logits[t] for t in token_ids])
logprobs = choice_logits - numpy.logaddexp.reduce(choice_logits)
probabilities = numpy.exp(logprobs)
for name, scores in (
("Logits", choice_logits),
("Log probabilities", logprobs),
("Probabilities", probabilities),
):
values = numpy.round(scores.astype(float), 3).tolist()
print(f"{name}:", dict(zip(choices, values, strict=True)))
The script prints three dictionaries, for example:
Logits: {'Legitimate': 26.254, 'Spam': 27.262, 'Phishing': 29.614}
Log probabilities: {'Legitimate': -3.482, 'Spam': -2.474, 'Phishing': -0.122}
Probabilities: {'Legitimate': 0.031, 'Spam': 0.084, 'Phishing': 0.885}
Why this matters
It demonstrates that Jev’s essential behavior – turning a natural‑language prompt with discrete options into calibrated probabilities – does not require a proprietary API, synthetic data, or RL‑based post‑training. The entire pipeline runs locally, incurs no network latency, and can be reproduced with any GGUF‑compatible model.
Community insights
Log‑probability caveats
"Going directly for the logprobs is always icky when you use a chat model as base, because they are trained to write prose as output. … you should add clear system instructions or use structured outputs to avoid the model wandering off." – sigmoid10
The comment warns that token‑level probabilities can be distorted if the model generates additional prose before the choice token. Adding explicit system prompts or constraining the assistant’s output format mitigates this risk.
Prompt ordering effects
"Because of masked attention, if you put the options before the body, the transformer already knows what it needs to look for and can allocate more tokens to the task." – antirez
Placing the choice list earlier in the prompt can improve the model’s focus on the classification task.
Structured output alternatives
"Instead of letting a model generate only "A", "B", or "C" and looking at the probs, have it directly generate "Legitimate", "Spam" or "Phishing" … you can also have it assign probabilities in words or numbers." – sigmoid10
Using a JSON or plain‑text schema for the assistant’s response can simplify downstream parsing and reduce reliance on token‑level logits.
Calibration concerns
"The probabilities are not always correct; calibrated decisions usually require RL‑based fine‑tuning (RLCD)." – original post
Many commenters note that raw logits are not guaranteed to be well‑calibrated. Techniques such as temperature scaling, Brier‑loss fine‑tuning, or post‑hoc calibration curves can improve reliability.
Speed vs. accuracy trade‑off
"Why would you not want ‘reasoning’ in a classifier? Speed and cost are obvious reasons, but isn’t this a trade‑off?" – brap
The script sacrifices any chain‑of‑thought reasoning for latency. For high‑stakes decisions, a slower model with explicit reasoning may yield higher accuracy.
Practical considerations
Model choice
The example uses Qwen/Qwen3-0.6B-Q8_0.gguf, a 0.6 B‑parameter model that fits on modest hardware. Faster inference can be achieved with quantized or smaller models, but accuracy may vary across domains.
Latency measurement
The post does not provide benchmark numbers. Community members request concrete latency and error‑rate figures (e.g., “<200 ms for 45 questions”). Measuring end‑to‑end time on the target hardware is essential before adopting the approach in production.
Error handling
Parsing failures can occur if the model emits unexpected tokens. Adding a strict output schema (e.g., JSON with a choice field) and retry logic reduces malformed responses.
Extensibility
The same pattern works for any multi‑class classification task: replace labels, choices, and email with the appropriate domain data, and adjust the prompt to reflect the new context.
Alternatives and open implementations
- OpenJev – a more complete open‑source reference implementation.
- openjev‑sglang – integrates with the
sglangruntime for higher throughput. - OpenJev on DiffusionGemma – demonstrates the approach with a different backbone model.
- Laya – an open‑weight model explicitly tuned for System‑One style decisions (see comment by SylonZero).
Bottom line
The 25‑line script proves that Jev’s core idea—prompt‑based classification with probability extraction—can be reproduced with a generic LLM, no proprietary training, and minimal code. However, practitioners should be aware of calibration limits, prompt design nuances, and the need for robust output parsing before relying on such a lightweight pipeline in production.
Sources
Related
- Dispatch
- Project
- Project
- Dispatch
- Project