Jeeves: Improving Decision Models with Reasoning and Diffusion Drafting
Overview
Jeeves is a reasoning-capable classifier designed to improve the accuracy of "Jev-like" decision models. By integrating a reasoning chain before the final decision, Jeeves addresses the common trade-off in decision models where calibrated probabilities are provided but often at the cost of low accuracy.
Built on Qwen3.5-9B with LoRA and a pointer head, Jeeves utilizes Supervised Fine-Tuning (SFT) and CISPO to enable the model to "think" before it decides, resulting in superior performance on out-of-domain tasks and JevBench hard benchmarks.
Performance Benchmarks
Jeeves demonstrates significant accuracy gains over previous decision models like Jev and Kev-9B, particularly in complex or unseen scenarios.
Accuracy Comparison
| Benchmark | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (out-of-domain/held-out) | 0.822 | 0.857 | 0.889 |
| JevBench overall (231 public items) | 0.715* | 0.866 | 0.935 |
| JevBench hard (111 public items) | 0.451* | 0.730 | 0.865 |
| PAWS | 0.763 | 0.788 | 0.875 |
| Held-out rule structures | 0.896 | 0.885 | 1.000 |
| Contrastive policies | 0.900 | 0.963 | 1.000 |
Note: Kev-9B results for JevBench are based on Kev-8B (Qwen3) data.
Trade-offs and Limitations
While Jeeves excels in accuracy, it introduces a latency penalty due to the reasoning process.
- Knowledge Gap: Jeeves trails Jev in pure knowledge benchmarks, such as MMLU (0.793 vs 0.900) and MMLU-Pro (0.739 vs 0.840).
- Latency: Full reasoning chains can result in a p90 latency of 17 seconds. Without thinking, the model responds in approximately 0.3 seconds.
- Interpretability: Because no language consistency reward was used during training, the reasoning chains are not highly interpretable.
Technical Architecture
Decision Mechanism
Jeeves uses a specific prompt format to separate state, instructions, and options. The model generates a reasoning chain within <think> tags. After the reasoning block, the model is prompted again with the instructions and options, followed by a <decide> token.
A pointer head then calculates the decision by performing a scaled dot product between a query projection of the hidden state at the <decide> token and a key projection of the hidden state at the corresponding option's </opt> token. The final probabilities are derived via a softmax over these scores, adjusted by a temperature fitted on the development set.
Training Pipeline
- SFT: Two epochs of Supervised Fine-Tuning using LoRA (r=16) on Qwen3.5-9B and the pointer head, utilizing 19,126 questions from public datasets and synthetic policy data.
- CISPO: A reinforcement learning schedule using 9,992 RL questions with 8 rollouts each, capped at 2,560 thinking tokens.
- Calibration: A final temperature fitting on the dev set to ensure calibrated probabilities.
Diffusion Drafter
To mitigate the latency of greedy decoding, Jeeves implements a diffusion drafter inspired by Orthrus. This drafter supports Qwen3.5's Gated DeltaNet layers by allowing mask tokens to cross-attend to post-convolution keys and values.
Speed Improvements:
- Plain greedy decoding (one question): 109 tokens/sec
- Block 4 drafter (one question): 176 tokens/sec (1.6x increase)
- Block 8 drafter (one question): 193 tokens/sec (1.76x increase)
Implementation and Usage
Jeeves is compatible with the Jev API and supports three types of questions in a single request:
- noul: Yes/no questions returning a calibrated probability.
- choice: Multiple-choice questions.
- score: Rating questions based on a provided legend.
Latency Optimization
Users can tune the balance between speed and accuracy using the following options:
think: Toggle reasoning on/off.max_think: Limit the number of reasoning tokens (e.g., 768 tokens reduces median latency from 3.3s to 2.0s).nothink_threshold: Skip reasoning if the initial "no-think" confidence exceeds this value.
Community Insights
Discussion among developers highlights a tension between the high accuracy of reasoning models and the speed requirements of decision-class models.
"What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast."
Other contributors suggest that a hybrid approach—using Jev-like models for simple tasks and reasoning-capable models like Jeeves for complex ones—may be the most viable production path. Some users have reported that on consumer hardware (e.g., M5 Pro), the thinking process is significantly slower, taking over 30 minutes for 100 tweets in a specific irony-detection benchmark.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch