Jeeves: Improving Decision Models with Reasoning and Diffusion Drafting

Overview

Jeeves is a reasoning-capable classifier designed to improve the accuracy of "Jev-like" decision models. By integrating a reasoning chain before the final decision, Jeeves addresses the common trade-off in decision models where calibrated probabilities are provided but often at the cost of low accuracy.

Built on Qwen3.5-9B with LoRA and a pointer head, Jeeves utilizes Supervised Fine-Tuning (SFT) and CISPO to enable the model to "think" before it decides, resulting in superior performance on out-of-domain tasks and JevBench hard benchmarks.

Performance Benchmarks

Jeeves demonstrates significant accuracy gains over previous decision models like Jev and Kev-9B, particularly in complex or unseen scenarios.

Accuracy Comparison

Benchmark Kev-9B Jev Jeeves
Test overall (out-of-domain/held-out) 0.822 0.857 0.889
JevBench overall (231 public items) 0.715* 0.866 0.935
JevBench hard (111 public items) 0.451* 0.730 0.865
PAWS 0.763 0.788 0.875
Held-out rule structures 0.896 0.885 1.000
Contrastive policies 0.900 0.963 1.000

Note: Kev-9B results for JevBench are based on Kev-8B (Qwen3) data.

Trade-offs and Limitations

While Jeeves excels in accuracy, it introduces a latency penalty due to the reasoning process.

  • Knowledge Gap: Jeeves trails Jev in pure knowledge benchmarks, such as MMLU (0.793 vs 0.900) and MMLU-Pro (0.739 vs 0.840).
  • Latency: Full reasoning chains can result in a p90 latency of 17 seconds. Without thinking, the model responds in approximately 0.3 seconds.
  • Interpretability: Because no language consistency reward was used during training, the reasoning chains are not highly interpretable.

Technical Architecture

Decision Mechanism

Jeeves uses a specific prompt format to separate state, instructions, and options. The model generates a reasoning chain within <think> tags. After the reasoning block, the model is prompted again with the instructions and options, followed by a <decide> token.

A pointer head then calculates the decision by performing a scaled dot product between a query projection of the hidden state at the <decide> token and a key projection of the hidden state at the corresponding option's </opt> token. The final probabilities are derived via a softmax over these scores, adjusted by a temperature fitted on the development set.

Training Pipeline

  1. SFT: Two epochs of Supervised Fine-Tuning using LoRA (r=16) on Qwen3.5-9B and the pointer head, utilizing 19,126 questions from public datasets and synthetic policy data.
  2. CISPO: A reinforcement learning schedule using 9,992 RL questions with 8 rollouts each, capped at 2,560 thinking tokens.
  3. Calibration: A final temperature fitting on the dev set to ensure calibrated probabilities.

Diffusion Drafter

To mitigate the latency of greedy decoding, Jeeves implements a diffusion drafter inspired by Orthrus. This drafter supports Qwen3.5's Gated DeltaNet layers by allowing mask tokens to cross-attend to post-convolution keys and values.

Speed Improvements:

  • Plain greedy decoding (one question): 109 tokens/sec
  • Block 4 drafter (one question): 176 tokens/sec (1.6x increase)
  • Block 8 drafter (one question): 193 tokens/sec (1.76x increase)

Implementation and Usage

Jeeves is compatible with the Jev API and supports three types of questions in a single request:

  • noul: Yes/no questions returning a calibrated probability.
  • choice: Multiple-choice questions.
  • score: Rating questions based on a provided legend.

Latency Optimization

Users can tune the balance between speed and accuracy using the following options:

  • think: Toggle reasoning on/off.
  • max_think: Limit the number of reasoning tokens (e.g., 768 tokens reduces median latency from 3.3s to 2.0s).
  • nothink_threshold: Skip reasoning if the initial "no-think" confidence exceeds this value.

Community Insights

Discussion among developers highlights a tension between the high accuracy of reasoning models and the speed requirements of decision-class models.

"What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast."

Other contributors suggest that a hybrid approach—using Jev-like models for simple tasks and reasoning-capable models like Jeeves for complex ones—may be the most viable production path. Some users have reported that on consumer hardware (e.g., M5 Pro), the thinking process is significantly slower, taking over 30 minutes for 100 tweets in a specific irony-detection benchmark.

Sources

Related