Kev: Open-Source Decision Models built on Qwen3.5

Kev is a family of small decision models designed for high-speed, structured classification and rating. Built on Qwen3.5 base models, Kev implements an architecture inspired by Jev, allowing users to run multiple decision-making questions (yes/no, multiple-choice, and rating) against a single piece of input text in a single forward pass.

Architecture and Implementation

Kev utilizes a rank-16 LoRA adapter and a custom pointer head on top of Qwen base models. The system is designed to maximize efficiency by processing the input state once and reusing that computation for multiple independent questions.

Question Isolation and Processing

For attention-only base models (like Qwen3), Kev uses a specific attention mask that allows a token to read the input state and its own question, but prevents it from reading other questions. This ensures that questions remain independent and do not influence one another.

For Qwen3.5 models, which incorporate Gated DeltaNet layers (recurrent layers that ignore attention masks), Kev processes each question as an independent row. The server computes the state once and reuses the cache for every row, maintaining exact isolation between questions.

The Pointer Head

Instead of generating text tokens, Kev uses a pointer head to score the hidden state of each option's closing tag (</opt>) against the hidden state of the question's decision token (<decide>). A softmax function then converts these scores into probabilities, providing a calibrated confidence level for each answer.

Model Family and Performance

Kev is available in three sizes: 0.8B, 4B, and 9B. All three are trained on the same data and settings, allowing users to choose the model based on their memory and accuracy requirements.

Benchmarks and Accuracy

Kev-9B is the most capable model, achieving 0.852 accuracy on the test set. While it trails the hosted Jev model by 3.5 points on new-source development sets (0.822 vs 0.857), it demonstrates strong generalization to unseen policy rule types.

Model Base Accuracy (New Sources - Test) Brier Score (New Sources - Test)
Kev-0.8B Qwen3.5-0.8B 0.684 0.460
Kev-4B Qwen3.5-4B 0.837 0.255
Kev-9B Qwen3.5-9B 0.852 0.237

Serving Performance

Latency varies significantly by hardware. On CUDA (H100), a five-question request takes tens of milliseconds. On Apple Silicon, the lack of fast kernels for DeltaNet layers means the 9B model can take approximately 2 seconds per request, whereas the 4B model takes roughly 779ms. For low-latency requirements on Mac, the previous generation Qwen3-based models are recommended.

API and Integration

Kev's API matches TypeSafe's System One, enabling seamless integration with the TypeSafe Python SDK. The /v1/systemone endpoint accepts a state (the text to evaluate) and a set of questions.

Supported Question Types

  • noul: A binary yes/no question returning the probability of "yes".
  • choice: A multiple-choice question returning the most likely option, probabilities for all options, and a confidence score.
  • score: A rating question returning a mean level index, a legend, and probabilities.

Training and Fine-Tuning

Kev is trained using cross-entropy on the correct answer, with the base weights remaining fixed while the adapter and pointer head are trained together.

Custom Fine-Tuning

Users can fine-tune Kev on their own domain-specific data using a JSONL format. The author recommends using the --init_from flag to load a released checkpoint's adapter and pointer head. This preserves the general decision-making capabilities of the model while adding domain-specific knowledge. In one test, fine-tuning from the base model scored 0.33 on Kev's evaluation set, whereas fine-tuning from a released checkpoint maintained 0.83 accuracy and reached 0.88 on the new domain.

Limitations and Considerations

  • Calibration: Raw probabilities can be over-confident on new sources. Using KEV_TEMPERATURE=2.0 can reduce confident errors (wrong answers with probability $\ge 0.9$) from 8.7% to 4.4% for Kev-9B.
  • Date Arithmetic: The models struggle with raw date subtraction. Setting KEV_DATE_FACTS=1 appends the number of days between absolute dates, improving Kev-9B's accuracy on deadline policy questions from 0.80 to 0.90.
  • Option Order: Changing the order of options within a single question can still affect the answer, despite question isolation.
  • Knowledge Gap: There is a a significant gap in general knowledge (MMLU) compared to Jev, which is attributed to the base model's limitations.

Community Insights

Discussion among developers suggests that Jev-like models are particularly useful for internal routing logic, data labeling, and reducing tool-call overhead in agentic workflows. One user noted that they use a proxy to route prompts to Jev to narrow down tools before sending them to a frontier model, resulting in 60% fewer tool calls.

However, some critics argue that the core value of Jev lies in its training data rather than its architecture, suggesting that open-source derivatives may struggle to match the performance of the original hosted model without similar high-quality datasets.

Sources

Related

  • Project
  • Project
  • Dispatch
  • Project
  • Dispatch