jaredpalmer/kev

Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own

What it solves

kev is a decision model designed to replace traditional text-generation for classification and scoring tasks. Instead of generating prose (which can be slow and inconsistent), it provides calibrated probabilities for specific question types, allowing for faster, more reliable, and structured decision-making from a single document.

How it works

The project uses a Qwen base model (ranging from 0.5B to 8B parameters) with a LoRA adapter and a custom pointer readout head. It processes a document and multiple questions in a single "prefill" pass without any decoding.

Key technical mechanisms include:

  • Block-Causal Masking: This ensures that while every question can see the document, questions cannot see each other, maintaining strict isolation.
  • Packing: The state (document) and all questions are packed into one sequence, allowing the model to answer many questions in parallel.
  • Pointer Readout: A specialized head scores option tokens against a decision token and applies a softmax to produce probabilities rather than text.
  • Training: The model is trained using cross-entropy on labeled outcomes, ensuring the output probabilities are learned and calibrated.

Who it’s for

It is intended for developers and researchers who need a high-performance, structured decision-making API (compatible with TypeSafe's System One contract) that can run locally on laptops or be trained in the cloud.

Highlights

  • Three Question Types: Supports noul (yes/no), choice (multiple options), and score (ordered levels).
  • High Efficiency: Answers multiple questions in one forward pass with no decoding, significantly reducing latency.
  • Strict Isolation: Guarantees that the answer to one question does not influence another.
  • Drop-in API: Compatible with the typesafe-sdk, allowing for easy integration into existing workflows.
  • Hardware Flexible: Small models (0.5B) can train on a Mac, while larger ones (4B/8B) can be served on a 32GB Mac in bf16.

Written about in

Related

  • Project
  • Project
  • Project
  • Project