firelex/jeff
Millisecond decisions, any domain: a 0.8B open "System 1" model that picks between your options with calibrated probabilities. One base, swappable LoRA adapters, on your own hardware.
What it solves
Jeff provides a way to integrate extremely fast, zero-shot classification into local code. It replaces the need for large, slow LLMs or complex parsing of generated text by returning calibrated probabilities for a set of options described in plain words. This allows developers to slot a small, high-performance decision model into their applications for tasks like intent classification, moderation labels, or voice commands without the need for cloud APIs.
How it works
Jeff consists of a series of small, fine-tuned versions of Qwen3.5 and Gemma 4 (ranging from 0.8B to 2B parameters). Unlike traditional LLMs that generate text, Jeff performs a single forward pass to return a probability for each provided option. It uses a trained answer readout and a fitted temperature for calibration. The models are trained on synthetic data generated by a local teacher model (Qwen3.8-Flash-Next) and include a leak filter to ensure benchmark integrity.
Who it’s for
Developers who need a low-latency, local decision-making agent that can handle arbitrary categories (zero-shot) and can be fine-tuned for specific domain-specific tasks to further increase accuracy.
Highlights
- Extreme Speed: Decisions are made in as little as 22ms on an RTX PRO 6000 and 28ms on an Apple M4 Max.
- Zero-Shot Capability: Can classify into categories that were not present in the training data by describing them in plain words.
- Local-First: Built and trained entirely on local hardware without relying on closed-model outputs in the training data.
- Fine-Tunable: Supports rapid fine-tuning on custom examples to significantly boost accuracy for specific use cases (e.g., moving from 31.7% to 95.8% accuracy in voice navigation).
- Multi-Question Support: Can answer several independent questions in a single request.
- Apple Silicon Optimized: Includes MLX serving for faster performance on Mac hardware.
Related
- Project
- Project
- Project
- Project