Jeff: Jev-compatible 0.8B and 2B Decision Models for Zero-Shot Classification
Jeff is a collection of small, open-weight language models fine-tuned for zero-shot classification. By utilizing a single forward pass to return calibrated probabilities for a set of options without generating text, Jeff provides a high-speed alternative to traditional LLM-based classification. The 0.8B parameter model achieves decision latencies of approximately 22ms on an NVIDIA RTX PRO 6000 and 28ms on an Apple M4 Max (via MLX).
High-Speed Decision Making
Jeff's primary value proposition is extreme speed and cost-effectiveness. Unlike generative LLMs that iterate over tokens, Jeff functions as a classifier that slots directly into application code.
Performance Metrics
| Model | Parameters | NVIDIA RTX PRO 6000 | Apple M4 Max (MLX) | CPU (32 threads) |
|---|---|---|---|---|
| Jeff-Qwen3.5-0.8B | 0.8B | 22 ms | 28 ms | 463 ms |
| Jeff-Qwen3.5-2B | 2B | 24 ms | 60 ms | 708 ms |
| Jeff-Gemma4-E2B | 2B effective | 29 ms | N/A | 1.0 s |
These speeds allow for real-time decision loops, such as voice navigation or game AI, where traditional API-based models (like Jev) may introduce too much latency (e.g., 114–212 ms per call).
Zero-Shot Classification and Benchmarks
Jeff is designed for zero-shot tasks where the user describes a situation and a list of options in plain words. The model then picks the most likely option based on the description.
Benchmark Results
Across five public benchmarks, the Jeff-Qwen3.5-2B model achieved an overall accuracy of 83.1%, matching the published figures for Jev (83.0%). While Jeff excels in classification and grounding tasks—outperforming Jev on the Financial PhraseBank (96.3% vs 77.0%) and RAGTruth (88.9% vs 77.3%)—it lags behind in reasoning-heavy benchmarks:
- BBH: 64-68% (Jeff) vs 94.3% (Jev)
- JudgeBench: 60-64% (Jeff) vs 78.6% (Jev)
- JevBench Hard: 47-53% (Jeff) vs 73.3% (Jev)
This performance gap is expected given the small size of the models (0.8B to 2B parameters), as they are intended as "System 1" judgement-callers rather than multi-step reasoners.
Local Training and Implementation
Jeff was built entirely on local hardware, utilizing one RTX PRO 6000 workstation GPU. The training process involved full-weight fine-tuning for one epoch with batches of 256 and cross-entropy over option letters.
- Synthetic Data: Training data was generated by Qwen3.8-Flash-Next on two DGX Sparks, with a closed model used only for spot-checking quality.
- Training Time: The 0.8B model trains in approximately 2 hours, and the 2B model in 3.5 hours.
- Fine-tuning: For domain-specific tasks, a short fine-tune can significantly boost accuracy. In one example, a voice-navigation task moved from 31.7% to 95.8% accuracy in under 30 minutes on a single GPU.
Practical Application: Game Testing
To test zero-shot capabilities on non-benchmark data, Jeff was used to play three games: Doom, Frogger, and Pac-Man.
- Doom: Jeff-0.8B achieved 6.55 kills per episode, matching the performance of a hand-coded rule bot and Jev's published runs.
- Frogger: Jeff-0.8B achieved 10.3 crossings, significantly outperforming the untrained base model (1.0).
- Pac-Man: Jeff-0.8B collected 57 pellets, roughly 60% of the rule bot's performance.
Usage Guidelines and Caveats
To maximize Jeff's effectiveness, developers should follow these guidelines:
- Reason in Code, Decide with Jeff: Use Jeff as a classifier, not a planner. Provide the consequences of a move (e.g., "this move gets you hit by a car") rather than asking it to forecast future states.
- Wording Matters: Consistent and descriptive wording for options is critical. Small changes in how a goal is described can lead to significant improvements in performance.
- Option Keys: Use short keys (e.g.,
{"1": "Refund request"}) to minimize latency and token overhead.
Community Feedback
While the project has been praised for its local-first approach and speed, some users have noted that zero-shot accuracy can vary wildly depending on the use case. One user reported a significant gap in accuracy (70% vs 94%) when compared to Jev in their specific application, highlighting the importance of fine-tuning for specialized tasks.
Project History and Licensing
Jeff is a fork of AutoJev, an open-source recipe for fine-tuning models to return Jev-style decisions. The weights are released under the Apache 2.0 license, and the code is under the MIT license.
Sources
Related
- Dispatch
- Project
- Dispatch
- Dispatch
- Dispatch