Jeff: Jev-compatible 0.8B and 2B Decision Models for Zero-Shot Classification

Jeff is a collection of small, open-weight language models fine-tuned for zero-shot classification. By utilizing a single forward pass to return calibrated probabilities for a set of options without generating text, Jeff provides a high-speed alternative to traditional LLM-based classification. The 0.8B parameter model achieves decision latencies of approximately 22ms on an NVIDIA RTX PRO 6000 and 28ms on an Apple M4 Max (via MLX).

High-Speed Decision Making

Jeff's primary value proposition is extreme speed and cost-effectiveness. Unlike generative LLMs that iterate over tokens, Jeff functions as a classifier that slots directly into application code.

Performance Metrics

Model Parameters NVIDIA RTX PRO 6000 Apple M4 Max (MLX) CPU (32 threads)
Jeff-Qwen3.5-0.8B 0.8B 22 ms 28 ms 463 ms
Jeff-Qwen3.5-2B 2B 24 ms 60 ms 708 ms
Jeff-Gemma4-E2B 2B effective 29 ms N/A 1.0 s

These speeds allow for real-time decision loops, such as voice navigation or game AI, where traditional API-based models (like Jev) may introduce too much latency (e.g., 114–212 ms per call).

Zero-Shot Classification and Benchmarks

Jeff is designed for zero-shot tasks where the user describes a situation and a list of options in plain words. The model then picks the most likely option based on the description.

Benchmark Results

Across five public benchmarks, the Jeff-Qwen3.5-2B model achieved an overall accuracy of 83.1%, matching the published figures for Jev (83.0%). While Jeff excels in classification and grounding tasks—outperforming Jev on the Financial PhraseBank (96.3% vs 77.0%) and RAGTruth (88.9% vs 77.3%)—it lags behind in reasoning-heavy benchmarks:

  • BBH: 64-68% (Jeff) vs 94.3% (Jev)
  • JudgeBench: 60-64% (Jeff) vs 78.6% (Jev)
  • JevBench Hard: 47-53% (Jeff) vs 73.3% (Jev)

This performance gap is expected given the small size of the models (0.8B to 2B parameters), as they are intended as "System 1" judgement-callers rather than multi-step reasoners.

Local Training and Implementation

Jeff was built entirely on local hardware, utilizing one RTX PRO 6000 workstation GPU. The training process involved full-weight fine-tuning for one epoch with batches of 256 and cross-entropy over option letters.

  • Synthetic Data: Training data was generated by Qwen3.8-Flash-Next on two DGX Sparks, with a closed model used only for spot-checking quality.
  • Training Time: The 0.8B model trains in approximately 2 hours, and the 2B model in 3.5 hours.
  • Fine-tuning: For domain-specific tasks, a short fine-tune can significantly boost accuracy. In one example, a voice-navigation task moved from 31.7% to 95.8% accuracy in under 30 minutes on a single GPU.

Practical Application: Game Testing

To test zero-shot capabilities on non-benchmark data, Jeff was used to play three games: Doom, Frogger, and Pac-Man.

  • Doom: Jeff-0.8B achieved 6.55 kills per episode, matching the performance of a hand-coded rule bot and Jev's published runs.
  • Frogger: Jeff-0.8B achieved 10.3 crossings, significantly outperforming the untrained base model (1.0).
  • Pac-Man: Jeff-0.8B collected 57 pellets, roughly 60% of the rule bot's performance.

Usage Guidelines and Caveats

To maximize Jeff's effectiveness, developers should follow these guidelines:

  • Reason in Code, Decide with Jeff: Use Jeff as a classifier, not a planner. Provide the consequences of a move (e.g., "this move gets you hit by a car") rather than asking it to forecast future states.
  • Wording Matters: Consistent and descriptive wording for options is critical. Small changes in how a goal is described can lead to significant improvements in performance.
  • Option Keys: Use short keys (e.g., {"1": "Refund request"}) to minimize latency and token overhead.

Community Feedback

While the project has been praised for its local-first approach and speed, some users have noted that zero-shot accuracy can vary wildly depending on the use case. One user reported a significant gap in accuracy (70% vs 94%) when compared to Jev in their specific application, highlighting the importance of fine-tuning for specialized tasks.

Project History and Licensing

Jeff is a fork of AutoJev, an open-source recipe for fine-tuning models to return Jev-style decisions. The weights are released under the Apache 2.0 license, and the code is under the MIT license.

Sources

Related