JevBench v1.3.0 benchmark ranks Jev 1.13.0 as top decision model and highlights open‑source alternatives
JevBench v1.3.0 puts Jev 1.13.0 on top of a 52‑model leaderboard
Result: The official JevBench Score (geometric mean of Intelligence, Calibration, Speed and Cost) ranks Jev 1.13.0 at 74.4 – the highest of all tested systems. The next best models are SemIf (73.1) and djev (73.0), both open‑source projects that were added only days after the benchmark was released.
How the score is computed
| Axis | Definition | Scoring range |
|---|---|---|
| Intelligence | Chance‑corrected accuracy across easy, standard, judge and hard tiers (hard tier weighted 30 %). | 0 = random, 100 = perfect |
| Calibration | For hard‑tier probability items: (a) Expected Calibration Error and (b) total‑variation distance to the gold distribution. | 0 = uncalibrated, 100 = perfect |
| Speed | Median and 95th‑percentile latency of a single‑request run (0.1 s = 100, each 10× slower loses 20 points). Self‑hosted endpoints are adjusted ×2 + 0.15 s to approximate production load. | 0 = slow, 100 = fast |
| Cost | US $ per 1,000 decisions (not per token). $0.001 = 100, each 10× higher loses 30 points. Prices are either the provider’s published tariff or an estimate based on size‑class reference rates. | 0 = expensive, 100 = free |
The final JevBench Score = (Intelligence × Calibration × Speed × Cost)^{1/4}. Systems with Intelligence < 50 receive an additional penalty ((I/50)^2).
---\n## Top‑10 ranked systems (official weighting 25 % each)
| Rank | System | Score | Intelligence | Calibration | Speed | Cost (USD/1k decisions) |
|---|---|---|---|---|---|---|
| 1 | Jev 1.13.0 (TypeSafe) | 74.4 | 85.7 | 82.7 | 83.3 | $0.040 |
| 2 | SemIf (formerly OpenJev, Qwen3.5‑4B) | 73.1 | 79.0 | 72.6 | 83.7 | ~$0.022 |
| 3 | djev (Maisa, diffusion‑gemma) | 73.0 | 82.7 | 65.4 | 91.4 | $0.026 |
| 4 | Winnow‑12B Q8 | 71.2 | 82.0 | 72.0 | 82.3 | ~$0.037 |
| 5 | reflex 4B | 70.3 | 80.1 | 75.2 | 68.0 | ~$0.022 |
| 6 | jqv (Qwen‑32B zero‑shot) | 68.6 | 79.3 | 79.0 | 74.6 | ~$0.056 |
| 7 | decision‑machine‑1 (milliseconds.ai) | 68.3 | 62.1 | 70.4 | 92.9 | $0.035 |
| 8 | decider‑35b‑a3b (Mapika) | 67.6 | 79.6 | 71.5 | 80.8 | ~$0.067 |
| 9 | open‑alternative‑jev (IkerMoel) | 67.0 | 64.0 | 63.2 | 83.5 | ~$0.022 |
| 10 | system‑one‑open (Gemma 4 E2B LoRA) | 66.6 | 69.5 | 56.7 | 77.0 | ~$0.015 |
All scores are based on the unchanged 534 decision set (220 hard items) measured on 21 Sept 2026 from a server in Germany.
Open‑source Jev‑class models that beat many commercial offerings
| Model | Base model | License | JevBench Score | Notable traits |
|---|---|---|---|---|
| SemIf | Qwen‑3.5‑4B | MIT (code) + Apache‑2.0 (weights) | 73.1 | Runs in the browser; no fine‑tuning beyond the base checkpoint. |
| djev | DiffusionGemma‑26B‑A4B (Apache‑2.0) | Apache‑2.0 | 73.0 | Provides a one‑step inference path; experimental "thinking" mode is slower and more costly. |
| Winnow‑12B Q8 | Gemma‑12B (Quantized) | Apache‑2.0 | 71.2 | Quantized GGUF, full GPU offload; private training corpus not released. |
| reflex 4B | Qwen‑3.5‑4B + LoRA | MIT | 70.3 | Uses per‑primitive calibration files; state is encoded once and each option read from label logits. |
| jqv | Qwen‑3‑32B (stock) | Apache‑2.0 | 68.6 | Zero‑shot; temperature fitted on MMLU validation. |
These open entrants are fully reproducible – the code and weights are publicly available, and the benchmark harness (MIT‑licensed) can be run locally.
Cost analysis – why price matters in the ranking
JevBench reports cost per 1,000 decisions, not per token. A decision averages ~950 input tokens for Jev 1.13.0, so its $0.040 cost corresponds to $0.042 / M input tokens (the public TypeSafe tariff). The cheapest entry is classifier.dev (fast tier) at $0.0033 per 1k decisions, but it is listed as an honorable mention because it simply wraps Jev 1.13.0 rather than being an independent model.
The cost column influences the final score dramatically: a model that is 10× cheaper gains 30 points on the Cost axis. For example, kev 0.6B scores 76.1 on Cost (the cheapest tier) but only 62.5 overall because its Intelligence (51.9) and Calibration (51.1) are low.
How the ranking shifts with alternative weightings
The benchmark provides four preset weightings. Below are the top‑3 positions under each view (only systems that met the ≥95 % decision‑coverage threshold are considered):
| Weighting | #1 | #2 | #3 |
|---|---|---|---|
| Official (25 % each) | Jev 1.13.0 (74.4) | SemIf (73.1) | djev (73.0) |
| Balanced (33 % Intelligence, 33 % Speed, 33 % Cost) | djev (75.8) | Jev 1.13.0 (71.8) | SemIf (73.3) |
| Emphasis on Accuracy | Jev 1.13.0 (77.1) | djev (78.5) | SemIf (75.5) |
| Emphasis on Speed | djev (81.7) | Jev 1.13.0 (76.2) | SemIf (77.3) |
| Emphasis on Cost | kev 0.6B (70.4) | Jev 1.13.0 (63.1) | SemIf (67.4) |
These tables illustrate that Intelligence‑heavy weightings favor Jev 1.13.0, while Cost‑heavy weightings push cheap models like kev 0.6B to the top.
Accuracy by subject area (full‑suite view)
The benchmark breaks down performance by topic. Jev 1.13.0 is strongest on Everyday language (100 % correct) and Safety & security (100 %). Its weakest area is Finance & commerce (73 % correct). SemIf excels on Coding & software (96 % correct) but lags on Rules, policy & law (64 %). djev shows the highest Hard‑tier public accuracy (67 % correct) among the top three.
| Topic | Jev 1.13.0 | SemIf | djev |
|---|---|---|---|
| Math & numbers | 87.6 % | 79.1 % | 67.6 % |
| Coding & software | 83.9 % | 96.4 % | 71.3 % |
| Rules, policy & law | 83.6 % | 64.2 % | 71.3 % |
| Finance & commerce | 73.4 % | 60.9 % | 71.3 % |
| Support & operations | 89.1 % | 87.4 % | 88.3 % |
| Everyday language | 100 % | 100 % | 100 % |
| Safety & security | 100 % | 75 % | 95 % |
Honorable mentions – services that run another entrant’s model
- classifier.dev (fast tier) – runs Jev 1.13.0 under a separate API, scoring 83.6 (higher than any ranked model) but listed as an honorable mention to avoid double‑counting the same underlying model.
- Other services (e.g., Qwen 3.8 27B via Chutes, Needle 3 in tool‑call mode) are shown as partial runs because they did not answer ≥95 % of the decision set.
Why some models were not measured
The benchmark reports a partial‑run or excluded status only when the model could not be evaluated under the same conditions (hardware, licensing, or endpoint availability). Examples:
- open‑jev (MLX on Apple Silicon) – cannot run on the NVIDIA GPUs used for the benchmark.
- mini‑jev – only implements a subset of the task types (Choice/Noul) and lacks a Score head.
- Needle 3 – returns a label plus confidence rather than a full probability distribution, so it is listed only as a partial run.
- Several models required proprietary licences (e.g., Gemma under Google’s terms) that the benchmark team could not accept on behalf of the authors.
Community reaction on Hacker News
- @sean_pedersen questioned duplicate results and model set differences.
- @swyx linked a discussion where Jev’s CEO explained the decision to avoid public benchmarking for privacy reasons.
- @hbrn accused Jev of being a “scam” because its performance matches a quickly‑built Qwen‑based open‑source clone (SemIf) while costing roughly twice as much.
- @adityamishra241 asked how the benchmark guards against over‑fitting to the public 534 questions; the authors note that a held‑out hard tier (109 items) is kept secret during evaluation.
- @tamimio expressed surprise at the rapid emergence of ~80 competing models within a week of the benchmark’s release.
How to submit a new model or custom evaluation
- Open an issue in the JevBench repository (https://github.com/fstandhartinger/jevbench) with a reproducible endpoint URL or a runnable Docker image.
- Provide the exact model version, licence, and whether any public JevBench items were used during training.
- The benchmark team will run the model on the full 534‑decision suite and publish the results in the next version.
- For private data, the benchmark offers a custom evaluation service (https://benchmarkheaven.com/jev-models/custom-evaluation) that runs your model on the same harness behind your firewall.
Takeaways
- Jev 1.13.0 remains the best‑overall decision model when Intelligence, Calibration, Speed and Cost are weighted equally.
- Open‑source rebuilds (SemIf, djev, Winnow‑12B Q8) are highly competitive, often beating commercial APIs on speed or cost while staying within a few points of the top score.
- Cost matters: the cheapest models achieve respectable scores, and under a cost‑heavy weighting they can outrank Jev.
- Transparency: the benchmark publishes the full harness, task definitions, and per‑model cost calculations, enabling reproducibility and fair comparison.
- Community scrutiny is intense – several commenters challenge the value of the benchmark, the pricing methodology, and the potential for over‑fitting, underscoring the need for continued independent evaluation.
Quick reference table (top‑5)
| Rank | System | Score | Cost (USD/1k decisions) | Speed (p95) |
|---|---|---|---|---|
| 1 | Jev 1.13.0 | 74.4 | $0.040 | 0.72 s |
| 2 | SemIf | 73.1 | ~$0.022 | 0.78 s |
| 3 | djev | 73.0 | $0.026 | 0.31 s |
| 4 | Winnow‑12B Q8 | 71.2 | ~$0.037 | 0.98 s |
| 5 | reflex 4B | 70.3 | ~$0.022 | 3.75 s |
All numbers are taken directly from the JevBench v1.3.0 release (21 Sept 2026) and have not been altered.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch