Jevstiller Guarantees 98% Agreement with Jev via Finite‑Sample Bounds
TL;DR – What Jevstiller achieves
Jevstiller learns a lightweight classifier from Jev’s own predictions and, by applying a rigorously calibrated confidence threshold, guarantees that no more than 2 % of all requests receive a label that Jev would not have given. The local model answers in ~15 ms on a CPU, cutting the 300 ms network latency of a direct Jev call.
How the system is built
Model architecture
- Encoder: frozen
bge‑smallsentence encoder (384‑dimensional, ONNX Runtime on CPU). - Head: multinomial logistic regression (single linear layer + softmax) trained on Jev’s full probability distribution, not just the top label.
- Size & latency: a few hundred kilobytes; inference ≈15 ms on a single CPU core.
- Training cadence: retrained every 2,000 new Jev answers, using a few thousand rows, finished in seconds.
Routing components
- k‑NN OOD scorer – detects out‑of‑distribution inputs using the training embeddings.
- Confidence router – applies a threshold plus the OOD cutoff to decide whether the local head may answer.
Both components are described in the project’s DESIGN.md (§7.3‑§7.6).
The contract in one equation
For a given task over a traffic window:
c= coverage (fraction of requests answered locally).e= local disagreement rate (fraction of answered requests whose label differs from Jev’s).- Overall agreement with Jev:
A = 1 – c·e.
Setting a target agreement A* = 98 % yields a budget β = 1 – A* = 2 %. The router must keep c·e ≤ β.
Key point: The contract says nothing about truth accuracy; it only bounds disagreement with Jev.
Why a naïve confidence‑threshold sweep fails
A common approach:
- Hold out calibration data.
- Sweep confidence thresholds.
- Pick the loosest threshold whose empirical disagreement is ≤ β.
On five public classification tasks (20 random splits each), this point‑estimate rule broke the 2 % budget in 6‑12 of the 20 splits per task, sometimes exceeding the budget by up to 1 %.
| Task | Coverage (point) | Mean disagreement | Worst disagreement | Budget broken |
|---|---|---|---|---|
| Banking77 | 79.8 % | 1.90 % | 2.70 % | 9/20 |
| CLINC150 | 84.3 % | 2.09 % | 2.85 % | 12/20 |
| AG News | 90.3 % | 2.06 % | 2.65 % | 11/20 |
| TweetEval sentiment | 28.2 % | 1.99 % | 2.60 % | 8/20 |
| TweetEval offensive | 36.8 % | 1.81 % | 2.70 % | 6/20 |
The failure is statistical: the chosen threshold is the one whose sample disagreement happened to be low due to noise. On fresh traffic the true disagreement can be higher.
Jevstiller’s statistically sound alternative
Four disciplined choices make the 98 % guarantee hold with ≥ 95 % confidence for every production version.
- Loss = contract – For each calibration row (unseen by training) we record a binary loss
1if the local model would answer and disagree with Jev, otherwise0. The average loss is exactlyc·e. Because it is a binomial proportion, we can apply the exact Clopper–Pearson confidence interval. - Fixed‑grid, strict‑first testing – Candidate thresholds are predefined. We evaluate them from the strictest to the loosest, stopping at the first that violates the upper bound of the 95 % interval. This fixed‑sequence testing ("Learn Then Test") controls the family‑wise error without any multiple‑testing correction.
- No calibration leakage – The OOD cutoff is derived solely from the training set; the calibration set is used only for testing the thresholds. This avoids the classic point‑estimate pitfall where the same data both selects and evaluates a threshold.
- Headroom & shadow audit – The chosen threshold is fitted to 85 % of the budget. A second check on a pooled set of calibration rows and live shadow traffic validates the full budget. Continuous audit traffic (2 % of all requests) is always sent to Jev, providing an unbiased live estimate of agreement. If the live bound drops below target, routing falls back to Jev and retraining restarts.
The first implementation missed steps 1 and 2, leading to occasional budget violations; after correction, coverage on the Banking77 replay improved from 70.6 % to 70.7 % with the guarantee intact.
Extending the guarantee to Jev’s own confidence floor
Jev returns a confidence score per answer. Many applications treat low confidence as "unsure" and route such cases to a human reviewer.
- Prior to version 0.4.0, Jevstiller only guaranteed agreement on labels.
- Since 0.4.0, a task can specify a
confidence_floor. During calibration and audit, a local answer is counted as a disagreement also when Jev’s confidence falls below this floor. - The same 2 % budget now covers both label disagreement and Jev‑unsure cases.
- Impact on coverage: on the five benchmark tasks, a floor of 0.6 reduces coverage by 4‑12 % points.
Maintaining the guarantee in production
Two dynamics threaten a static bound:
- Model drift – The local head is periodically retrained on new Jev answers.
- Data drift – Traffic distribution or Jev’s behavior may change.
Jevstiller mitigates both by:
- Keeping a permanent 2 % audit slice that always queries Jev.
- Measuring live agreement on this slice per version; if the confidence interval crosses the target, the audit rate is increased and the system falls back to Jev.
- Demonstrated rapid recovery: when a synthetic Jev change occurred at hour 12, local coverage fell from 90 % to 9 % within four minutes, then rebounded to 90 % in 49 minutes after retraining on the new answers.
Coverage vs. agreement trade‑offs
- The bound is conservative: on the benchmark tasks it sacrifices 4‑8 % coverage compared to the naïve point‑estimate rule.
- Training on Jev’s full probability distribution (instead of just the top label) regains 2‑3 % coverage on many‑class tasks.
- Adjusting the target agreement (e.g., from 98 % to 95 %) roughly doubles coverage on the tweet‑sentiment tasks and lifts intent‑classification coverage from ~70 % to the low 80 % range.
- Across the full 90‑99 % agreement spectrum, Jevstiller’s accuracy against human labels stays within ±1 % of Jev’s accuracy, because when the local model disagrees with Jev it is roughly as often correct as Jev.
- Coverage correlates with Jev’s consistency: on TweetEval tasks Jev’s agreement with human labels is only 64‑74 %, limiting the local model’s confident share.
What Jevstiller does not promise
- Accuracy: the guarantee is about agreement with Jev, not correctness against ground truth.
- Per‑class budgets: the current contract is aggregate; rare classes may dominate the disagreement budget.
- Static traffic assumption: the finite‑sample bound assumes calibration rows are a random sample of the future traffic. Significant distribution shift invalidates the bound, which is why the live audit is essential.
Related work
- Selective classification with risk guarantees – Geifman & El‑Yaniv (2017).
- Fixed‑sequence testing – "Learn Then Test" (2021).
- Cascades that distill large models – OCaTS (EMNLP 2023), Cache & Distil (ACL 2024).
- Finite‑sample agreement contracts – BARGAIN (2025) for batch processing; vCache (ICLR 2026) for semantic caching.
Jevstiller’s novelty lies in combining these ideas into a continuously retrained, audited, production‑ready system with a provable agreement contract.
Getting started
git clone https://github.com/tomerglick57/Jevstiller && cd Jevstiller && pip install .
# Reproduce the Banking77 result (no API key, ~10 min)
bash experiments/reproduce.sh
# Run the full benchmark (all five tasks, ~1‑2 h)
bash experiments/bench.sh --no-record
Run the Docker image as a drop‑in proxy for Jev:
docker run -d -p 8080:8080 -v jevstiller-data:/data ghcr.io/tomerglick57/jevstiller
Configure TYPESAFE_BASE_URL to point to the container. After a few thousand requests the service prints the achieved bound and the live audit interval.
The code is released under the Apache 2.0 license.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch