jkudish/jev-mcp
Fast, cheap, typed judgments from TypeSafe's Jev model, as MCP tools.
Jev MCP – Typed Judgment Tools for AI Agents
What it is – A small Node‑based server that wraps TypeSafe’s Jev language model and exposes a set of “mechanical‑check‑point” (MCP) tools. Each tool takes a concrete request (e.g., a claim to verify, a page to screen, a list of candidates to rank) and returns a typed result: a probability distribution, a confidence score, and a flag indicating whether the answer can be accepted automatically or needs human review. The goal is to give autonomous agents cheap, fast checks that would otherwise require running a large LLM on every piece of data.
How you get it – Install the package from npm and run it as a local server:
npx -y @jkudish/jev-mcp # starts the MCP server
You need Node 20+ and a TypeSafe API key (TYPESAFE_API_KEY). The README even shows how to register the server with a variety of MCP‑compatible clients (Amp, Claude Code, Codex, OpenCode, etc.).
Core tools
| Tool | What it does | Typical output |
|---|---|---|
jev_verify |
Checks each claim against supplied evidence. | Verdict per claim (verified, contradicted, …), full probability distribution, confidence, auto‑accept flag. |
jev_screen |
Detects prompt‑injection, substance, and relevance before an agent reads a page. | Probabilities for injection, substance, relevance; recommendation (pass, review, block, skip). |
jev_find |
Scores a list of candidate texts against a natural‑language query and tells you if any answer exists. | Existence probability, top‑k candidate IDs with probabilities. |
jev_rerank |
Returns a relevance score for every candidate and sorts them. | Ranked list with a relevance probability for each item. |
jev_classify |
Batch‑labels items against a user‑provided catalog of classes. | Class ID, margin, confidence, auto‑accept flag per item. |
jev_decide |
Makes a bounded decision among 2‑6 candidates, checking each against explicit requirements. | Selected candidate, escape‑hatch flags, per‑requirement support/contradiction. |
jev_compare |
Compares two passages (overall and per‑aspect) to say whether they state the same fact, contradict, or are unrelated. | Relation per aspect + overall, with confidence and auto‑review flag. |
jev_extract |
Runs a user‑supplied regex, then lets Jev pick the correct match and return the exact string from the source. | Extracted value, status (auto/review/not_found), diagnostics. |
jev_review |
Scores a proposed code diff against evidence, specs, tests, etc. (described in the intro). | |
jev_gate |
Combines a patch‑review and a claim‑verification step in one call. |
Why it matters – Large LLMs are expensive and slow (seconds to minutes per call). Jev MCP trades a tiny fraction of a cent and ~150‑500 ms latency for typed judgments, making it practical to:
- Fact‑check reports or PR descriptions claim‑by‑claim.
- Block pages that contain hidden prompt‑injection before they reach an agent.
- Find or rerank documents without building embeddings or a vector index.
- Auto‑label support tickets, inbox messages, or log entries.
- Make simple decisions (e.g., choose a deployment strategy) while still surfacing ambiguous cases for a human.
Design choices & limits
- Typed responses: every tool enforces a strict schema (probabilities must sum to 1, confidence in
[0,1], etc.). Invalid or malformed model answers are reported asinvalid_responserather than silently accepted. - Auto‑accept thresholds: each tool has a default confidence cut‑off (
auto_accept≈ 0.8) that can be tuned; below that the result is flagged for review. - Batch limits: up to 250 candidates/classes per call, 64 items for classification, 2 000‑character text truncation, and a 100 000‑character aggregate budget for rerank.
- No embeddings or indexes: the system relies on the Jev model’s semantic scoring, so you don’t need to maintain a separate vector store.
- Early‑stage software: the README warns of “rough edges” and encourages contributions.
Typical workflow
- Start the server (
npx -y @jkudish/jev-mcp). - Configure your MCP client (e.g.,
amp mcp add jev …). - Call a tool from your agent code, passing JSON arguments as shown in the README.
- Inspect the typed result – if
decision/actionisauto, proceed; otherwise pause for human review.
Bottom line – Jev MCP is a practical, open‑source wrapper that lets autonomous agents perform cheap, fast, and well‑structured judgments using TypeSafe’s Jev model. It’s squarely in the AI‑agent tooling space and can be dropped into any workflow that already supports MCP‑style servers.
Related
- Project
- Project
- Project
- Project