jkudish/jev-mcp

Fast, cheap, typed judgments from TypeSafe's Jev model, as MCP tools.

Jev MCP – Typed Judgment Tools for AI Agents

What it is – A small Node‑based server that wraps TypeSafe’s Jev language model and exposes a set of “mechanical‑check‑point” (MCP) tools. Each tool takes a concrete request (e.g., a claim to verify, a page to screen, a list of candidates to rank) and returns a typed result: a probability distribution, a confidence score, and a flag indicating whether the answer can be accepted automatically or needs human review. The goal is to give autonomous agents cheap, fast checks that would otherwise require running a large LLM on every piece of data.

How you get it – Install the package from npm and run it as a local server:

npx -y @jkudish/jev-mcp   # starts the MCP server

You need Node 20+ and a TypeSafe API key (TYPESAFE_API_KEY). The README even shows how to register the server with a variety of MCP‑compatible clients (Amp, Claude Code, Codex, OpenCode, etc.).

Core tools

Tool What it does Typical output
jev_verify Checks each claim against supplied evidence. Verdict per claim (verified, contradicted, …), full probability distribution, confidence, auto‑accept flag.
jev_screen Detects prompt‑injection, substance, and relevance before an agent reads a page. Probabilities for injection, substance, relevance; recommendation (pass, review, block, skip).
jev_find Scores a list of candidate texts against a natural‑language query and tells you if any answer exists. Existence probability, top‑k candidate IDs with probabilities.
jev_rerank Returns a relevance score for every candidate and sorts them. Ranked list with a relevance probability for each item.
jev_classify Batch‑labels items against a user‑provided catalog of classes. Class ID, margin, confidence, auto‑accept flag per item.
jev_decide Makes a bounded decision among 2‑6 candidates, checking each against explicit requirements. Selected candidate, escape‑hatch flags, per‑requirement support/contradiction.
jev_compare Compares two passages (overall and per‑aspect) to say whether they state the same fact, contradict, or are unrelated. Relation per aspect + overall, with confidence and auto‑review flag.
jev_extract Runs a user‑supplied regex, then lets Jev pick the correct match and return the exact string from the source. Extracted value, status (auto/review/not_found), diagnostics.
jev_review Scores a proposed code diff against evidence, specs, tests, etc. (described in the intro).
jev_gate Combines a patch‑review and a claim‑verification step in one call.

Why it matters – Large LLMs are expensive and slow (seconds to minutes per call). Jev MCP trades a tiny fraction of a cent and ~150‑500 ms latency for typed judgments, making it practical to:

  • Fact‑check reports or PR descriptions claim‑by‑claim.
  • Block pages that contain hidden prompt‑injection before they reach an agent.
  • Find or rerank documents without building embeddings or a vector index.
  • Auto‑label support tickets, inbox messages, or log entries.
  • Make simple decisions (e.g., choose a deployment strategy) while still surfacing ambiguous cases for a human.

Design choices & limits

  • Typed responses: every tool enforces a strict schema (probabilities must sum to 1, confidence in [0,1], etc.). Invalid or malformed model answers are reported as invalid_response rather than silently accepted.
  • Auto‑accept thresholds: each tool has a default confidence cut‑off (auto_accept ≈ 0.8) that can be tuned; below that the result is flagged for review.
  • Batch limits: up to 250 candidates/classes per call, 64 items for classification, 2 000‑character text truncation, and a 100 000‑character aggregate budget for rerank.
  • No embeddings or indexes: the system relies on the Jev model’s semantic scoring, so you don’t need to maintain a separate vector store.
  • Early‑stage software: the README warns of “rough edges” and encourages contributions.

Typical workflow

  1. Start the server (npx -y @jkudish/jev-mcp).
  2. Configure your MCP client (e.g., amp mcp add jev …).
  3. Call a tool from your agent code, passing JSON arguments as shown in the README.
  4. Inspect the typed result – if decision/action is auto, proceed; otherwise pause for human review.

Bottom line – Jev MCP is a practical, open‑source wrapper that lets autonomous agents perform cheap, fast, and well‑structured judgments using TypeSafe’s Jev model. It’s squarely in the AI‑agent tooling space and can be dropped into any workflow that already supports MCP‑style servers.

Related

  • Project
  • Project
  • Project
  • Project