ProvenanceGuard source-aware verification for MCP agents
TL;DR
Hugging Face released ProvenanceGuard, a verification layer for Model Context Protocol (MCP) agents that ensures each factual claim is backed by the exact source the answer attributes, addressing the cross‑source conflation failure mode.
The problem: pooled evidence hides source mis‑attribution
Traditional factuality checkers (RAGAS Faithfulness, MiniCheck, AlignScore, SummaC) evaluate whether a claim is supported by any evidence after pooling all retrieved passages. They do not verify that the claim’s stated source matches the evidence that actually supports it. This leads to cross‑source conflation: a true claim is marked as correct even though it is attributed to the wrong source (e.g., a refund policy cited as coming from an account record). In data‑sensitive domains such as customer support or clinical assistance, incorrect provenance can be as harmful as a factual error.
"A claim can be supported by one MCP source while the answer attributes it to another. Source‑blind scoring sees support in the pooled evidence and passes it; ProvenanceGuard separately checks whether the supporting source matches the one the answer states or implies." – paper Figure 1
What ProvenanceGuard does
ProvenanceGuard operates after an MCP agent generates an answer. It never collapses tool outputs into an anonymous context. Instead, it:
- Decomposes the answer into individual claims.
- Routes each claim to the most relevant MCP source using a similarity model (MiniLM in the paper).
- Verifies support with a natural‑language‑inference (NLI) model (DeBERTa‑v3‑base‑mnli‑fever‑anli).
- Checks attribution by comparing the source used for verification with the source explicitly or implicitly mentioned in the answer.
- Emits a per‑claim source verdict and a global allow/block decision.
If a claim is blocked, a RARR‑style repair loop attempts a source‑grounded rewrite or falls back to a safe response, after which the verifier re‑evaluates the revised answer.
"The verification flow. Source identity is preserved through decomposition, routing, support scoring, attribution checking, and repair, rather than being pooled." – paper Figure 2
The architecture is model‑agnostic: the paper used local models for reproducibility, but any hosted embeddings, NLI, or claim‑splitting service can replace them after appropriate calibration.
Empirical results
The authors evaluated ProvenanceGuard on a medical MCP agent that accessed patient records, research articles, and other tools, yielding 281 real traces and 361 human‑annotated claims from 40 answers.
- Blocking performance: 138 of 139 expert‑identified bad claims were blocked (F1 = 0.802).
- Source identification: Correct source selected for 86 % of claims with identifiable provenance.
- Comparison to baselines: ProvenanceGuard outperformed MiniCheck, RAGAS Faithfulness, AlignScore, and SummaC‑ZS on the paper’s reject/block F1 metric while also providing claim‑to‑source IDs.
| Verifier | Reject/block F1 | Emits claim‑to‑source ID |
|---|---|---|
| ProvenanceGuard | 0.802 | Yes |
| MiniCheck | 0.783 | No |
| RAGAS Faithfulness | 0.758 | No |
| AlignScore | 0.662 | No |
| SummaC‑ZS | 0.436 | No |
Harder source‑disambiguation tests
- On a set with many similar sources, ProvenanceGuard achieved 0.846 F1 for blocking but identified the exact source correctly only 50.3 % of the time, highlighting a remaining challenge.
- In a controlled swap experiment (50 cases where the named source was deliberately altered), ProvenanceGuard caught all 50 mis‑attributions.
Repairing blocked answers
The RARR‑style repair loop resolved every blocked answer in the full‑trace run. Most (144/173) ended in a safe fallback rather than a substantive rewrite, reflecting a conservative policy that prefers “no answer” over an unverifiable one. In a multi‑source stress test, 59 initially blocked answers were repaired with only two fallbacks.
- Latency: Approximately 0.5 s per answer on the local configuration; NLI and routing calls each take tens of milliseconds.
Fit with Multiverse Computing and broader adoption
As agents transition from single‑passage retrieval‑augmented generation (RAG) to multi‑tool MCP workflows, provenance becomes integral to factuality. ProvenanceGuard provides a plug‑in‑style verification stage that respects the agent’s trace without retraining the model.
- Adoption example: NVIDIA’s NVFlow finance agent incorporated an optional grounding‑verification stage based on ProvenanceGuard, checking answers against SEC excerpts while preserving the original rollout.
- Community exposure: Presented as a poster at the Agentic AI Summit 2026 (UC Berkeley).
How to get ProvenanceGuard
The full technical details—including routing heuristics, NLI derivations, calibration ablations, and complete result tables—are available in the paper on Hugging Face and arXiv (arXiv:2606.18037). Teams can implement the pipeline with the same local models or substitute cloud services, provided they perform their own calibration and testing.
For a deeper dive, read the paper on Hugging Face or contact the Multiverse Computing team for integration support.