vLLM Introduces Distortion-Free Gumbel-Max Watermarking

TL;DR

vLLM now supports distortion‑free watermarking using a Gumbel‑max algorithm that embeds a secret key‑derived signal into token selection, enabling reliable provenance detection with virtually no impact on throughput or output quality.

Why Text Provenance Matters

Establishing the origin of generated text is essential for trust and accountability. Traditional watermarking methods for images or audio cannot be directly applied to text because text is discrete. Instead, vLLM influences the generation process itself, creating a detectable pattern without modifying the expected output distribution.

Core Requirements for Practical Text Watermarking

  • Non‑distortion: The watermark must not bias the model toward particular words, styles, or solutions.
  • Robustness: The signal should survive common edits, truncations, and re‑phrasing by other LLMs.
  • Speed: Minimal added latency and memory overhead are required for high‑throughput serving.
  • Minimal detection dependencies: Detection should rely only on the text, tokenizer, and secret key, without extra metadata.

These criteria are often at odds: a stronger, easily detectable signal can increase distortion, while avoiding extra information limits algorithmic options.

Exploiting Randomness via the Gumbel‑Max Trick

At each step a language model samples a token from a categorical distribution. The Gumbel‑max trick adds Gumbel‑distributed noise to the logits and selects the argmax, reproducing the exact original distribution while allowing the noise to be generated deterministically.

Keyed Pseudorandom Noise

vLLM replaces the independent uniform draws (u) with a pseudorandom function (PRF) that takes three inputs:

  • a secret watermark key (k),
  • the recent watermarking context (default: last 4 tokens), and
  • the candidate token ID.

The PRF outputs a uniform value that is transformed into Gumbel noise. Because the same context and key appear in the final text, a detector can reconstruct the exact noise values used during generation.

Non‑Distortion Guarantee

In expectation over keys, the probability of selecting any token (t) remains exactly its original probability (p_t). Empirically, quality benchmarks on Qwen3.5‑27B show negligible differences (e.g., GSM8K 93.0 % vs. 94.2 %).

Detection Procedure

Detection reverses the generation process:

  1. Tokenize the unknown text.
  2. For each token, recompute the keyed PRF value using the preceding context and secret key.
  3. Convert each value into a token‑level score; larger scores indicate alignment with the watermark.
  4. Sum the scores. Under the null hypothesis (no watermark) the sum follows a Gamma distribution (\Gamma(k,\theta)) with shape (k) equal to the number of scored tokens and scale (\theta = 1).
  5. Compute a one‑sided p‑value. A low p‑value (e.g., <0.01) signals the presence of the watermark.

Multiple‑key or multi‑tokenizer testing requires a multiple‑testing correction, which raises the detection threshold and reduces power.

Empirical Detection Power

  • Creative writing reaches near‑100 % true‑positive rate (TPR) after ~100 tokens.
  • Code‑generation benchmarks (MBPP) achieve ~69 % TPR at 400 tokens for a single‑key test, decreasing with more candidate keys.

Integration into vLLM’s Sampling Pipeline

The watermarking logic is embedded in Model Runner v2’s GPU sampler:

  • GPUWatermarkSampler delegates to a Watermarker implementation.
  • A fused GPU kernel combines PRF generation, Gumbel transformation, and argmax reduction, avoiding a full ([batch, vocab]) noise tensor.
  • Philox generates four PRF values per call, processing four consecutive token IDs together.
  • A per‑row mask enables mixed batches of watermarked and unwatermarked requests.

Throughput Impact

Across batch sizes 1–256 on a single H100 (Qwen3.5‑27B, MTP‑3), watermarked and unwatermarked throughput curves overlap. Mean matched throughput changes range from –1.1 % to +2.0 % with no statistically significant slowdown.

Handling Speculative Decoding

Speculative decoding proposes draft tokens from a fast model and accepts them against a target distribution. Applying the same watermark to both distributions harms acceptance rates. vLLM solves this with a dual‑key scheme:

  • Key (k_d) for draft tokens.
  • Key (k_t) for target residual and bonus tokens. The final text may contain tokens watermarked by either key, so detection combines scores from both keys. This mitigates acceptance‑rate loss but dilutes the overall detection signal.

Preserving Output Diversity

A fixed key can cause repeated contexts to generate identical PRF values, leading to repetitive loops (e.g., "1 + 1 + 1 + …"). To prevent this, vLLM implements generation‑time context deduplication:

  • When a context repeats, watermarking is skipped for that step.
  • This restores single‑sequence non‑distortion and adds less than 0.19 % overhead on Qwen3.5‑27B. Dual‑key routing also introduces randomness that reduces repetition, and random key selection can be used even without speculative decoding.

Getting Started

Enable Gumbel‑max watermarking when launching a vLLM server:

vllm serve mistralai/Mistral-7B-Instruct-v0.3 \
  --watermark-config '{"algorithm":"gumbel","key":42}'

A minimal HTTP detection server is provided in the repository. See the official documentation for full configuration details.

Conclusion

vLLM’s new watermarking feature embeds a provenance signal directly into token sampling using a keyed Gumbel‑max process. It guarantees non‑distortion in expectation, adds negligible latency, supports speculative decoding via a dual‑key design, and includes safeguards (context deduplication) to maintain output diversity.

References

  1. Ingemar J. Cox et al., Digital Watermarking and Steganography, 2nd ed., Morgan Kaufmann, 2008.
  2. S. Dathathri et al., “Scalable watermarking for identifying large language model outputs,” Nature 634, 2024.
  3. J. Kirchenbauer et al., “A Watermark for Large Language Models,” Proceedings of the 40th International Conference on Machine Learning, 2023.
  4. S. Aaronson & H. Kirchner, “Watermarking GPT outputs,” 2023.
  5. T. Sander et al., “TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection,” arXiv:2605.12456, 2026.

Sources