vLLM PD Serving of Qwen3.8-2.4T achieves 5K TPS/GPU and 180 gen tok/s/user
TL;DR
vLLM’s PD serving of the Qwen3.8-2.4T model hits 5,000 total token throughput per GPU in high‑throughput mode and 180 generated tokens per user in low‑latency mode, establishing a complete pareto frontier for the model on a GB300 NVL72 cluster.
Overview of the Announcement
The blog post details how vLLM achieved the reported performance on a GB300 NVL72 cluster using an 8K/1K workload. It provides reproducible srt‑slurm recipes, a step‑by‑step tuning methodology, and a full pareto frontier that balances total token throughput (TPS) against per‑user interactivity (gen TPS per user). The methodology is presented as more valuable than the raw numbers because it can be applied to any model.
Maximizing Throughput: Concurrency and KV‑Cache Limits
Conclusion
Total token throughput per GPU is limited primarily by KV‑cache capacity, which in turn is governed by the size of the per‑request GDN state.
Technical Details
- Model architecture: Qwen3.8-2.4T has 92 layers (69 GDN, 23 Full‑Attn) with a MoE block of 512 experts per layer.
- Weight footprint: 91 GiB non‑expert weights, 1,242 GiB expert weights.
- KV‑cache composition:
- Full‑Attn state: 2 KiB per token per layer.
- GDN state (per request): 4 MiB (SSM) + 120 KiB (conv) ≈ 4.216 MiB.
- Block sizing: GDN state dominates, leading to a block that holds 2,112 tokens (4.125 MiB per block).
- GPU memory baseline: GB300 provides 279 GB; after driver, CUDA context, and an 8 % safety reserve, ~254 GiB remains for weights, activations, CUDA graphs, and KV‑cache.
- Peak activation memory: Measured per topology (e.g., TP8 with 20 sequences uses 0.57 GiB, TP4DP4 with 272 sequences uses up to 2.24 GiB).
- CUDA‑graph reservation: Estimated conservatively; actual usage often lower, but the reservation prevents OOM.
- KV‑cache availability: After accounting for non‑weight memory (~20 GiB), the remaining memory determines how many requests can be concurrent. TP4DP4 topology provides the most KV‑cache space, enabling the highest concurrency.
Prefill Performance Measurements
Conclusion
For low concurrency, the TP4DP2+EP topology yields the best prefill throughput; at higher concurrency, TP2DP4+EP becomes dominant.
Results
Prefill throughput was measured with ISL/OSL = 8192/2. The curve shows a clear crossover point where TP2DP4+EP overtakes TP4DP2+EP as concurrency grows.
Decode Performance Measurements
Conclusion
MTP (speculative decoding) dramatically improves decode throughput until KV‑cache becomes the bottleneck; the best overall decode topology is TEP8 with MTP, followed by TP4DP4+EP at very high concurrency.
Results
Decode tests used ISL/OSL = 1/1000. Enabling MTP with three speculative tokens raised per‑GPU throughput, after which the TEP8 topology (with MTP) leads until KV‑cache limits are reached, after which TP4DP4+EP takes the lead.
Disaggregated Pareto Frontier
Conclusion
Combining the optimal prefill and decode configurations yields a final pareto frontier where the model achieves 5K total token throughput per GPU and 180 generated tokens per user.
Environment & Reproducibility
- Hardware: GB300 cluster, NVLink72 interconnect.
- Workload: ISL = 8192, OSL = 1024, concurrency 1–2560.
- Model:
Inferact/Qwen3.8-2.4T-A95B-NVFP4from HuggingFace. - Software stack:
- vLLM Docker image
vllm/vllm-openai:nightly-a9a17(revisionv0.26.1rc1.dev1177+ga9a17e709). - Dynamo 1.2.0.dev20260526.
- srt‑slurm v1.0.98.
- AIPerf v0.12.0.
- vLLM Docker image
- Recipes: All launch scripts are in the
srt‑slurm‑recipesrepository underrecipes/multi-node/Qwen3.8/GB300/8k1k/vllm/disagg.
Accuracy Verification
All configurations were validated on the GSM8K benchmark, achieving 95 % accuracy across the board, confirming that performance gains do not sacrifice model quality.
Performance Visuals
- Figure 3 – Pareto curves for each disaggregated configuration (individual sweeps over concurrency).
- Figure 4 – Combined pareto frontier showing the optimal trade‑off between total TPS and per‑user generation speed.
Implications for Practitioners
Conclusion
The presented tuning workflow—estimating KV‑cache, measuring activations and CUDA‑graph overhead, and selecting topologies per workload—enables practitioners to replicate frontier‑level performance for any large‑scale model on vLLM.
Practical Takeaways
- KV‑cache sizing is the primary limiter; compute the per‑request GDN state accurately to estimate maximum concurrency.
- Reserve enough GPU memory (default 8 % safety margin) to accommodate CUDA‑graph estimation errors.
- Choose decode topology based on KV‑cache headroom: TP4DP4+EP maximizes cache space, while TEP8 excels when cache is abundant.
- Enable MTP for decode workloads to boost throughput until KV‑cache becomes saturated.
- Use the provided srt‑slurm recipes to reproduce the exact environment and verify results on your own cluster.
Acknowledgements
Artem Perevedentsev (NVIDIA), Vadim Gimpelson (NVIDIA), Xin Li (NVIDIA), and the broader vLLM community for contributions and review.
Sources
- OriginalPD Serving of Qwen3.8-2.4T