SemiAnalysisAI/InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
What it solves
InferenceX addresses the problem of stale benchmarks in the rapidly evolving LLM inference ecosystem. Because software stacks (like vLLM and SGLang) improve daily through kernel optimizations and scheduling innovations, static benchmarks quickly become outdated and fail to represent the actual performance achievable with the latest software releases.
How it works
It functions as an automated, continuous benchmarking research platform. It tracks the real-time performance of popular open-source inference frameworks across a wide array of high-end hardware, including NVIDIA (Blackwell, Hopper) and AMD (MI300 series) GPUs, providing a live indicator of performance progress via a public dashboard.
Who it’s for
This platform is designed for operators of large-scale "token factories" (such as OpenAI, Meta, Microsoft, and Oracle), ML researchers, and maintainers of inference frameworks who need to track real-time performance gains and hardware comparisons.
Highlights
- Continuous Tracking: Captures software-driven performance gains in near real-time rather than at fixed points in time.
- Broad Hardware Support: Officially supports a vast range of cutting-edge accelerators, including GB200, B200, MI325X, and H100.
- Framework Agnostic: Benchmarks a variety of industry-standard stacks including SGLang, vLLM, and TensorRT-LLM.
- Public Transparency: Provides a free, open-source live dashboard for the community to monitor inference performance curves.