SemiAnalysisAI/InferenceX

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

What it solves

InferenceX addresses the problem of stale AI benchmarks. Because LLM inference software (like vLLM and SGLang) evolves daily through kernel optimizations and scheduling innovations, benchmarks taken at a single point in time quickly become outdated and fail to reflect the actual performance of the latest software releases.

How it works

It functions as an automated, continuous research platform that tracks and benchmarks popular open-source inference frameworks in near real-time. By continuously running evaluations across a wide array of high-end hardware (including NVIDIA Blackwell, Hopper, and AMD Instinct GPUs), it captures incremental performance gains as software stacks improve, providing a live indicator of the performance frontier.

Who it’s for

This platform is designed for operators of large-scale "token factories" (such as OpenAI, Meta, and Microsoft), ML researchers, and developers using inference frameworks who need accurate, up-to-date data on hardware and software performance.

Highlights

  • Continuous Tracking: Moves at the speed of the software ecosystem to provide real-time performance updates.
  • Broad Hardware Support: Supports a vast range of cutting-edge GPUs, including GB300 NVL72, B200, MI355X, and H100.
  • Framework Integration: Analyzes major stacks including SGLang, vLLM, TensorRT-LLM, CUDA, and ROCm.
  • Public Dashboard: Offers a free, open-source live dashboard for public performance tracking.

Related

  • Project
  • Project
  • Project
  • Project
  • Project