xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

What it solves

Large language model deployment can be costly and inefficient on standard hardware. xLLM provides an enterprise-grade solution to enable high-throughput, low-latency inference specifically optimized for Chinese AI accelerators, reducing deployment costs for large-scale applications.

How it works

It utilizes a decoupled architecture where a service layer manages scheduling and availability, while an engine layer handles the actual computation. The framework includes advanced features such as hybrid KV cache management with intelligent offloading and prefetching to maximize efficiency.

Who it’s for

It is designed for enterprises and developers needing to deploy large language models on specialized Chinese hardware, such as Ascend NPUs, Cambricon MLUs, Moore Threads GPUs, and other regional AI accelerators.

Highlights

  • High Performance: Delivers top-tier throughput and low latency.
  • Specialized Hardware Support: Deeply optimized for a wide range of Chinese AI accelerators including Ascend, Cambricon, Moore Threads, Hygon, MetaX, and Iluvatar.
  • Decoupled Architecture: Separates service management from computational execution.
  • Battle-tested: Proven at scale within JD.com's core retail business operations.