51

vLLM AMD ROCm Attention Backends Optimization

vLLM introduces optimized attention backends for AMD ROCm, delivering up to 4.4x higher throughput for MHA and 1.5x for MLA models on Instinct MI300X, MI325X, and MI355X GPUs.

52

vLLM Multi-LoRA Serving for MoE Models

vLLM version 0.15.0 introduces optimized Multi-LoRA serving for Mixture of Experts (MoE) models, enabling multiple fine-tuned adapters to share a single GPU to reduce idle compute capacity.