vllm-metal v0.28.0 release notes / what's new

vllm-metal integrates vLLM's advanced scheduling and memory management into the Apple Silicon ecosystem, allowing Mac users to run an OpenAI-compatible server that handles overlapping requests efficiently. By combining the vLLM V1 scheduler and paged KV cache with MLX and Metal for execution, vllm-metal transforms local model execution into a scalable serving problem.

Core Architecture and Integration

vllm-metal acts as a plugin for upstream vLLM, leveraging the following components:

  • vLLM Frontend: Provides the V1 scheduler, paged KV block management, chunked prefill, sampling, and an OpenAI-compatible server with tool-call parsing and streaming.
  • MLX Execution: Uses mlx_lm for model implementations and MLX for hardware execution on Apple Silicon unified memory.
  • Custom Metal Kernels: While vllm-metal reuses mlx_lm's token-wise layers (RMSNorm, linear, MoE, and MLP), it replaces standard attention with a paged varlen Metal kernel. This kernel is a port of vLLM's unified Triton kernel, implementing a binary search over cu_seqlens to identify request ownership for each query token.

Memory Management and Paged KV Cache

vllm-metal implements a memory guard using the --gpu-memory-utilization flag to set a predictable inference budget, ensuring headroom for macOS and other applications. The system runs a warmup pass at startup to account for weights and activations, then allocates the remaining budget to a fixed KV cache.

To optimize throughput, vllm-metal uses packed queries and paged KV storage:

  • Packed Queries: Unlike mlx_lm's padded batches, vllm-metal packs all scheduled query tokens into a single [total_q, H, D] tensor. This allows the engine to process mixed steps (prefill and decode) in one model forward pass without padding to the longest request.
  • Paged KV: KV cache is stored in fixed-size pages addressed by per-request block tables, allowing requests to grow without requiring the reshaping of contiguous buffers.

Performance Benchmarks

Concurrent Serving and Latency

Using the SiliconBench agent split (100 multi-turn prompts, ~4.6K input tokens each) on an M5 Pro with 64 GB memory, vllm-metal demonstrated superior performance in several areas:

  • Qwen3.8-27B: vllm-metal achieved the lowest Time to First Token (TTFT) and end-to-end latency at concurrency levels 2 and 4.
  • Gemma 4 E4B: vllm-metal maintained low TTFT across a sweep up to concurrency 16.
  • Qwen3.6-35B-A3B: At concurrency 4, vllm-metal's throughput and end-to-end latency were comparable to oMLX with a RAM cache, though vllm-metal showed a significant advantage in TTFT.

Multi-Token Prediction (MTP)

For Gemma 4 E4B, vllm-metal supports speculative decoding via an MTP assistant model. At concurrency 1, MTP reduced wall time by 15% and increased output throughput by 20% compared to generation without MTP.

M5 Hardware Acceleration

On M5 Macs, vllm-metal automatically employs the NAX attention kernel, which utilizes the GPU's tensor hardware for compatible prefill batches, resulting in faster prefill times compared to tiled attention.

Feature Set in v0.28.0

The v0.28.0 release introduces several key capabilities:

  • Model Support: Support for GGUF checkpoints and hybrid-attention models from the Qwen3.5, Qwen3.6, Qwen3.8, and Qwen3-Next families.
  • Speculative Decoding: Three methods are supported: Gemma 4 MTP, separate draft models, and prompt-lookup n-grams.
  • Advanced Serving: LoRA adapters, structured outputs, and pipeline parallelism across multiple Macs via the MLX ring backend.
  • Experimental Features: Support for vision-language models, text embeddings, reranking, and speech-to-text.
  • Prefix Caching: For Qwen3.5-style hybrid models, vllm-metal supports align mode, saving GDN recurrent state at block boundaries to allow resuming from a cached prefix.

Installation and Usage

vllm-metal requires macOS 15 or later and can be installed via Homebrew:

brew tap vllm-project/vllm-metal https://github.com/vllm-project/vllm-metal
brew install vllm-project/vllm-metal/vllm-metal

An OpenAI-compatible server can be launched using the vllm serve command, specifying the model and memory utilization budget.

Sources