LFM2.5-VL-3B DSpark draft model released with up to 3.13x faster vision-language inference

TL;DR

Hugging Face announced the LFM2.5‑VL‑3B‑DSpark draft model, a speculative decoding add‑on that speeds up vision‑language inference by up to 3.13× on edge devices and 2.66× on H100 GPUs, while increasing the total parameter count by only 8.9 %.


What is LFM2.5‑VL‑DSpark?

LFM2.5‑VL‑DSpark is an experimental draft model that sits alongside the base vision‑language model LFM2.5‑VL‑3B. It implements speculative decoding: a lightweight drafter proposes a block of candidate tokens, which the target model then verifies. The approach preserves exact output quality because every token is ultimately validated by the full model.

Key properties:

  • Memory impact: adds 280 M parameters (≈8.9 % of the 3 B‑parameter target).
  • Speed gains: up to 3.13× faster token decoding on Apple silicon, 2.66× on an NVIDIA H100, with end‑to‑end latency reductions of 2.62× (device) and 2.27× (GPU).
  • Day‑one integration: compatible with llama.cpp, MLX‑VLM, and SGLang.

How speculative decoding works for vision‑language models

The drafter shares the same architecture as the text‑only LFM2.5‑DSpark drafters. It taps hidden states from a fixed set of layers in the target model, projects both image patches and text tokens into a shared representation, and drafts a block of k candidate tokens.

Because the projection yields vectors of identical dimensionality for both modalities, the inference algorithm is unchanged from the text case. The drafter therefore operates purely on hidden‑state vectors, regardless of whether the input is visual, textual, or a mixture of both.

DSpark‑Vision architecture diagram


Training recipe and architecture details

  • Data mix: a curated mixture of vision‑language supervised‑fine‑tuning (SFT) data weighted toward anticipated deployment workloads.
  • Layer depth: ablations across 3, 4, and 5 layers identified a 4‑layer attention‑only drafter as optimal.
  • Block size: training used a block size of 9; inference recommends a block size of 8 or 9 depending on hardware.
  • Training duration: 10 epochs over the final data mixture; acceptance rates improved with more tokens before plateauing.

Parameter breakdown (all numbers in millions):

Component Parameters
Decoder stack (4 layers) 193.0
Hidden‑state projection 21.0
Markov head 65.5
Norms + confidence head 0.006
Total 279.5

Measured inference speedups

On‑device (Apple silicon)

  • MLX on M5 Max: decoding 2.30×–3.13× faster; end‑to‑end latency 1.56×–2.62× faster.
  • llama.cpp on M3 Ultra: decoding 1.57×–2.14× faster; end‑to‑end latency 1.30×–1.77× faster.

On‑device speedup screenshots

GPU (NVIDIA H100)

  • Decoding speedup: 2.04×–2.66×.
  • End‑to‑end latency improvement: 1.64×–2.27×.

GPU speedup screenshots

All measurements use a DSpark block size of 8 and evaluate six MMSpec benchmark tasks (general VQA, text VQA, image captioning, chart VQA, complex reasoning, multi‑turn conversation).


Limitations of speculative decoding for vision workloads

Speculative decoding only accelerates the decode phase; it does not speed up image encoding or the prefill stage where the vision encoder processes image patches. On edge devices, prefill can dominate wall‑clock time, so overall latency gains are bounded by Amdahl’s law. Consequently, even large decode speedups may translate to modest end‑to‑end improvements when encoding or prefill already consume a substantial portion of the runtime.


Getting started with LFM2.5‑VL‑DSpark

SGLang

python -m sglang.launch_server \
  --model-path LiquidAI/LFM2.5-VL-3B \
  --speculative-algorithm DSPARK \
  --speculative-draft-model-path LiquidAI/LFM2.5-VL-3B-DSpark \
  --speculative-draft-attention-backend flashinfer \
  --speculative-dspark-block-size 9 \
  --disable-radix-cache

Query the OpenAI‑compatible endpoint at http://localhost:30000/v1.

llama.cpp

llama-server -m models/LFM2.5-VL-3B-F16.gguf \
  --mmproj models/mmproj-LFM2.5-VL-3B-F16.gguf \
  -md LFM2.5-2.6B-DSpark-F16.gguf \
  --spec-type draft-dspark --spec-draft-n-max 8 --spec-draft-n-min 0 \
  -fa on -ngl 99 -c 8192

MLX‑VLM

mlx_vlm.server --model LiquidAI/LFM2.5-VL-3B --draft-model LiquidAI/LFM2.5-VL-3B-DSpark

All three integrations read the block size from the draft model’s config.json. The speculative decoding is exact: the target model validates every drafted token, guaranteeing identical output to running the target alone.


Availability and licensing

The draft model is hosted on Hugging Face in both Safetensors and GGUF formats:

LFM2.5 models are open‑weight, allowing unrestricted download, fine‑tuning, and deployment.


Citation

Please cite the announcement as follows:

Liquid AI, "LFM2.5-VL-DSpark: Accelerating vision-language models on edge and beyond", Liquid AI Blog, Sep 2026.

Or use the provided BibTeX entry:

@article{liquidAI2026vldspark,
  author = {Liquid AI},
  title = {LFM2.5-VL-DSpark: Accelerating vision-language models on edge and beyond},
  journal = {Liquid AI Blog},
  year = {2026},
  note = {www.liquid.ai/blog/lfm2-5-vl-dspark},
}

Sources