ROCm/FastFlowLM

Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.

What it solves

FastFlowLM (FLM) allows users to run large language models (LLMs) and vision-language models (VLMs) directly on AMD Ryzen™ AI NPUs. This removes the need for a dedicated GPU, significantly reducing power consumption and freeing up CPU/GPU resources.

How it works

FLM provides an ultra-lightweight (17 MB) NPU-first runtime and a single-command CLI. It uses optimized binary kernels (built with IRON and AIE-MLIR) to execute models on XDNA2 NPUs. It supports a variety of model types, including Vision, Audio, Embedding, and Mixture-of-Experts (MoE), and can be deployed as a local terminal chat or a REST/OpenAI-compatible API server.

Who it’s for

Developers and users with AMD Ryzen™ AI Series chips (specifically those with XDNA2 NPUs) who want to run local AI models privately and efficiently without relying on a GPU.

Highlights

  • NPU-Native Execution: Runs fully on the AMD Ryzen™ AI NPU, bypassing CPU and GPU load.
  • High Efficiency: Over 10x more power-efficient than traditional methods.
  • Long Context Support: Handles context windows up to 256k tokens.
  • Developer Friendly: Features a simple CLI, fast installation (under 20 seconds), and an OpenAI-compatible API server.
  • Broad Model Support: Compatible with LLMs, VLMs, and robotics policies like SmolVLA.

相关

  • 项目
  • 项目
  • 项目
  • 项目
  • 项目