FastFlowLM/FastFlowLM

Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.

What it solves

FastFlowLM (FLM) allows users to run large language models (LLMs) and vision-language models (VLMs) locally on AMD Ryzen™ AI NPUs without requiring a GPU or CPU load. It provides a power-efficient, lightweight alternative to GPU-based inference, supporting context lengths up to 256k tokens.

How it works

FLM acts as an NPU-first runtime that optimizes model kernels for AMD's XDNA2 NPU architecture (supporting Strix, Strix Halo, Kraken, and Gorgon Point chips). It offers a CLI for running models directly in the terminal and a local server mode that provides REST and OpenAI-compatible APIs for integration into other applications.

Who it’s for

Developers and users with AMD Ryzen™ AI Series chips who want to run local AI models privately and efficiently without relying on heavy GPU resources.

Highlights

  • NPU-Optimized: Runs fully on AMD Ryzen™ AI NPUs, reducing power consumption and freeing up the GPU/CPU.
  • Ultra-Lightweight: The runtime is only 17 MB and installs in under 20 seconds.
  • Developer-Friendly: Features a CLI and API similar to Ollama, making it easy to deploy and switch models.
  • Broad Model Support: Supports Vision, Audio, Embedding, and MoE models.
  • High Capacity: Supports context windows up to 256k tokens.