FastFlowLM/FastFlowLM
Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.
What it solves
FastFlowLM (FLM) allows users to run large language models (LLMs) and vision-language models (VLMs) locally on AMD Ryzen™ AI NPUs without requiring a GPU or CPU load. It provides a power-efficient, lightweight alternative to GPU-based inference, supporting context lengths up to 256k tokens.
How it works
FLM acts as an NPU-first runtime that optimizes model kernels for AMD's XDNA2 NPU architecture (supporting Strix, Strix Halo, Kraken, and Gorgon Point chips). It offers a CLI for running models directly in the terminal and a local server mode that provides REST and OpenAI-compatible APIs for integration into other applications.
Who it’s for
Developers and users with AMD Ryzen™ AI Series chips who want to run local AI models privately and efficiently without relying on heavy GPU resources.
Highlights
- NPU-Optimized: Runs fully on AMD Ryzen™ AI NPUs, reducing power consumption and freeing up the GPU/CPU.
- Ultra-Lightweight: The runtime is only 17 MB and installs in under 20 seconds.
- Developer-Friendly: Features a CLI and API similar to Ollama, making it easy to deploy and switch models.
- Broad Model Support: Supports Vision, Audio, Embedding, and MoE models.
- High Capacity: Supports context windows up to 256k tokens.