ROCm/FastFlowLM
Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.
What it solves
FastFlowLM (FLM) allows users to run large language models (LLMs) and vision-language models (VLMs) directly on AMD Ryzen™ AI NPUs. This removes the need for a dedicated GPU, significantly reducing power consumption and freeing up CPU/GPU resources.
How it works
FLM provides an ultra-lightweight (17 MB) NPU-first runtime and a single-command CLI. It uses optimized binary kernels (built with IRON and AIE-MLIR) to execute models on XDNA2 NPUs. It supports a variety of model types, including Vision, Audio, Embedding, and Mixture-of-Experts (MoE), and can be deployed as a local terminal chat or a REST/OpenAI-compatible API server.
Who it’s for
Developers and users with AMD Ryzen™ AI Series chips (specifically those with XDNA2 NPUs) who want to run local AI models privately and efficiently without relying on a GPU.
Highlights
- NPU-Native Execution: Runs fully on the AMD Ryzen™ AI NPU, bypassing CPU and GPU load.
- High Efficiency: Over 10x more power-efficient than traditional methods.
- Long Context Support: Handles context windows up to 256k tokens.
- Developer Friendly: Features a simple CLI, fast installation (under 20 seconds), and an OpenAI-compatible API server.
- Broad Model Support: Compatible with LLMs, VLMs, and robotics policies like SmolVLA.
相关
- 项目
- 项目
- 项目
- 项目
- 项目