InfiniTensor/InfiniTensor
InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation.
What it solves
InfiniTensor is designed to provide a high-performance inference engine for GPUs and AI accelerators, simplifying the process of deploying AI models and conducting academic validation of new optimizations.
How it works
It functions as a backend for AI model execution, supporting a wide range of hardware targets including NVIDIA GPUs, Cambricon MLUs, Kunlunxin XPUs, Ascend NPUs, and Intel CPUs. The project integrates tensor program optimizers like PET (which uses partially equivalent transformations) and is moving towards a new framework design called RefactorGraph.
Who it’s for
It is primarily aimed at researchers and developers who need to optimize AI model inference on diverse hardware accelerators.
Highlights
- High-performance inference engine tailored for GPUs and AI accelerators.
- Broad hardware support including NVIDIA, Cambricon, Kunlunxin, Ascend, and Intel.
- Integration with tensor program optimizers like PET.
- Python frontend for easier accessibility.
Related
- Project
- Project
- Project
- Project
- Project