InfiniTensor/InfiniTensor
InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation.
What it solves
InfiniTensor 是为 GPU 和 AI 加速器设计的高性能推理引擎,旨在简化 AI 模型部署以及对新优化技术进行学术验证的过程。
How it works
它作为 AI 模型执行的后端,支持包括 NVIDIA GPUs、Cambricon MLUs、Kunlunxin XPUs、Ascend NPUs 和 Intel CPUs 在内的广泛硬件支持。该项目集成了像 PET(使用部分等价转换)这样的张量程序优化器,并正朝着名为 RefactorGraph 的新框架设计迈进。
Who it's for
它主要针对需要在各种硬件加速器上优化 AI 模型推理的研究人员和开发者。
Highlights
- 高性能的 GPU 和 AI 加速器推理引擎,专为这些硬件量身定制。
- 广泛的硬件支持,包括 NVIDIA、Cambricon、Kunlunxin、Ascend 和 Intel。
- 集成了像 PET 这样的张量程序优化器。
- 提供 Python frontend 以便于使用。
相关
- 项目
- 项目
- 项目
- 项目
- 项目