InfiniTensor/InfiniTensor

InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation.

What it solves

InfiniTensor 是为 GPU 和 AI 加速器设计的高性能推理引擎,旨在简化 AI 模型部署以及对新优化技术进行学术验证的过程。

How it works

它作为 AI 模型执行的后端,支持包括 NVIDIA GPUs、Cambricon MLUs、Kunlunxin XPUs、Ascend NPUs 和 Intel CPUs 在内的广泛硬件支持。该项目集成了像 PET(使用部分等价转换)这样的张量程序优化器,并正朝着名为 RefactorGraph 的新框架设计迈进。

Who it's for

它主要针对需要在各种硬件加速器上优化 AI 模型推理的研究人员和开发者。

Highlights

  • 高性能的 GPU 和 AI 加速器推理引擎,专为这些硬件量身定制。
  • 广泛的硬件支持,包括 NVIDIA、Cambricon、Kunlunxin、Ascend 和 Intel。
  • 集成了像 PET 这样的张量程序优化器。
  • 提供 Python frontend 以便于使用。

相关

  • 项目
  • 项目
  • 项目
  • 项目
  • 项目