aws-neuron/aws-neuron-sdk

Powering AWS purpose-built machine learning chips. Blazing fast and cost effective, natively integrated into PyTorch and TensorFlow and integrated with your favorite AWS services

What it solves

It provides a high-performance environment for developing, profiling, and deploying deep learning and generative AI workloads specifically for AWS's own AI hardware accelerators (Inferentia and Trainium).

How it works

The SDK includes a graph compiler and runtime that translates machine learning models into a format the hardware can execute. It integrates with popular frameworks like PyTorch and JAX, and supports vLLM Neuron for inference serving. For advanced users, the Neuron Kernel Interface (NKI) allows for the direct programming of NeuronCore for custom kernels.

Who it’s for

Machine learning engineers and developers who are using AWS EC2 instances (such as Inf1, Inf2, Trn1, Trn2, and Trn3) to train or run inference on large-scale AI models.

Highlights

  • High-performance deep learning and generative AI support
  • Integration with PyTorch, JAX, and vLLM
  • Direct hardware programming via the Neuron Kernel Interface (NKI)
  • Dedicated developer tools including Neuron Explorer

Related

  • Project
  • Project
  • Project
  • Dispatch
  • Dispatch