TurboGPT: High-Speed Tiny Transformer Training in CUDA C++

TurboGPT is a specialized implementation for training tiny, byte-level GPT models using CUDA C++. The project demonstrates extreme training efficiency, enabling a 22KiB transformer to be trained in approximately 13 seconds.

Technical Implementation and Performance

TurboGPT is written primarily in C++ (40.6%) and CUDA (38.5%), with Python used for verification tests. It is designed specifically for NVIDIA GPUs, requiring CUDA 13.4 and Visual Studio 2022 C++ tools for Windows builds, or Nix for Linux/NixOS builds.

Key technical features include:

  • Byte-level Training: The model operates on bytes rather than subword tokens.
  • Rotary Positional Embeddings (RoPE): Recent updates to the codebase replaced learned positions with partial RoPE to improve positional encoding.
  • PyTorch Compatibility: The system generates checkpoints in a format compatible with torch.load, containing the model, optimizer, scheduler, and trainer state.
  • Observability: Training logs are TensorBoard-compatible, with reports generated per batch and capped at 8 million reports.

In terms of performance, the project reports a result of 2.5295 BPB (Bits Per Byte) on the hn1g dataset after 1.5 billion training tokens.

Build and Execution Workflow

TurboGPT requires a GPU with a specific compute capability (CudaArch). Users can build the project using the following methods:

  • Linux/NixOS: Using nix-build -o build/nix-result.
  • Windows: Using .\build.ps1 -CudaArch <arch>, where <arch> corresponds to the GPU's compute capability.

To execute a training run, the user specifies the data file and a log directory: .\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4

Training can be resumed from a previous state using the --load CHECKPOINT.pt flag.

Community Discussion and Critique

While the project showcases technical speed, community feedback on Hacker News highlights a debate regarding the utility of such "tiny" implementations compared to established educational resources.

One contributor questioned the motivation behind the proliferation of these projects, stating:

I've seen a hundred of them at this point- and each of them is probably worse and has less learning value than the one Andrej Karpathy made to teach people the building blocks involved in a GPT

Other technical observations included a critique of the trend of labeling models based on their disk size (e.g., 22KiB) rather than their count of learnable parameters, and a suggestion that at such a small scale, traditional optimization problems (like KKT conditions) might be more applicable than iterative gradient descent.

Sources

Related

  • Project
  • Dispatch
  • Dispatch
  • Project
  • Dispatch