TurboGPT: High-Speed Tiny Transformer Training in CUDA C++
TurboGPT is a specialized implementation for training tiny, byte-level GPT models using CUDA C++. The project demonstrates extreme training efficiency, enabling a 22KiB transformer to be trained in approximately 13 seconds.
Technical Implementation and Performance
TurboGPT is written primarily in C++ (40.6%) and CUDA (38.5%), with Python used for verification tests. It is designed specifically for NVIDIA GPUs, requiring CUDA 13.4 and Visual Studio 2022 C++ tools for Windows builds, or Nix for Linux/NixOS builds.
Key technical features include:
- Byte-level Training: The model operates on bytes rather than subword tokens.
- Rotary Positional Embeddings (RoPE): Recent updates to the codebase replaced learned positions with partial RoPE to improve positional encoding.
- PyTorch Compatibility: The system generates checkpoints in a format compatible with
torch.load, containing the model, optimizer, scheduler, and trainer state. - Observability: Training logs are TensorBoard-compatible, with reports generated per batch and capped at 8 million reports.
In terms of performance, the project reports a result of 2.5295 BPB (Bits Per Byte) on the hn1g dataset after 1.5 billion training tokens.
Build and Execution Workflow
TurboGPT requires a GPU with a specific compute capability (CudaArch). Users can build the project using the following methods:
- Linux/NixOS: Using
nix-build -o build/nix-result. - Windows: Using
.\build.ps1 -CudaArch <arch>, where<arch>corresponds to the GPU's compute capability.
To execute a training run, the user specifies the data file and a log directory:
.\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4
Training can be resumed from a previous state using the --load CHECKPOINT.pt flag.
Community Discussion and Critique
While the project showcases technical speed, community feedback on Hacker News highlights a debate regarding the utility of such "tiny" implementations compared to established educational resources.
One contributor questioned the motivation behind the proliferation of these projects, stating:
I've seen a hundred of them at this point- and each of them is probably worse and has less learning value than the one Andrej Karpathy made to teach people the building blocks involved in a GPT
Other technical observations included a critique of the trend of labeling models based on their disk size (e.g., 22KiB) rather than their count of learnable parameters, and a suggestion that at such a small scale, traditional optimization problems (like KKT conditions) might be more applicable than iterative gradient descent.
Sources
Related
- Project
- Dispatch
- Dispatch
- Project
- Dispatch