maanHimself/OpenDLSS-NR
A Vulkan reimplementation of NVIDIA's DLSS 5 Neural Rendering network, bit-exact against the original.
OpenDLSS‑NR – Vulkan re‑implementation of NVIDIA’s DLSS 5 Neural Rendering
What it is
- A faithful, bit‑exact recreation of NVIDIA’s DLSS 5 Neural Rendering (NR) network, written in C++20 and Vulkan. It runs the same 71‑block Swin‑ViT U‑Net that NVIDIA ships in DLSS 5 build 310.8.0, using FP8 (E4M3) arithmetic on tensor cores.
- The repository also contains a pure WebGPU version that runs in a browser (no tensor cores, no FP8) and a demo that plugs the network into the open‑source Filament renderer.
Why it matters
- DLSS 5’s NR is a generative neural‑rendering step that adds detail, corrects tone and improves skin quality without changing the output resolution (it is not an up‑scaler). OpenDLSS‑NR lets researchers and developers experiment with the exact same model on their own hardware, inspect intermediate activations, and benchmark performance.
- All intermediate tensors and the final image match the proprietary implementation byte‑for‑byte, which is rare for a reverse‑engineered AI model.
Key components
| Component | What it provides |
|---|---|
src/ |
Host side C++20 code: Vulkan setup, model loading, weight layout, a CPU reference implementation, and the dlss5vk command‑line tool. |
shaders/ |
GLSL kernels that implement the exact FP8 GEMM, attention, MLP, and other blocks. These are the reference path (no fusion). |
scripts/ptx/ |
Python generators that emit low‑level PTX kernels (using mma.sync FP8 × FP16) for the high‑performance path with fused operations and counter‑based chaining. |
demo/ |
A Filament‑based viewer that loads glTF scenes, shows ImGui controls, and runs the network on each frame (including the temporal feedback loop). |
ports/browser-webgpu/ |
A WebGPU port that reproduces the same numerical results in a browser, albeit much slower (≈72 ms at 512×512). |
How to get it running
- Prerequisites – Windows, an NVIDIA Ada‑generation GPU (or newer) with the required Vulkan extensions, Visual Studio 2022, Python 3, Node.js, and the Vulkan‑Headers/GLSLang toolchain (the repo can fetch these automatically).
- Fetch tools –
scripts\fetch_tools.ps1(adds glslang, Volk, CMake, Ninja, etc.). - Build the engine –
scripts\build.ps1compiles shaders, PTX and producesbuild\dlss5vk.exe. - Optional demo setup – fetch and patch Filament (
fetch_filament.ps1/build_filament.ps1), then build the demo (build_demo.ps1). - Run –
dlss5vk.exe bench --model <model_dir> --width 768 --height 768for a quick timing.dlss5vk.exe parity …to compare against reference fixtures (bit‑exactness test).demo\dlss5-demo.exeto launch the interactive viewer (drag‑and‑drop a glTF or pick a built‑in scene).
Performance snapshot (RTX 4070 SUPER, median over 40 frames, whole network per frame):
| Resolution | Time per frame |
|---|---|
| 768×768 | 2.8 ms |
| 1920×1080 | 7.8 ms |
| 2560×1440 | 12.6 ms |
| 3840×2160 | 29.3 ms |
Model handling
- The repository does not ship the 141 MiB of FP8 weights. Users must supply a model directory that follows the
manifest.jsonlayout described in the README. The loader validates the manifest, re‑layouts the packed bytes into the matrix forms expected by the kernels, and rejects any model with a different block count.
Verification & exactness
parityandverifycommands compare the GPU output against NVIDIA‑provided capture fixtures (not included). The comparison is bit‑exact; failures are reported as sign‑of‑zero differences, one‑code deviations (for 8‑bit captures), or full mismatches. The PTX‑generated kernels and the GLSL reference route are both capable of passing these tests when the appropriate switches are used.
Tuning knobs
- Environment variables (
DLSS5VK_UNFUSED,DLSS5VK_CHAIN,DLSS5VK_PTX_GEMM, etc.) let you toggle between the fast fused PTX path and the slower but easier‑to‑debug GLSL path, enable/disable split‑K GEMMs, control barrier insertion, and more—all while preserving numerical identity.
Limitations
- Only the NR network is implemented; the DLSS‑SR (super‑resolution) branch is absent.
- Requires a recent Ada‑generation NVIDIA GPU and the specific Vulkan extensions; it will not run on older hardware or on non‑Windows platforms.
- Model weights must be obtained separately and are subject to NVIDIA’s licensing.
License
- The code in the repository is MIT‑licensed. Third‑party components (Filament, glslang, etc.) are listed in
NOTICE.
OpenDLSS‑NR gives the community a transparent, reproducible platform for studying NVIDIA’s latest generative neural‑rendering technology, complete with low‑level PTX kernels, a reference GLSL path, and a ready‑to‑run demo.
Related
- Project
- Project
- Project
- Project