OpenDLSS-NR: Vulkan Reimplementation of NVIDIA DLSS 5 Neural Rendering (Bit‑Exact)
TL;DR
OpenDLSS‑NR reproduces NVIDIA’s DLSS 5 neural‑rendering network on Vulkan with bit‑exact output, matching every intermediate tensor of the original 71‑block Swin/ViT model, and runs at sub‑10 ms per 1080p frame on an RTX 4070 SUPER.
What the project delivers
- Exact network replica – 71 transformer blocks (six pooling levels) identical to DLSS‑NR build 310.8.0, using FP8 (E4M3) activations and FP16 accumulation. All 75 block boundaries match the original byte‑for‑byte.
- Vulkan‑native implementation – C++20 host, GLSL reference kernels, and hand‑written PTX kernels that exploit NVIDIA cooperative‑matrix FP8 GEMMs, cp.async rings, and barrier‑free counter chaining.
- WebGPU fallback – A second port runs the same network in a browser without tensor cores or FP8, achieving functional parity (bit‑exactness is defined by the specification, not hardware).
- Demo application – Integrates the network into a patched Filament renderer, exposing ImGui controls and a library of glTF scenes.
- Open‑source tooling – Scripts to fetch Vulkan headers, glslang, PTX generators, and Filament; no full Vulkan SDK required.
Network architecture at a glance
The model is a U‑Net composed of shifted‑window transformer blocks topped by a global ViT. Key specifications:
| Attribute | Value |
|---|---|
| Blocks | 71 (75 block boundaries) |
| Transformer type | Swin‑window + ViT |
| Precision | FP8 (E4M3) activations, FP16 accumulation |
| Weight size | 141 MiB |
| Input | Low‑dynamic‑range proxy frame, three Gaussian‑noise lanes, reprojected previous output, five conditioning scalars |
| Output | Four f32 channels per pixel: RGB residual + temporal‑blend logit |
The network does not upscale; it re‑renders the same‑resolution frame, adding detail and adjusting tone.
Build, run, and verification workflow
- Fetch toolchain –
scripts\fetch_tools.ps1downloads glslang 16.6.0, Vulkan‑Headers v1.4.363, Volk, CMake 3.31, and Ninja 1.13. - Compile –
scripts\build.ps1builds shaders, PTX, and thedlss5vk.exedriver. - Prepare Filament demo –
scripts\fetch_filament.ps1andscripts\build_filament.ps1clone Filament v1.77.0, apply the motion‑vector patch, and build the demo binary. - Run –
build\dlss5vk.exe bench --model <dir> --width 768 --height 768measures performance;paritychecks bit‑exactness against a fixture;verifyperforms kernel‑by‑kernel bisects. - Model provision – Users supply a directory containing
manifest.jsonand packed E4M3 weight files; the repository does not provide weights.
Performance numbers (RTX 4070 SUPER)
The whole network is executed as 241 dispatches per frame. Median timings over 40 frames are:
| Resolution | Time per frame |
|---|---|
| 768×768 | 2.8 ms |
| 1920×1080 | 7.8 ms |
| 2560×1440 | 12.6 ms |
| 3840×2160 | 29.3 ms |
GPU clock throttling can raise medians a few percent; the table reports minima for a fair comparison.
How exactness is guaranteed
- Parity mode –
dlss5vk parity --fixture <dir>runs the network against captured reference outputs from NVIDIA’s implementation. It validates every tensor at each of the 75 block boundaries. - Numerics contract – The
docs/numerics.mdfile defines four verdict categories: bit‑exact (pass), sign‑of‑zero (acceptable), within‑one‑code (8‑bit capture tolerance), and mismatch. - Kernel‑by‑kernel verification –
verifyrequires the fixture to contain the block‑0 reference and input features, allowing a fine‑grained bisect of any discrepancy. - Switchable routes – Environment variables (e.g.,
DLSS5VK_UNFUSED=1) let developers replace PTX kernels with the GLSL reference while preserving output identity, useful for debugging or hardware without FP8 support.
Community insights from Hacker News
"bit‑exact against the original – what kind of sorcery is this? Very impressive work!" – TheJCDenton
"Almost 8 ms on 1080p seems extremely expensive; does the original also eat into the rendering budget as much?" – flohofwoe
"I'm surprised it doesn't take the z‑buffer as input; that would seem useful for a neural renderer." – Lerc
"How useful is this without weights? NVIDIA ships per‑game models; is a single generalized model feasible?" – franticgecko3
These comments highlight two recurring themes: the technical achievement of byte‑level parity, and practical questions about performance impact and the availability of pretrained weights.
Limitations and missing pieces
- No DLSS‑SR support – The repository implements only the neural‑rendering (NR) path; the super‑resolution variant is absent.
- Weights are external – Users must obtain or generate their own E4M3 weight files; the repo contains no proprietary NVIDIA weights.
- Temporal history only in demo – The
dlss5vktool processes single frames without history; the demo uses reprojected previous output for temporal stability. - Hardware requirements – Requires an Ada‑generation NVIDIA GPU with Vulkan extensions
VK_KHR_cooperative_matrix,VK_NV_cooperative_matrix2,VK_EXT_shader_float8, andVK_NV_cuda_kernel_launch.
Why this matters
OpenDLSS‑NR proves that a complex, proprietary AI‑accelerated graphics pipeline can be faithfully reproduced in open‑source Vulkan, enabling independent research, cross‑platform experimentation (including WebGPU), and transparent performance benchmarking. The project's rigorous parity testing sets a new standard for reproducibility in neural rendering.
Getting started checklist
- Hardware – NVIDIA Ada‑generation GPU (RTX 40‑series or newer) with the required Vulkan extensions.
- Toolchain – Windows, Visual Studio 2022 (C++ x64), Git, Python 3, Node.js/npm, Pillow.
- Clone repo –
git clone https://github.com/maanHimself/OpenDLSS-NR.git. - Run fetch scripts – Execute the PowerShell scripts in
scripts/. - Provide model – Place a directory with
manifest.jsonand packed weight files; set--model <dir>orDLSS5VK_MODEL. - Run demo –
build\dlss5vk.exe bench …or double‑click the demo executable.
License
The repository is released under the MIT license; third‑party components are listed in NOTICE.
Sources
Related
- Project
- Project
- Project
- Project
- Project