npunlock enables custom C kernels on Intel Core Ultra NPUs
TL;DR
npunlock provides a workflow to compile arbitrary C code into machine‑code for Intel's SHAVE cores inside the Core Ultra NPU3720 and execute those kernels inside a standard OpenVINO‑style graph, extending the NPU beyond Intel’s officially supported operation set.
What npunlock does
npunlock reconstructs the missing path from user‑written C to a runnable NPU kernel. It:
- Compiles C source with Intel/Movidius MoviTools into ACT‑SHAVE machine code.
- Packages the compiled kernel as a custom operation that can be inserted into an OpenVINO‑compatible graph.
- Uses Intel’s existing driver and compiler for the surrounding graph, so only the custom node is handled by npunlock.
- Emits OpenVINO‑format IR, allowing the rest of the pipeline to remain unchanged.
"Intel ships programmable SHAVE cores inside its NPUs, but the public stack exposes only graph‑level programming.
npunlockreconstructs the missing path from custom C code to a runnable NPU kernel." – npunlock README
Why this matters
Intel’s NPU software only accepts graphs composed of operations that the proprietary compiler knows about. There is no public API for supplying a hand‑written C implementation for a new operation, effectively locking developers out of the SHAVE cores. npunlock opens that lock, enabling:
- Research on novel operators that are not yet (or never will be) supported by Intel.
- Fine‑grained performance tuning by writing hand‑optimized kernels.
- Exploration of mixed‑precision pipelines that combine FP32 unary and FP16 binary custom branches in a single graph (a breakthrough announced on 2026‑09‑23).
Quick start example (FP32 GELU)
The following Python snippet demonstrates a complete end‑to‑end flow:
import numpy as np, npunlock as npu
npu.configure(movi_dll_dir=r"C:\path\to\MVC_DEPEND")
gelu_c = b"""
#define MLIBM_DEFINE_LINK_COMPAT 1
#include <npunlock/npu3720_kernel.h>
void controlled_act(unsigned layerParams) {
act_abi_invocation invocation;
ACT_ABI_LOAD_INVOCATION32_OR_RETURN(layerParams, invocation);
const float *in = ACT_ABI_INPUT_PTR32(const float, invocation, 0u);
float *out = ACT_ABI_OUTPUT_PTR32(float, invocation, 1u);
const float SQRT_2_DIV_PI = 0.7978845608028654f;
for (unsigned i = 0; i < invocation.element_count; ++i) {
float x = in[i];
float w = x + 0.044715f * x * x * x;
w = tanhf(w * SQRT_2_DIV_PI);
out[i] = 0.5f * x * (1.0f + w);
}
}
"""
N = 2048
x = npu.input("x", shape=(1, N), dtype="f32")
y = npu.custom(x, source=gelu_c, carrier="Abs", _name="y")
program = npu.compile(npu.Graph(inputs=[x], outputs=[y], name="gelu_f32_example"))
input_value = np.linspace(-4, 4, N, dtype=np.float32).reshape(1, -1)
output = program.run({"x": input_value})["y"]
reference = 0.5 * input_value * (1.0 + np.tanh(np.sqrt(2.0/np.pi) * (input_value + 0.044715 * input_value**3)))
print(f"maximum absolute error: {np.max(np.abs(output - reference)):g}")
Running the script on a Windows x64 machine with a Meteor Lake CPU and NPU3720 reports a tiny maximum absolute error, confirming functional correctness.
Supported features (as of the latest release)
- Custom kernel compilation: C → ACT‑SHAVE machine code via MoviTools.
- Graph integration: Custom nodes can coexist with Intel‑provided ops.
- Data types: Static dense FP16 unary/binary kernels and a verified FP32 unary path.
- Mixed‑precision graphs: One graph can contain independent FP32‑unary and FP16‑binary custom branches.
- Math library: Non‑linear functions such as
tanhfare available through the bundledmlibm.asymbol inventory. - APIs: Python, CLI, and native C interfaces are provided.
Current limitations
- Platform: Windows x64 only; Linux support is untested.
- Hardware: Verified exclusively on Meteor Lake / Intel NPU3720. Other generations have not been validated.
- Static shapes: Only static tensor shapes are supported; dynamic shapes are not yet handled.
- ACT carriers: Only compatible ACT carriers (e.g.,
Abs) can host custom kernels. - Mixed‑precision conversion groups: Not automatically discoverable; the mixed‑precision example uses separate branches.
"Support is experimental and currently limited to Windows x64, Meteor Lake / NPU3720, static shapes, compatible ACT carriers, and known tensor layouts." – npunlock README
Getting started
Prerequisites
- Hardware: Windows x64 machine with a Meteor Lake CPU and Intel NPU3720.
- Drivers: Install the official Intel NPU driver for the device.
- Toolchain: Python 3.10+, CMake 3.24+, MSVC toolchain, and the MoviTools
MVC_DEPENDpackage (extracted from a legacy Lenovo driver pack, do not install the driver itself).
Installation steps
# Clone the repo and install the Python package
git clone https://github.com/hsfzxjy/npunlock.git
cd npunlock
python -m pip install .
The package bundles npunlock.dll and npunlock_worker.exe, so no additional native path configuration is required beyond pointing to MVC_DEPEND.
Running the GELU example
$env:NPUNLOCK_MOVITOOLS_DIR = 'C:\path\to\MVC_DEPEND'
python examples\example_gelu_f32.py
The script compiles the custom kernel, inserts it into a graph, executes on the NPU, and prints the maximum absolute error against NumPy.
Community feedback from Hacker News
- @ur-whale noted the Windows‑only requirement as a barrier.
- @Gigachad asked about practical use‑cases; the project’s author responded that it enables "bare‑metal NPU programming" for inference tasks Intel does not officially support.
- @Bayard_ne highlighted the excitement of having direct access to custom, unconventional inference workloads.
- @alex7o suggested extending the approach to Qualcomm Hexagon, indicating interest in broader applicability.
These comments underscore both enthusiasm for low‑level NPU hacking and the desire for cross‑platform support.
Contributing to Linux support and newer NPUs
The repository invites contributions for:
- Linux port – testing whether a Windows‑generated SHAVE image runs unchanged on Linux, and building a Linux‑compatible MoviTools wrapper.
- New hardware – checking if the existing
3720xxSHAVE image works on later Intel NPUs or if OEM driver packages provide matching toolchains.
Both efforts require hardware validation, driver/firmware version tracking, and numerical comparison against a host oracle. Detailed guidance is available in the Porting to Linux and newer NPUs wiki page.
Documentation overview
- Getting MoviTools – how to obtain the compiler without installing the legacy driver.
- Python API – constructing, compiling, and executing graphs.
- Writing custom kernels – entry‑point conventions, tensor handling, and example kernels.
- How npunlock works – internals of graph compilation and kernel injection.
- Reverse‑engineering breakthroughs – the experiments that made custom kernels possible.
- Current limitations – full compatibility matrix.
- Development and native APIs – build system, testing, and C interfaces.
All documentation resides in the repository’s wiki folder and is linked from the README.
License
npunlock is released under the Apache License 2.0. Proprietary dependencies such as MoviTools and Intel/Movidius libraries are not redistributed and remain under their original licenses.
Sources
Related
- Project
- Dispatch
- Project
- Project
- Project