Go 1.27 introduces experimental platform‑independent SIMD package
TL;DR
Go 1.27 ships an experimental, platform‑independent simd package that abstracts away fixed‑size vector types, enabling write‑once vector code that runs at near‑assembly performance on AVX/AVX2/AVX512 (amd64), NEON (arm64), and WebAssembly, while falling back to a fast software emulation on CPUs without SIMD support.
Why a portable SIMD layer?
Different CPU families expose SIMD in wildly divergent ways:
- Vector width can be fixed (e.g., 128 bits on wasm, PowerPC, s390x) or variable (RISC‑V V‑extension, ARM SVE).
- Masking semantics differ: some architectures have dedicated mask registers (AVX‑512, RVV), others use vector booleans (NEON, AVX2), and some lack masks entirely.
- Instruction sets vary not only by family but also by feature flags (AVX vs AVX2 vs AVX‑512, NEON vs SVE, etc.).
Because Go’s earlier archsimd package had to expose these quirks directly, writing portable SIMD code required per‑architecture conditionals and hand‑rolled assembly. The new simd package removes fixed‑size vectors from the type system and implements only the intersection of operations that all supported platforms can perform, filling the gaps with efficient emulation.
How to enable the experimental API
Set the environment variable GOEXPERIMENT=simd when building or running your program, just like the older archsimd experiment. The compiler rewrites calls to simd types into architecture‑specific implementations or emulated fallbacks based on the detected hardware at program start.
Core design principles
- Intersection‑first API – Only operations common to all target platforms are exposed. Missing primitives are provided via emulation that maps to existing SIMD instructions.
- Near‑assembly performance – When the source operation matches the hardware capability, the generated code is as fast as hand‑written assembly.
- Graceful fallback – On CPUs without SIMD support (or when
GODEBUG=simd=0is set), the same code runs using a pure‑Go implementation. - Readability – Types are named
simd.Uint8s,simd.Float32s, etc., and methods read like ordinary Go methods, making the code approachable for humans and large language models alike.
Example: Vectorized inner product
func innerProduct(x, y []float32) float32 {
var acc simd.Float32s
var i int
for i = 0; i < len(x)-acc.Len()+1; i += acc.Len() {
u := simd.LoadFloat32s(x[i : i+acc.Len()])
v := simd.LoadFloat32s(y[i : i+acc.Len()])
acc = u.MulAdd(v, acc) // acc += u*v
}
if i < len(x) {
u, _ := simd.LoadFloat32sPart(x[i:])
v, _ := simd.LoadFloat32sPart(y[i:])
acc = u.MulAdd(v, acc)
}
return simd.ReduceSum(acc) // added in Go 1.28
}
The loop processes acc.Len() elements per iteration, where acc.Len() is the runtime‑detected vector width (128, 256, or 512 bits). The same source compiles to AVX‑512 on a modern x86, NEON on Apple Silicon, or a pure‑Go loop when SIMD is unavailable.
Supported operations (Go 1.27)
The table below summarizes the methods that are currently available. V denotes a vector type, M a mask type, and E a scalar element type.
Load / Broadcast
LoadV([]E) V– load a full slice into a vector.LoadVPart([]E) (V, int)– load a partial slice, returning the number of elements actually loaded.BroadcastV(E) V– broadcast a scalar to all lanes.
Store / String
Store([]E)– write a vector back to a slice.StorePart([]E) int– store as many lanes as fit.String() string– human‑readable representation.
Arithmetic (selected)
Add,Sub,Mul,Div(float only),Neg,Abs(float only).MulAdd– fused multiply‑add for floats.IfElse(mask, y)– select elements fromywheremaskis true.Masked(mask)– zero out lanes where mask is false.
Boolean & Masking
And,Or,Xor,AndNoton vectors.Masktypes (Mask8s,Mask16s, …) supportAnd,Or,String,ToIntWs.
Comparisons
Equal,NotEqual,Greater,GreaterEqual,Less,LessEqual– produce a mask of the same element width.
Conversions & Reshaping
ConvertToFloatW,ConvertToIntW,ConvertToUintW.ToBits,ReshapeToUint*,BitsToFloatW,BitsToIntW– zero‑cost reinterpretations.
Shifts & Rotates (integer vectors only)
ShiftAllLeft/Right,RotateAllLeft/Rightfor 16‑, 32‑, and 64‑bit lanes.
Note: The first release lacks a generic reduction (
ReduceSum) and some mask‑centric utilities; these are slated for Go 1.28.
Bridging to architecture‑specific code
When a needed operation is not yet in the portable API, developers can fall back to the archsimd package via ToArch() and FromArch helpers. Example for a missing Int8s.OnesCount implementation on amd64:
//go:build goexperiment.simd && amd64
func OnesCount(v simd.Int8s) simd.Int8s {
switch x := v.ToArch().(type) {
case archsimd.Int8x16:
// Use a lookup‑table trick for AVX/AVX2.
lut := archsimd.LoadInt8x16Array(&popcnt4x16)
mask := archsimd.BroadcastInt8x16(0x0f)
lo := x.And(mask)
hi := x.ToBits().ReshapeToUint16s().ShiftAllRight(4).
ReshapeToUint8s().BitsToInt8().And(mask)
return simd.Int8sFromArch(lut.PermuteOrZero(lo).Add(lut.PermuteOrZero(hi)))
default:
return OnesCountEmulated(v)
}
}
The compiler specializes the type switch away because the binary knows at compile time which archsimd types exist for the target platform.
Runtime control with GODEBUG
Developers can force a particular vector width or full emulation for testing:
GODEBUG=simd=0– always use the emulated path.GODEBUG=simd=128,256,512– require that width; panic if unavailable.GODEBUG=simd=+128(or+256,+512) – allow the width even if some optional instructions are missing; using an unsupported instruction will panic at runtime.
Implementation notes
The compiler performs an AST rewrite that generates size‑specialized copies of every function that mentions a simd type. These copies live in simd/internal/bridge and are instantiated as concrete archsimd types (e.g., Int8x16). Dispatch occurs once at program start, after which the hot loops execute the specialized version without further branching. Adding a dummy reference to a simd type (e.g., var _ simd.Uint64s) can raise the dispatch point to a surrounding loop, ensuring the specialized version is used.
Community reaction (selected HN comments)
"Portable SIMD is ~11 % slower than non‑portable SIMD in this case, but both are ~5× faster than non‑SIMD." – ImJasonH (benchmark on wasm)
"C++ is getting
std::simd; I’m all aboard writing vector code with the least amount of intrinsics." – beached_whale
"Nobody was asking for this, but they took the time to do it right and continue to push Go as a memory‑safe, high‑level systems language." – u8
"This opens many doors for low‑level performance in Go projects that already run multicore. Few languages have std‑lib SIMD support." – qprofyeh
"The interface conversion and type switch look inefficient, but the compiler specializes them away; the switch is resolved at start‑up, not per‑iteration." – vlovich123
What’s next?
- Go 1.28 will add SVE support to
archsimdand eventually to the portablesimdpackage. - New operations such as
OnesCount, richer mask utilities, reduction primitives, and vector shuffles are planned. - Feature‑variant flags will let code run on hardware that lacks only a few optional instructions (e.g., Raspberry Pi’s NEON without PMULL).
- A dedicated blog post on
archsimdinternals is forthcoming.
Bottom line
Go’s experimental SIMD layer removes the historical barrier of per‑architecture assembly, offering a clean, size‑agnostic API that delivers near‑native performance on a wide range of CPUs while guaranteeing a functional fallback on any platform. The design mirrors the philosophy of Highway for C++ and anticipates future extensions such as SVE, making Go a more attractive choice for high‑performance, cross‑platform workloads.
Sources
Related
- Project
- Project
- Project
- Project