Go 1.27 shipped a portable SIMD package on August 19, and on September 24 the Go team published the design behind it. The package is called simd, it hides behind GOEXPERIMENT=simd, and it lets you write one vector loop that compiles to AVX2 on your x86 servers and to NEON on your Graviton boxes. The same loop builds to 128-bit wasm SIMD if you target the browser. The blog post is by David Chase and Junyang Shao, and it drew 371 points on Hacker News by Friday.

For as long as Go has existed, its users have reached for assembly whenever a hot loop needed to go faster. This is the first credible attempt to end that.

What shipped and where it came from

The story starts in Go 1.26, back in February. That release added simd/archsimd, a set of intrinsics for amd64 only, with fixed types like Int8x16 and Float64x8 that map straight onto AVX, AVX2 and AVX-512 registers. You wrote code against a specific vector width and the compiler emitted the instructions. Useful, but you had to write it three times if you cared about three widths, and it did nothing on ARM.

Go 1.27 did two things. It revised the amd64 archsimd API and added 128-bit support for arm64 NEON and WebAssembly. Then it added the second layer: a package named plain simd with types like Int8s, Uint32s and Float32s. Note the plural. Those types have no fixed length. You call Len() on one and get whatever the target hardware gives you, from 16 bytes on NEON up to 64 bytes on an AVX-512 machine.

The release notes describe the portable package as supporting "a scalable subset of the operations present in the simd/archsimd package that are hardware-supported or easily emulated across architectures and vector widths." That sentence is the whole design. The Go team took the intersection of what amd64, NEON and wasm can all do, made that the API, and wrote emulation for the small gaps. Unsigned compares become a signed compare plus an XOR. Scalar shifts become vector shifts. Carryless multiply, which crypto code wants, gets a constant time fallback where the hardware lacks it.

The API is, in the authors' words, "loosely based on Highway," the C++ library Google uses for the same problem. Jan Wassenberg, who leads Highway, showed up in the Hacker News thread to say the Highway team "shared some advice on the API." That lineage matters because Highway already handles SVE and RISC-V vectors, where the register width is unknown until runtime. Go picked a design that has a path to those.

How the compiler makes it fast

A portable vector type sounds like it should cost you a dispatch on every operation. It doesn't, because the compiler cheats in a sensible way.

Any function that touches a simd type gets cloned at compile time into specialized copies. The post names them with suffixes: @simd128, @simd256, @simd512, and @simd0 for pure emulation. Each copy is compiled as if the vector width were fixed, so inside the loop there is no branching on hardware. The dispatch happens once, at the entry to the SIMD code, based on what the CPU reports. If you've ever seen how klauspost/compress or the standard library's hash packages pick an implementation at init time, this is that pattern, done for you by the toolchain.

The example in the post is an inner product:

func innerProduct(x, y []float32) float32 {
    var a simd.Float32s
    var i int
    for i = 0; i < len(x)-a.Len()+1; i += a.Len() {
        u := simd.LoadFloat32s(x[i : i+a.Len()])
        v := simd.LoadFloat32s(y[i : i+a.Len()])
        a = u.MulAdd(v, a)
    }
    if i < len(x) {
        u, _ := simd.LoadFloat32sPart(x[i:])
        v, _ := simd.LoadFloat32sPart(y[i:])
        a = u.MulAdd(v, a)
    }
    return sum(a)
}

Two things to notice. The tail is handled with a partial load, so you don't write the usual scalar cleanup loop. And that final sum(a) is a function you have to write yourself, because the portable package has no horizontal reduction yet. The authors say ReduceSum is planned for Go 1.28. If your first SIMD job is a dot product, and for anyone doing embedding search in Go it will be, you'll be writing that helper for the next six months.

There's an escape hatch. Every portable type has a ToArch() method that returns the underlying archsimd value, and a matching FromArch to go back. The post uses this to implement a population count that the portable API lacks. On amd64 it's a type switch over three widths with a lookup table, and on arm64 and wasm it's a direct instruction. Everything else gets a scalar fallback. You write the portable loop once and drop to the metal only where the intersection fails you.

For testing, there are GODEBUG knobs. GODEBUG=simd=0 forces emulation everywhere. GODEBUG=simd=128 pins a width and panics if the hardware can't do it. That second one is how you check, on a developer laptop with AVX-512, that your code still works on the NEON box in production.

The cost of portability

Nobody should expect the portable layer to match hand tuned intrinsics, and the early numbers say it doesn't. One commenter, ImJasonH, posted a palette swap benchmark in the Hacker News thread: the portable version ran about 11% slower than the archsimd version, and both were roughly 5x faster than the scalar loop. Another commenter reported a 30% speedup on a foreground estimation task after switching. Those are anecdotes, not a benchmark suite, but the ratio is what I'd expect from a Highway style design. You give up a tenth to avoid maintaining three copies.

The sharper criticism in the thread was binary size. Cloning every SIMD function into four variants is fine for a dot product. For a codebase with hundreds of vectorized functions it could add up, and the post doesn't publish numbers on that. I'd want to see the size of a real compression library built this way before I believed it was free.

The other cost is that the flag is still there. Cherry Mui filed a proposal on April 27 to turn archsimd on by default for amd64, arguing that user feedback on the experiment had been "very positive" and that no major flaws had turned up. That proposal is on hold. The portable package is newer and less baked. So today, if you want any of this, every build of your project needs GOEXPERIMENT=simd set, which means every CI config, every Dockerfile, every Makefile, and every teammate's shell. Modules published to the proxy that depend on it will fail to build for anyone who forgot. That is a hard line for a library author, and it's the reason most Go library authors will keep their assembly files for now.

What I'd do with it

If you run Go services, the obvious win is anywhere you currently ship an _amd64.s file next to a _generic.go fallback. Hashing, checksums, base64, JSON scanning, string search. Those files were written for one architecture, and when the fleet started mixing in Graviton, the ARM half of the fleet fell back to the generic path. The portable package removes that asymmetry. One loop, both architectures, and wasm for free if you ever compile to it.

The second place is the vector math that's crept into ordinary services over the last two years. Cosine similarity over float32 embeddings is a dot product and a couple of norms. Most Go shops doing this either call out to a C library or accept a scalar loop. With this package it's twenty lines of plain Go, and the 5x from the benchmark above is the difference between a similarity check being free and it showing up in the profile.

What I wouldn't do is put it in a library you publish. The experiment flag makes it a private tool until it lands by default. Use it in your own binaries where you control the build, measure it against your current assembly, and read the SVE roadmap before you assume the widths you see today are the widths you'll see on the next generation of ARM servers.

My prediction: archsimd on amd64 goes default in Go 1.28 next February, the portable package stays behind the flag until 1.29, and by the end of 2027 the assembly stubs in the standard library's crypto packages start getting replaced with it. If that last one happens, the Go team will have done for vector code what it did for the garbage collector: made the default good enough that most people stop thinking about it.