Any processor made this century carries more arithmetic hardware than its instruction decoder can feed. The decoder is the bottleneck, and the workaround has a name: single instruction, multiple data. Hand the chip a whole batch of numbers — a vector — and one operation costs about what one operation costs. On recent x86 parts the batches run to 512 bits, which yields, in theory, an 8x speedup on f64 math, 64x on u8. In theory. “In practice,” notes the author of a new survey of the SIMD landscape in Rust, “it can run both slower and faster.”
The survey, published September 28 on the author’s own blog, is the second annual edition, and it arrives with a disclosure: since the first one, its author has become a maintainer of Fearless SIMD, one of the libraries under comparison. The authors of std::simd, wide, pulp and macerator were invited to review the draft; editorial control was retained, and “all mistakes are my own.”
SIMD arrives late to every architecture, so it comes with marketing. ARM calls its version NEON and made it standard on every 64-bit ARM chip. WebAssembly, the survey observes, has no marketing department, and so its extension is called the “WebAssembly 128-bit packed SIMD extension.” Sixty-four-bit x86 shipped with SSE2 and its 128-bit vectors; later came SSE 4.2 with more operations, AVX and AVX2 with 256-bit vectors, AVX-512 with 512-bit. “The word ‘later’ in the above paragraph,” the author writes, “creates a problem.”
The problem is assumption. A compiler building for x86_64 may not assume anything past SSE2, because anything past SSE2 will fail on some machine. If you run your own servers or your own cloud, you can simply declare every processor at least AVX2 — a feature introduced over ten years ago — with RUSTFLAGS=’-C target-cpu=x86-64-v3’, and let the program crash or misbehave anywhere older. If you ship binaries to strangers, you do function multiversioning instead: the same function compiled for several instruction sets, the choice made at runtime after checking what the silicon actually has. ARM mostly escapes this; NEON was made mandatory, and ARM, the survey says, “hasn’t really added useful SIMD extensions after that.” WebAssembly asks for two binaries and a JavaScript capability check.
There are three ways in. Write plain Rust — &[i32].sum() — and let the compiler’s heuristics vectorize it. Write a portable abstraction — i32x4 + i32x4. Or call the platform intrinsics directly, the option whose conditional-compilation headers, the author warns, need a bigger code block.
Automatic vectorization requires no dependencies and reaches every instruction set a compiler knows, however obscure. It is also, the survey says, unreliable: the larger the function, the worse the odds, and the compiler version alone can swing performance. Practitioners iterate over &[i32].as_chunks() and read the assembly to verify. Floats were long beyond the heuristics’ reach, because reordering float arithmetic changes the observable result; Rust 1.98 changed that by stabilizing algebraic operations such as algebraic_add(), “a less dangerous -ffast-math,” though the code must still be rewritten to use them.
For runtime dispatch there is the multiversion crate: annotate a function with #[multiversion(targets = “simd”)] and go. The survey flags an undocumented pitfall — each annotated call carries a little overhead, under a dozen instructions, which looms large when the function itself is small. The rule of thumb: a function with a loop gets #[multiversion]; one processing a handful of values gets #[inline(always)], so long as a #[multiversion] sits somewhere up the call chain. The crate is also the only one that names exact CPU extensions rather than a preset SIMD level, though that rarely matters for autovectorized code. And for AVX-512 it checks presence, not speed, which may hurt performance in practice; the workaround is boilerplate on every function, a target string long enough to end with the note “Ice Lake and later.”
The portable-abstraction crates are graded, table by table, against the wish list: fixed-width vectors of known size, hardware-width vectors, generics over element type, generics over vector width. One entry, macerator, also supports LoongArch, because its author was — the survey says, quoting — “bored.”
std::simd is described not as a complete solution but as the building blocks that belong in the standard library, everything else left to the ecosystem. The largest drawback is that it is nightly-only and still undergoes infrequent breaking API changes: one day you update the compiler and the code stops compiling, and you go and fix it. For anyone at peace with that, needing only fixed-width vectors and perhaps multiversioning, the survey calls it “pretty great.”
On Hacker News, where the survey circulated, the footnotes acquired footnotes. ARM’s baseline is genuinely comprehensive, wrote the commenter ack_complete, but some optional extensions — the crypto extension, the newer dot-product instructions — matter situationally, “and unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting” them. And then the blunt version, from Archit3ch: “Hot take: there is no portable SIMD. You can either have performance (=write manual ASM for each platform), or portability, but not both.” What the portable libraries actually deliver, the comment says, is “portable auto-vectorization.” In theory, and in practice. “The microbenchmarks will look great, though. ;)”

