Hot take: there is no portable SIMD.<p>You can either have performance (=write manual ASM for each platform), or portability, but not both.<p>What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
You've got a point but are overstating it considerably. There is a big gap between <i>just</i> autovectorization and the portable primitives a library like Highway or Fearless SIMD will give you. For example, I haven't seen autovectorization do select or swizzle.<p>But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.<p>So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.<p>Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.<p>As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.
Getting 2x or 4x performance in your inner loops using a reasonable SIMD library is infinitely better than theoretically getting 8x performance with hand-coded nonportable intrinsics, because the latter is never going to happen in most programs, so the actual point of comparison is scalar code, or autovectorized code at best.
No it's not because it sucks the air out from the actual solution. ISPC more than a decade ago managed to demonstrate close-to-linear speedups for increasing vector sizes, even for branchy code.<p>Nowadays you can even get AI to write intristics and it works just fine, the portable libraries/autovec aren't really a serious player here.<p>Portability is also overstated - see the recent shift where Spotify decided to make native Android/iOS apps again instead of React Native. Usually, the number of relevant platforms is somewhere between 2 and 3, so portability concerns are more theoretical than real.
ISPC gets pretty close, though. <a href="https://ispc.github.io/" rel="nofollow">https://ispc.github.io/</a>
I agree with that historically auto-vertorization does not seem to work reliably. I'm not sure about your broad claim.<p>Thoughts on an abstraction over ARM and x86, at 128, 256, and 512-bit widths which, either in a manual or automatic way (The latter more challenging) makes your floating point computations 4-16x faster with minimal restructuring? I think that's doable, and a nice goal of SIMD.
> there is no portable SIMD<p>Except in languages with a JIT compiler
That still just gets you autovectorization, and generally locks you out of the performance you could have with direct SIMD intrinsics.<p>Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
> That still just gets you autovectorization<p>You'll have to explain to me why that's a problem? Because, I don't quite understand why (but it is early on monday morning, so perhaps the coffee hasn't fully kicked in yet). The code generated (by the JIT compiler) is raw SIMD instructions that are supported by the CPU it is running on and it will spot patterns of usage to optimise those too. There's no <i>runtime</i> check to see what is supported. Maybe I'm missing something?
Which JIT languages use much simd in generated code, in general?
Starts to get a bit philosophical on what constitutes "portable" but JIT compilers would emit an opcode based off of whatever the frontend/IR is saying to do surely?
That assumes no per-platform optimisation, which most JIT compilers will do. I can only speak to the dotnet platform as that's what I am using the most at the moment, but its JIT will spot certain functions, like Vector512.Load and replace it with the the SIMD equivalent. So, it's not just IR op-code to CPU op-code, it's looking for <i>patterns-of-code</i> that can be made more efficient at JIT compile-time.
Sure, in the same way that there is no such thing as portable code at all. The result will be suboptimal, but it will still be better than not having it.
This is going to be a real problem in the future when x86 and ARM, SIMD moves ahead and becomes wider.
It's a continuum. Some things basically all SIMD implementations support. Want to add 2 4xf32 vectors together? That's pretty easy to do portably.<p>But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.<p>Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
What about numpy, numba, and torch.compile?