Add byte compress and expand operations - #395
Conversation
a0fa987 to
1f68c64
Compare
| Op::new( | ||
| "load_expand", | ||
| OpKind::AssociatedOnly, | ||
| OpSig::LoadExpand { merge: false }, | ||
| "Load consecutive bytes from `{arg0}` into the lanes selected by `{arg1}`.\n\n\ | ||
| The first selected-lane-count bytes are consumed. Unselected lanes are zero.", | ||
| ), |
There was a problem hiding this comment.
I admit I am having trouble understanding what this does, and why the load variant is needed as opposed to using a regular load into a vector followed by expand().
Could you provide a code example, or elaborate the description?
There was a problem hiding this comment.
It's the same thing as a load + expand(). The only difference is that AVX-512 has a native expand-from-memory instruction (vpexpandb with a memory operand) and this maps to it. Highway has LoadExpand for the same reason, on everything else it's also just Expand(LoadU(...)).
So it's really there for a little bit of extra performance on AVX-512.
|
Thank you for the PR! I agree the performance numbers are compelling. SSE2 is not really an optimization target due to being nearly extinct in the wild, and all the other x86 targets show significant improvements. I am not familiar with this class of operations, and I am having some difficulty understanding what they do based on the docs in this PR. Could you add code examples to all the ops demonstrating what they do? I know we haven't been very disciplined about doing this, but I think in this case is necessary. |
|
compress packs the lanes where the mask is set to the front, expand does the reverse: Compress is basically for anything where you filter stuff. You store the whole vector and advance the output by the number of set bits. Removing whitespace, filtering a column, getting the indices of matches, quicksort partitioning (that's how vqsort does it). Expand is the other way around. You have packed values and a bitmap of where they go, e.g. optional values that are stored without the empty slots, or decompression. I use both in a compressor I'm working on. Compress to sort symbols by code length for Huffman, load_expand to merge a packed byte stream with a constant. |
1f68c64 to
d8293e4
Compare
|
I messed around with load_expand and it doesn't seem to affect performance at all on my machine compared to load + expand. The AVX-512 vpexpandb actually has different semantics than load+expand: it performs a masked load, so only the number of bytes indicated by the mask are accessed. Sadly we cannot make use of this in the portable API, so we don't benefit from this form. Since there is no benefit to load_expand, I'd rather not include it, since it clutters the public API and adds |
d8293e4 to
b69e907
Compare
b69e907 to
1b31119
Compare
|
I removed the |
|
Thank you! I believe we can implement this for all vector widths by reusing the byte implementation: AVX-512 has dedicated instructions for it, and on all other platforms the masks are either all zeroes or all ones, so we can just route all other widths though the u8 implementation and it will still be correct. I can implement that myself. Are you OK with me pushing into this branch directly? |
|
Yes, please |
^ Conflicts: ^ fearless_simd_gen/src/mk_wasm.rs
1b31119 to
681b1cd
Compare
…SimdBase vectors can implement compress/expand
… for 1.3x to 1.8x speedup over the byte variant
…values and larger that is 3x faster than the generic byte codepath
…rs, same 3x gains seen there as for the 256-bit wide vectors
|
I've expanded the op coverage to all widths, made it accessible generically through SimdBase trait, and added optimizations: a generic one for 16 bits and up on 128-bit vectors for 1.3x to 1.8x improvement, and a specialized AVX2 one for 32 bits and up on 256-bit and 512-bit vectors that makes use of the native 256-bit dword shuffle instruction for a 3x improvement. The only reason I'm not scared of this code is the addition of exhaustive tests. @Novum I'd appreciate if you could take a look and let me know if this looks good to you. |
# Conflicts: # fearless_simd/src/generated/avx2.rs # fearless_simd/src/generated/sse2.rs # fearless_simd/src/generated/sse4_2.rs # fearless_simd/src/generated/wasm.rs # fearless_simd_gen/src/mk_x86.rs
Shnatsel
left a comment
There was a problem hiding this comment.
I've resolved the conflicts with main and I think it's good to go, but I'd appreciate a look over my work from someone who isn't me.
# Conflicts: # fearless_simd_gen/src/mk_neon.rs # fearless_simd_gen/src/mk_x86.rs
|
It is possible to have 2.5x higher throughput for |
|
I looked through your changes and they make sense. I think this is good to merge. |
|
Thank you! |
|
Excellent, thank you. |
Another try at this, only compress/expand this time. Totally fine if you don't want to take it, but I really didn't want to write scalar fallbacks for every single one of the situations I use those.
In #349 the argument was that these only exist on AVX-512. But you can do them pretty fast without it too, with a lookup table and PSHUFB. That's what Highway does as well (docs, implementation). They call it "potentially slow" for bytes, but it's still ~4x faster than scalar, see below.
This adds
compressandexpand(plus_mergevariants) foru8x16,u8x32andu8x64.vpcompressb/vpexpandbswizzle_dyn_preciseBenchmark on a Ryzen 9 9900X (locked at 4.4Ghz, no SMT), 1 MiB of random bytes, 50% of lanes selected. Compared against a branchless scalar loop, which is what you would write yourself as a fallback. Columns are the vector type used.
compress, GB/s:
The wider types on SSE4.2 are slower than u8x16 because of the splicing. Could probably be improved, but it's still faster than scalar.