Skip to content

Add byte compress and expand operations - #395

Merged
Shnatsel merged 10 commits into
linebender:mainfrom
Novum:compress-expand
Oct 3, 2026
Merged

Shnatsel merged 10 commits into
linebender:mainfrom
Novum:compress-expand

Conversation

@Novum

@Novum Novum commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Another try at this, only compress/expand this time. Totally fine if you don't want to take it, but I really didn't want to write scalar fallbacks for every single one of the situations I use those.

In #349 the argument was that these only exist on AVX-512. But you can do them pretty fast without it too, with a lookup table and PSHUFB. That's what Highway does as well (docs, implementation). They call it "potentially slow" for bytes, but it's still ~4x faster than scalar, see below.

This adds compress and expand (plus _merge variants) for u8x16, u8x32 and u8x64.

  • AVX-512: native vpcompressb/vpexpandb
  • SSE4.2/AVX2: lookup table + PSHUFB per 128-bit block, then the blocks get spliced together
  • Neon/Wasm: same tables with swizzle_dyn_precise
  • SSE2/fallback: scalar, branchless

Benchmark on a Ryzen 9 9900X (locked at 4.4Ghz, no SMT), 1 MiB of random bytes, 50% of lanes selected. Compared against a branchless scalar loop, which is what you would write yourself as a fallback. Columns are the vector type used.

compress_expand_bench

compress, GB/s:

u8x16 u8x32 u8x64
Scalar loop 3.3 3.3 3.3
SSE4.2 12.2 4.6 7.1
AVX2 13.0 13.1 9.5
AVX-512 31.5 55.5 56.7

The wider types on SSE4.2 are slower than u8x16 because of the splicing. Could probably be improved, but it's still faster than scalar.

Comment thread fearless_simd_gen/src/ops.rs Outdated
Comment on lines +766 to +772
Op::new(
"load_expand",
OpKind::AssociatedOnly,
OpSig::LoadExpand { merge: false },
"Load consecutive bytes from `{arg0}` into the lanes selected by `{arg1}`.\n\n\
The first selected-lane-count bytes are consumed. Unselected lanes are zero.",
),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I admit I am having trouble understanding what this does, and why the load variant is needed as opposed to using a regular load into a vector followed by expand().

Could you provide a code example, or elaborate the description?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's the same thing as a load + expand(). The only difference is that AVX-512 has a native expand-from-memory instruction (vpexpandb with a memory operand) and this maps to it. Highway has LoadExpand for the same reason, on everything else it's also just Expand(LoadU(...)).

So it's really there for a little bit of extra performance on AVX-512.

@Shnatsel

Copy link
Copy Markdown
Contributor

Thank you for the PR! I agree the performance numbers are compelling. SSE2 is not really an optimization target due to being nearly extinct in the wild, and all the other x86 targets show significant improvements.

I am not familiar with this class of operations, and I am having some difficulty understanding what they do based on the docs in this PR. Could you add code examples to all the ops demonstrating what they do? I know we haven't been very disciplined about doing this, but I think in this case is necessary.

@Novum

Novum commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

compress packs the lanes where the mask is set to the front, expand does the reverse:

mask = lanes 0, 2, 5
compress([10, 11, 12, 13, 14, 15, ...]) = [10, 12, 15, 0, 0, 0, ...]
expand([10, 12, 15, 0, 0, 0, ...])      = [10, 0, 12, 0, 0, 15, ...]

Compress is basically for anything where you filter stuff. You store the whole vector and advance the output by the number of set bits. Removing whitespace, filtering a column, getting the indices of matches, quicksort partitioning (that's how vqsort does it).

Expand is the other way around. You have packed values and a bitmap of where they go, e.g. optional values that are stored without the empty slots, or decompression.

I use both in a compressor I'm working on. Compress to sort symbols by code length for Huffman, load_expand to merge a packed byte stream with a constant.

@Shnatsel

Copy link
Copy Markdown
Contributor

I messed around with load_expand and it doesn't seem to affect performance at all on my machine compared to load + expand.

The AVX-512 vpexpandb actually has different semantics than load+expand: it performs a masked load, so only the number of bytes indicated by the mask are accessed. Sadly we cannot make use of this in the portable API, so we don't benefit from this form.

Since there is no benefit to load_expand, I'd rather not include it, since it clutters the public API and adds unsafe that we are trying to minimize.

@Novum Novum changed the title Add byte compress, expand and load_expand Add byte compress and expand operations Sep 27, 2026
@Novum

Novum commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

I removed the load variants.

@Shnatsel

Copy link
Copy Markdown
Contributor

Thank you!

I believe we can implement this for all vector widths by reusing the byte implementation: AVX-512 has dedicated instructions for it, and on all other platforms the masks are either all zeroes or all ones, so we can just route all other widths though the u8 implementation and it will still be correct.

I can implement that myself. Are you OK with me pushing into this branch directly?

@Novum

Novum commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Yes, please

^ Conflicts:
^	fearless_simd_gen/src/mk_wasm.rs
@Shnatsel

Shnatsel commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

I've expanded the op coverage to all widths, made it accessible generically through SimdBase trait, and added optimizations: a generic one for 16 bits and up on 128-bit vectors for 1.3x to 1.8x improvement, and a specialized AVX2 one for 32 bits and up on 256-bit and 512-bit vectors that makes use of the native 256-bit dword shuffle instruction for a 3x improvement.

The only reason I'm not scared of this code is the addition of exhaustive tests.

@Novum I'd appreciate if you could take a look and let me know if this looks good to you.

# Conflicts:
#	fearless_simd/src/generated/avx2.rs
#	fearless_simd/src/generated/sse2.rs
#	fearless_simd/src/generated/sse4_2.rs
#	fearless_simd/src/generated/wasm.rs
#	fearless_simd_gen/src/mk_x86.rs

@Shnatsel Shnatsel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've resolved the conflicts with main and I think it's good to go, but I'd appreciate a look over my work from someone who isn't me.

# Conflicts:
#	fearless_simd_gen/src/mk_neon.rs
#	fearless_simd_gen/src/mk_x86.rs
@Shnatsel

Shnatsel commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

It is possible to have 2.5x higher throughput for u32x8 compression specifically by avoiding the stack buffer and instead reusing the existing tables for a shuffle that joins the halves. But I don't feel special-casing one type on one ISA is worth the additional complexity.

@Novum

Novum commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

I looked through your changes and they make sense. I think this is good to merge.

@Shnatsel

Shnatsel commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Thank you!

@Shnatsel
Shnatsel added this pull request to the merge queue Oct 3, 2026
Merged via the queue into linebender:main with commit 9dd8537 Oct 3, 2026
23 checks passed
@Novum

Novum commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor Author

Excellent, thank you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants