Skip to content

Fix SimdRadixN losing the backend's target feature in the cross-FFT layers - #194

Merged
ejmahler merged 1 commit into
ejmahler:masterfrom
HEnquist:wasm_simd_cross_layer_feature
Sep 27, 2026
Merged

ejmahler merged 1 commit into
ejmahler:masterfrom
HEnquist:wasm_simd_cross_layer_feature

Conversation

@HEnquist

@HEnquist HEnquist commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Replaces the closure-passing cross_layer_chunks calls with a single cross_layer on
SimdVector, so the butterfly closure is created inside the backend's
#[target_feature] function and inherits it. A closure only inherits target features
from the function it is written in, and the chunk closure fft_helper_* takes is
written outside one, so on wasm every intrinsic in a layer became an out-of-line call.

  • 174 out-of-line intrinsic calls per RadixN monomorphization on wasm, now 0
  • 65536 point all-radix-4 f64 under wasmtime: 5.51ms -> 3.83ms
  • fcma gets the same attribute, since neon,fcma is not baseline either
  • neon passes #[inline(always)] instead, and is unchanged
  • simd-radix-4 needs the same two-line change in simd_radix4.rs and
    simd_radix4_otf.rs, which takes RadixN there from 12x slower than the old WASM
    Radix4 to slightly faster

Fixes #186

On a target where the backend's instruction set is not in the baseline, wasm's
simd128 being the case that showed up, a function only gets SIMD instructions if
it carries #[target_feature] itself or is inlined into one that does. SimdRadixN
had the attribute only on the fft_helper_* wrappers and relied on #[inline(always)]
chains to pull the rest inside.

That does not hold, because a closure inherits target features from the function
it is written in, and nothing else propagates them. The chunk closure the
fft_helper_* boundary takes is written in SimdRadixN's Fft methods, which carry no
attribute, so the closure and cross_ffts below it compiled without simd128 and
every intrinsic in the layer became an out-of-line call. Adding #[inline(always)]
to cross_ffts makes it worse, since it just clones those calls into each closure.

Add a single cross_layer method to SimdVector, implemented for every backend by
one macro that takes the attribute the backend needs. The per-radix match moves
out of cross_ffts and into that macro, so the butterfly closures are written
inside the target feature function and the whole layer lands inside it whatever
the inliner decides, with one non-inlinable boundary per layer rather than per
intrinsic. The match cannot live in shared code instead, since the closures
written there would again be outside the attribute.

Measured on wasm32-wasip1 under wasmtime, the RadixN path goes from 174
out-of-line intrinsic calls per monomorphization to zero, and an all-radix-4
65536 point f64 FFT drops from 5.51ms to 3.83ms. On ejmahler's simd-radix-4
branch, where more of the loop had been outlined, the same fix takes RadixN from
12x slower than the old WASM Radix4 to slightly faster.

fcma gets the same treatment, since neon,fcma is not baseline either. neon is
baseline on aarch64 so it passes #[inline(always)] instead and keeps the layer
inlined, with timings unchanged.

Fixes ejmahler#186
@ejmahler

ejmahler commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

I confirmed that this improves the performance of SimdRadixN on wasm simd. To be honest I am not a fan of this solution because it destroys the encapsulation of radixn.

However, I'm going to merge anyways because this solution will unblock benchmarking work, and we can investigate alternative approaches at our leisure. I have an idea for an alternative approach, which I'll submit after the radix4 and trait merging tasks are done.

@ejmahler
ejmahler merged commit 6dc993c into ejmahler:master Sep 27, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

wasm_simd: large performance regression in SimdRadixN vs the old wasm Radix4

2 participants