Fix SimdRadixN losing the backend's target feature in the cross-FFT layers - #194
Merged
Merged
Conversation
On a target where the backend's instruction set is not in the baseline, wasm's simd128 being the case that showed up, a function only gets SIMD instructions if it carries #[target_feature] itself or is inlined into one that does. SimdRadixN had the attribute only on the fft_helper_* wrappers and relied on #[inline(always)] chains to pull the rest inside. That does not hold, because a closure inherits target features from the function it is written in, and nothing else propagates them. The chunk closure the fft_helper_* boundary takes is written in SimdRadixN's Fft methods, which carry no attribute, so the closure and cross_ffts below it compiled without simd128 and every intrinsic in the layer became an out-of-line call. Adding #[inline(always)] to cross_ffts makes it worse, since it just clones those calls into each closure. Add a single cross_layer method to SimdVector, implemented for every backend by one macro that takes the attribute the backend needs. The per-radix match moves out of cross_ffts and into that macro, so the butterfly closures are written inside the target feature function and the whole layer lands inside it whatever the inliner decides, with one non-inlinable boundary per layer rather than per intrinsic. The match cannot live in shared code instead, since the closures written there would again be outside the attribute. Measured on wasm32-wasip1 under wasmtime, the RadixN path goes from 174 out-of-line intrinsic calls per monomorphization to zero, and an all-radix-4 65536 point f64 FFT drops from 5.51ms to 3.83ms. On ejmahler's simd-radix-4 branch, where more of the loop had been outlined, the same fix takes RadixN from 12x slower than the old WASM Radix4 to slightly faster. fcma gets the same treatment, since neon,fcma is not baseline either. neon is baseline on aarch64 so it passes #[inline(always)] instead and keeps the layer inlined, with timings unchanged. Fixes ejmahler#186
HEnquist
force-pushed
the
wasm_simd_cross_layer_feature
branch
from
September 26, 2026 20:11
659e8d0 to
eaec3b8
Compare
Owner
|
I confirmed that this improves the performance of SimdRadixN on wasm simd. To be honest I am not a fan of this solution because it destroys the encapsulation of radixn. However, I'm going to merge anyways because this solution will unblock benchmarking work, and we can investigate alternative approaches at our leisure. I have an idea for an alternative approach, which I'll submit after the radix4 and trait merging tasks are done. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces the closure-passing
cross_layer_chunkscalls with a singlecross_layeronSimdVector, so the butterfly closure is created inside the backend's#[target_feature]function and inherits it. A closure only inherits target featuresfrom the function it is written in, and the chunk closure
fft_helper_*takes iswritten outside one, so on wasm every intrinsic in a layer became an out-of-line call.
neon,fcmais not baseline either#[inline(always)]instead, and is unchangedsimd-radix-4needs the same two-line change insimd_radix4.rsandsimd_radix4_otf.rs, which takes RadixN there from 12x slower than the old WASMRadix4 to slightly faster
Fixes #186