Skip to content

Support odd f32 bases in SimdRadixN - #195

Open
HEnquist wants to merge 9 commits into
ejmahler:masterfrom
HEnquist:simd_radixn_odd_f32_base
Open

HEnquist wants to merge 9 commits into
ejmahler:masterfrom
HEnquist:simd_radixn_odd_f32_base

Conversation

@HEnquist

@HEnquist HEnquist commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Lets the f32 SimdRadixN use odd base lengths, by ending each cross-FFT layer with a
partial column.

  • SimdVector gains load_partial_lo and store_partial_lo, which map onto the partial
    complex load and store every backend already has. For f64 a vector is one complex
    number, so they forward to the full load and store and are never reached.
  • An inexact column count now finishes with one complex in the low half of a vector, high
    half zeroed, and only that half stored back. A full load would run past the end of the
    buffer on the last row, and a full store would overwrite column 0 of the next row.
  • Twiddle chunk counts round up. The extra chunk's unused entries are ordinary twiddles
    that go unread.
  • Odd bases cost 0 to 25% more per point than even ones on an M1 with NEON f32, depending
    on the recipe. The even path is unchanged: it measures 1.3% slower, but so does the
    unmodified code with a never-executed block of the same size added to cross_layer.
  • SimdRadix4 gets the same treatment, since it shares the cross-layer code with
    SimdRadixN.
  • Tests: factor_pairs and the composite base body now run odd bases on every backend,
    and the large_recipes f32 bases alternate parity so both paths are covered at every
    depth.

Fixes #189

HEnquist and others added 3 commits September 26, 2026 22:11
On a target where the backend's instruction set is not in the baseline, wasm's
simd128 being the case that showed up, a function only gets SIMD instructions if
it carries #[target_feature] itself or is inlined into one that does. SimdRadixN
had the attribute only on the fft_helper_* wrappers and relied on #[inline(always)]
chains to pull the rest inside.

That does not hold, because a closure inherits target features from the function
it is written in, and nothing else propagates them. The chunk closure the
fft_helper_* boundary takes is written in SimdRadixN's Fft methods, which carry no
attribute, so the closure and cross_ffts below it compiled without simd128 and
every intrinsic in the layer became an out-of-line call. Adding #[inline(always)]
to cross_ffts makes it worse, since it just clones those calls into each closure.

Add a single cross_layer method to SimdVector, implemented for every backend by
one macro that takes the attribute the backend needs. The per-radix match moves
out of cross_ffts and into that macro, so the butterfly closures are written
inside the target feature function and the whole layer lands inside it whatever
the inliner decides, with one non-inlinable boundary per layer rather than per
intrinsic. The match cannot live in shared code instead, since the closures
written there would again be outside the attribute.

Measured on wasm32-wasip1 under wasmtime, the RadixN path goes from 174
out-of-line intrinsic calls per monomorphization to zero, and an all-radix-4
65536 point f64 FFT drops from 5.51ms to 3.83ms. On ejmahler's simd-radix-4
branch, where more of the loop had been outlined, the same fix takes RadixN from
12x slower than the old WASM Radix4 to slightly faster.

fcma gets the same treatment, since neon,fcma is not baseline either. neon is
baseline on aarch64 so it passes #[inline(always)] instead and keeps the layer
inlined, with timings unchanged.

Fixes ejmahler#186
SimdRadixN::new asserted base_len % COMPLEX_PER_VECTOR == 0, which restricted f32
to even bases. The column count starts at base_len and only grows, so an odd base
means an odd count at every layer, and no factor reordering helps: the first layer
is the one with the fewest columns.

Give each layer a partial column instead. A single complex number is loaded into
the low half of a vector, the same column butterfly runs with the high half zeroed,
and only the low half is stored back. The butterflies work down the columns and the
halves never interact, so the result is the same one a full vector would produce.
Neither a full load nor a full store works there: the load would run past the end
of the buffer on the last row, and the store would overwrite the first column of
the next row. SimdVector gains load_partial_lo and store_partial_lo for this, which
every backend already had as load_partial_lo_complex and store_partial_lo_complex.
For f64 a vector is one complex number, so they forward to the full load and store
and nothing ever reaches the partial path.

Twiddle chunk counts round up, since the partial column needs a chunk of its own.
Only its first twiddle is used, and the rest are ordinary twiddles that go unread.

Measured on an M1, NEON f32, best of five, in ns per point:

  even base 4, 6*6*6:   1.61      odd base 3, 6*6*6:   1.89
  even base 6, 6*6*4:   1.94      odd base 5, 6*6*4:   1.90
  even base 2, 6*5*5:   1.69      odd base 1, 6*5*5:   2.13
  even base 4, 6*6*6*4: 1.90      odd base 3, 6*6*6*4: 2.12

The even path measures 1.3% slower than before, but so does the unmodified code
with a never-executed block of the same size added to cross_layer, so that is the
inlined layer's sensitivity to code size and not work the even path now does. The
wasm IR check is unchanged at 0 out-of-line intrinsic calls added.

Tests: every backend now runs the factor_pairs bodies over odd bases as well, the
composite base body runs its odd bases for f32 too, and the large_recipes f32 bases
alternate between even and odd so both paths are covered at every depth.

Fixes ejmahler#189
@ejmahler

Copy link
Copy Markdown
Owner

Merged master in now that 194 is merged, to make this easier to review

@ejmahler

Copy link
Copy Markdown
Owner

Looks good to me. Disappointing to hear that there's so much of a performance penalty, but not quite surprising. I'm assuming this is marked as draft because it was waiting for 194, but I'll hold off on merging until you mark it ready just in case.

@ejmahler

Copy link
Copy Markdown
Owner

I merged in simd radix4. Since that reuses some simd radixn code, I pushed a change here that makes the same updates to simd radix4.

@HEnquist
HEnquist marked this pull request as ready for review September 28, 2026 21:47
The merged SimdVector trait from ejmahler#197 already has load1_lo_complex and
store1_lo_complex, with load1_lo and store1_lo wrappers in simd_array.rs,
so the partial column tail uses those and this branch no longer adds any
trait methods or backend impls of its own.

The cross layer tail now runs after ejmahler#197's unroll / no-unroll split, so
both paths get it.
SimdRadixN's partial column path is the first caller of load1_lo and
store1_lo that is generic over the vector type, so say at the trait
declaration that only the multi-complex vectors implement them and that
a generic caller has to stay away unless COMPLEX_PER_VECTOR > 1.

That also makes the load1_lo and store1_lo dead_code allows stale, so
drop those two. load1_dup still has no caller and keeps its own.
- load1_dup_complex says what it actually is on a one complex vector,
  a plain load that nothing needs, instead of claiming it's impossible
- the scalar Radix4 and RadixN tests carried the SIMD backends' comment
  about f32 needing an even base, which doesn't apply to a scalar vector
- simd/mod.rs still said SimdRadixN isn't used by the planner, which
  stopped being true when the scalar algorithms became adapters, so the
  comment and both dead_code allows are gone
- the SimdRadix4 and SimdRadixN module docs name the scalar adapter
  alongside f64 as the one complex per vector case
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support odd f32 bases in SimdRadixN

2 participants