Skip to content

Update multi-stark for hybrid CUDA scheduling - #612

Draft
arthurpaulino wants to merge 1 commit into
mainfrom
ap/gpu
Draft

Update multi-stark for hybrid CUDA scheduling#612
arthurpaulino wants to merge 1 commit into
mainfrom
ap/gpu

Conversation

@arthurpaulino

Copy link
Copy Markdown
Member

Point Aiur at the multi-stark revision that balances large mixed-height commitments across host CPUs and CUDA according to device capacity. The backend carries hybrid placement through lookup, quotient, openings, and FRI while preserving the existing Goldilocks/BLAKE3 protocol and CPU verification path.

No Ix source or build configuration changes are required: the existing opt-in cuda feature chain selects the new backend, while default builds remain CUDA-independent. Cargo.lock records the matching dependency graph.

Validated with release clippy across the non-CUDA CI feature set, a release all-targets CUDA check, and IX_CUDA=1 lake build bench-typecheck, including the final Rust static-library link.

@arthurpaulino

Copy link
Copy Markdown
Member Author

!benchmark fresh

@argument-ci-bot

argument-ci-bot Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

!benchmark — main vs 1de9691

backends: aiur=prove · envs: InitStd · baseline: fresh (benchmark products rebuilt, base-SHA run, bencher bypassed)

aiur · InitStd · prove — main from: base run @ 2911c16 (fresh — bencher bypassed)

7 constants · 4 with regressions · 5 with improvements (|Δ| > 3.0% on any metric).

IxVM on FRI (7 constants)
constant execute-time (main) execute-time (PR) Δ% prove-time (main) prove-time (PR) Δ% throughput (const/s) (main) throughput (const/s) (PR) Δ% peak-ram (main) peak-ram (PR) Δ% proof-size (main) proof-size (PR) Δ% verify-time (main) verify-time (PR) Δ% fft-cost (main) fft-cost (PR) Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append 8.813 s 8.852 s +0.4% 29.304 s 29.377 s +0.2% 94.700 94.460 -0.3% 72.85 GiB 72.83 GiB -0.0% 11.11 MiB 11.11 MiB +0.0% 63.2 ms 62.9 ms -0.5% 136.01B 136.01B +0.0%
Char.ofOrdinal_le_of_le 6.844 s 6.823 s -0.3% 24.993 s 25.849 s +3.4% ⚠️ 110.550 106.890 -3.3% ⚠️ 65.85 GiB 65.91 GiB +0.1% 11.11 MiB 11.11 MiB +0.0% 62.8 ms 65.5 ms +4.2% ⚠️ 104.21B 104.21B +0.0%
Array.extract_append 6.379 s 6.467 s +1.4% 22.227 s 22.199 s -0.1% 72.260 72.350 +0.1% 53.03 GiB 52.99 GiB -0.1% 11.02 MiB 11.02 MiB +0.0% 58.6 ms 56.6 ms -3.5% 🟢 98.00B 98.00B +0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq 3.516 s 3.568 s +1.5% 13.822 s 13.891 s +0.5% 135.080 134.400 -0.5% 34.99 GiB 35.03 GiB +0.1% 11.05 MiB 11.05 MiB +0.0% 56.3 ms 56.9 ms +1.1% 56.49B 56.49B +0.0%
Std.HashMap 4.066 s 3.984 s -2.0% 15.121 s 14.994 s -0.8% 135.040 136.190 +0.9% 37.37 GiB 37.35 GiB -0.1% 11.06 MiB 11.06 MiB +0.0% 61.1 ms 61.7 ms +1.0% 62.93B 62.93B +0.0%
String.append 429.0 ms 449.7 ms +4.8% ⚠️ 2.139 s 2.113 s -1.2% 152.910 154.730 +1.2% 4.96 GiB 4.95 GiB -0.1% 9.78 MiB 9.78 MiB +0.0% 50.9 ms 50.3 ms -1.2% 3.45B 3.45B +0.0%
Nat.add_comm 262.1 ms 263.5 ms +0.5% 973.2 ms 983.4 ms +1.0% 47.260 46.780 -1.0% 4.67 GiB 4.50 GiB -3.7% 🟢 8.95 MiB 8.95 MiB +0.0% 42.4 ms 42.9 ms +1.2% 314.33M 314.33M +0.0%
FRI verifier on FRI (7 constants)
constant execute-time (main) execute-time (PR) Δ% prove-time (main) prove-time (PR) Δ% throughput (const/s) (main) throughput (const/s) (PR) Δ% peak-ram (main) peak-ram (PR) Δ% proof-size (main) proof-size (PR) Δ% verify-time (main) verify-time (PR) Δ% fft-cost (main) fft-cost (PR) Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append 4.840 s 4.826 s -0.3% 30.484 s 29.998 s -1.6% 91.030 92.500 +1.6% 102.13 GiB 101.68 GiB -0.4% 4.00 MiB 4.00 MiB +0.0% 25.0 ms 33.1 ms +32.5% (1.32× slower) ⚠️ 209.69B 209.69B +0.0%
Char.ofOrdinal_le_of_le 4.898 s 4.945 s +1.0% 29.907 s 29.900 s -0.0% 92.390 92.410 +0.0% 101.72 GiB 101.65 GiB -0.1% 4.00 MiB 4.00 MiB +0.0% 27.6 ms 24.7 ms -10.7% (1.12× faster) 🟢 208.66B 208.66B +0.0%
Array.extract_append 4.631 s 4.642 s +0.2% 27.465 s 27.613 s +0.5% 58.480 58.160 -0.5% 91.76 GiB 91.81 GiB +0.1% 4.00 MiB 4.00 MiB +0.0% 23.1 ms 23.3 ms +0.9% 198.51B 198.51B +0.0%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq 4.809 s 4.889 s +1.7% 29.379 s 30.735 s +4.6% ⚠️ 63.550 60.750 -4.4% ⚠️ 99.53 GiB 99.49 GiB -0.0% 4.02 MiB 4.02 MiB +0.0% 38.6 ms 22.4 ms -41.8% (1.72× faster) 🟢 203.56B 203.56B +0.0%
Std.HashMap 4.883 s 4.786 s -2.0% 29.086 s 28.963 s -0.4% 70.210 70.500 +0.4% 97.27 GiB 97.34 GiB +0.1% 4.00 MiB 4.00 MiB +0.0% 37.4 ms 34.7 ms -7.3% (1.08× faster) 🟢 208.71B 208.71B +0.0%
String.append 3.875 s 3.858 s -0.4% 26.097 s 26.392 s +1.1% 12.530 12.390 -1.1% 89.25 GiB 89.35 GiB +0.1% 4.00 MiB 4.00 MiB +0.0% 23.4 ms 23.4 ms +0.0% 169.61B 169.61B +0.0%
Nat.add_comm 3.115 s 3.092 s -0.7% 17.280 s 17.329 s +0.3% 2.660 2.650 -0.4% 58.67 GiB 58.66 GiB -0.0% 4.01 MiB 4.01 MiB +0.0% 21.2 ms 21.8 ms +2.9% 129.00B 129.00B +0.0%
Aggregate flat join (7 constants)
constant execute-time (main) execute-time (PR) Δ% prove-time (main) prove-time (PR) Δ% peak-ram (main) peak-ram (PR) Δ% proof-size (main) proof-size (PR) Δ% verify-time (main) verify-time (PR) Δ% fft-cost (main) fft-cost (PR) Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a
Char.ofOrdinal_le_of_le n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a
Array.extract_append n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a
Std.HashMap n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a
String.append n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a
Nat.add_comm n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a
Pipeline total (7 constants)
constant total-time (main) total-time (PR) Δ% pipeline-throughput (const/s) (main) pipeline-throughput (const/s) (PR) Δ% pipeline-peak-ram (main) pipeline-peak-ram (PR) Δ%
ByteArray.utf8DecodeChar?_utf8EncodeChar_append 59.788 s 59.375 s -0.7% 46.410 46.740 +0.7% 102.13 GiB 101.68 GiB -0.4%
Char.ofOrdinal_le_of_le 54.900 s 55.749 s +1.5% 50.330 49.560 -1.5% 101.72 GiB 101.65 GiB -0.1%
Array.extract_append 49.691 s 49.812 s +0.2% 32.320 32.240 -0.2% 91.76 GiB 91.81 GiB +0.1%
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq 43.201 s 44.626 s +3.3% ⚠️ 43.220 41.840 -3.2% ⚠️ 99.53 GiB 99.49 GiB -0.0%
Std.HashMap 44.207 s 43.957 s -0.6% 46.190 46.450 +0.6% 97.27 GiB 97.34 GiB +0.1%
String.append 28.236 s 28.506 s +1.0% 11.580 11.470 -0.9% 89.25 GiB 89.35 GiB +0.1%
Nat.add_comm 18.254 s 18.313 s +0.3% 2.520 2.510 -0.4% 58.67 GiB 58.66 GiB -0.0%

Workflow logs

Point Aiur at the multi-stark revision that balances large mixed-height commitments across host CPUs and CUDA according to device capacity. The backend carries hybrid placement through lookup, quotient, openings, and FRI while preserving the existing Goldilocks/BLAKE3 protocol and CPU verification path.

No Ix source or build configuration changes are required: the existing opt-in cuda feature chain selects the new backend, while default builds remain CUDA-independent. Cargo.lock records the matching dependency graph.

Validated with release clippy across the non-CUDA CI feature set, a release all-targets CUDA check, and IX_CUDA=1 lake build bench-typecheck, including the final Rust static-library link.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants