Skip to content

Pull requests: pytorch/helion

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Disable CUDAGraph timing for ROCm benchmarks ciflow/rocm CLA Signed This label is managed by the Meta Open Source bot. module: rocm
#3471 opened Aug 26, 2026 by yushangdi Contributor Draft
[cutedsl] Add a reproducible attention backend comparison benchmark and fused-ReLU attention example kernels CLA Signed This label is managed by the Meta Open Source bot.
#3470 opened Aug 26, 2026 by jansel Contributor Loading…
[cutedsl] Refine CuTe flash autotuning finalists with a deterministic terminal refinement stage CLA Signed This label is managed by the Meta Open Source bot.
#3469 opened Aug 26, 2026 by jansel Contributor Loading…
[cutedsl] Generalize CuTe flash attention schedules across workloads and add autotuner search support for them CLA Signed This label is managed by the Meta Open Source bot.
#3468 opened Aug 26, 2026 by jansel Contributor Loading…
[cutedsl] Reuse warmed nested-input kernels under torch.compile and isolate backend tests from shared cache and capability state CLA Signed This label is managed by the Meta Open Source bot.
#3467 opened Aug 26, 2026 by jansel Contributor Loading…
[benchmark] Limit MI350X CUDAGraph repetitions CLA Signed This label is managed by the Meta Open Source bot.
#3466 opened Aug 25, 2026 by yushangdi Contributor Draft
[autotuner] harden B200 matmul heuristics for off-corpus kernels CLA Signed This label is managed by the Meta Open Source bot.
#3464 opened Aug 25, 2026 by calebmkim Contributor Loading…
[cute] Add pretuned B200/GB300 flash attention kernel CLA Signed This label is managed by the Meta Open Source bot.
#3463 opened Aug 25, 2026 by yushangdi Contributor Draft
Optimize tcgen05 AB consumer waits CLA Signed This label is managed by the Meta Open Source bot.
#3454 opened Aug 24, 2026 by yushangdi Contributor Draft
Optimize M64 scale_mm with four-CTA schedule CLA Signed This label is managed by the Meta Open Source bot.
#3453 opened Aug 24, 2026 by yushangdi Contributor Draft
[compiler] Allow torch.matmul rank broadcasting; add sparse-attention indexer example CLA Signed This label is managed by the Meta Open Source bot.
#3450 opened Aug 22, 2026 by MrlixiangWE Loading…
[Pallas] Optimize compact-worklist masks for emit pipelines CLA Signed This label is managed by the Meta Open Source bot.
#3439 opened Aug 21, 2026 by ethche Contributor Loading…
[Pallas] Enable GDN forward-h example on TPU CLA Signed This label is managed by the Meta Open Source bot.
#3438 opened Aug 21, 2026 by thcmbs Collaborator Loading…
FlyDSL: minimal elementwise-only backend CLA Signed This label is managed by the Meta Open Source bot.
#3435 opened Aug 20, 2026 by umechand-amd Collaborator Draft
Support nested function definitions CLA Signed This label is managed by the Meta Open Source bot.
#3433 opened Aug 20, 2026 by parsshar-RH Contributor Loading…
[Pallas] Enable Mamba2 chunk scan on TPU CLA Signed This label is managed by the Meta Open Source bot.
#3432 opened Aug 20, 2026 by thcmbs Collaborator Loading…
[Pallas] Support scalar stores to resident VMEM outputs CLA Signed This label is managed by the Meta Open Source bot.
#3431 opened Aug 20, 2026 by thcmbs Collaborator Loading…
[autotuner] extend the B200 matmul heuristic to multi-matmul kernels CLA Signed This label is managed by the Meta Open Source bot.
#3409 opened Aug 17, 2026 by calebmkim Contributor Loading…
[Pallas] Keep private remote-copy buffers in VMEM CLA Signed This label is managed by the Meta Open Source bot.
#3408 opened Aug 17, 2026 by AmesingFlank Contributor Loading…
[Pallas] Reuse aligned HBM windows as resident matmul Refs CLA Signed This label is managed by the Meta Open Source bot.
#3407 opened Aug 17, 2026 by AmesingFlank Contributor Loading…
[Pallas] Drain computed-source sends after device loops CLA Signed This label is managed by the Meta Open Source bot.
#3405 opened Aug 17, 2026 by AmesingFlank Contributor Loading…
[Pallas] Overlap read-only input DMAs with weight streams CLA Signed This label is managed by the Meta Open Source bot.
#3403 opened Aug 17, 2026 by AmesingFlank Contributor Loading…
[Pallas] Lower static advanced indices as basic slices CLA Signed This label is managed by the Meta Open Source bot.
#3402 opened Aug 17, 2026 by AmesingFlank Contributor Loading…
[Pallas] Preserve computed scalar indices across loop lowering CLA Signed This label is managed by the Meta Open Source bot.
#3401 opened Aug 17, 2026 by AmesingFlank Contributor Loading…
[Pallas] Pass TPU VMEM capacity to Mosaic CLA Signed This label is managed by the Meta Open Source bot.
#3400 opened Aug 17, 2026 by AmesingFlank Contributor Loading…
ProTip! Add no:assignee to see everything that’s not assigned.