Reproducible study of how tensor layout, reduction length, and pointer alignment drive vendor GEMM dispatch and latency cliffs.
benchmarking gpu cuda pytorch cuda-kernels performance-analysis reproducibility gemm gpu-profiling systems-research torchinductor deep-learning-compilers tensor-layout kernel-dispatch
-
Updated
Sep 2, 2026 - Python