GPU memory optimization and kernel fusion for the Alpaka backend - #54
Open
ioantreas wants to merge 82 commits into
Open
GPU memory optimization and kernel fusion for the Alpaka backend#54ioantreas wants to merge 82 commits into
ioantreas wants to merge 82 commits into
Conversation
Member
|
/runtest h100-47gb |
|
|
|
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR extends the SOFIE Alpaka GPU backend with two main optimizations: GPU intermediate-memory management and GPU kernel fusion. It also expands the benchmark infrastructure used to evaluate these changes against other GPU inference backends.
Main changes
GPU intermediate-memory allocation
GPU kernel fusion
RModel_Fusion_ALPAKA.cxxto keep the main Alpaka code generator separated from fusion planning.GPU operator/model support
Benchmarking
Expand the SOFIE GPU benchmark to support comparisons with:
The benchmark now includes:
The build configuration also supports optional custom paths for ONNX Runtime, TensorRT, benchmark models, and the Mamba
selective_scan_cudalibrary.PyTorch AOTInductor and Mamba
.pt2packages.Documentation
Update the benchmark README with: