GPU-accelerated linear solvers based on the conjugate gradient (CG) method, supporting NVIDIA and AMD GPUs with GPU-aware MPI, NCCL, RCCL or NVSHMEM
-
Updated
Mar 14, 2026 - C
GPU-accelerated linear solvers based on the conjugate gradient (CG) method, supporting NVIDIA and AMD GPUs with GPU-aware MPI, NCCL, RCCL or NVSHMEM
Interactive web visualization for understanding collective communication algorithms (as used in NCCL, RCCL, MPI). Learn how AllReduce, Broadcast, Reduce, AllGather and more work step by step.
Pure-Python drop-in subset of Ray's API that makes vLLM distributed inference work across nodes. ~1,400 LoC, one dependency, bind-mount into the stock vLLM image. NVIDIA + AMD, validated cross-node.
Multi-GPU vLLM on consumer Radeon (RX 7900 XT/XTX, gfx1100): root cause and fix for the RCCL "operation cannot be performed in the present state" crash, plus 292 benchmark measurements across five model architectures
Royal Caribbean custom integration for Home Assistant
Add a description, image, and links to the rccl topic page so that developers can more easily learn about it.
To associate your repository with the rccl topic, visit your repo's landing page and select "manage topics."