MSVC+CUDA llama.cpp fork for Windows: KDA/GDN, long-context KV placement, TurboQuant/TCQ when measured, speculative draft. Lab defaults — TriAttention opt-in only, not recommended.
-
Updated
Jul 28, 2026 - C++
MSVC+CUDA llama.cpp fork for Windows: KDA/GDN, long-context KV placement, TurboQuant/TCQ when measured, speculative draft. Lab defaults — TriAttention opt-in only, not recommended.
Pre-built PyTorch wheels and build scripts for NVIDIA DGX Spark (GB10, sm_121, Blackwell, CUDA 13.0, ARM64)
Measured sm_110 / sm_110a facts for NVIDIA Jetson AGX Thor: tensor-core capability matrix (2:4 sparsity works on plain sm_110; tcgen05 + NVFP4 block-scale is sm_110a-only), CUDA 11/12->13 migration breaks, CPU ISA, and the bandwidth roofline. With probes.
Build the GPU inference stack from source, repeatably: for Blackwell sm_120 on CUDA 13.x and Python 3.14. Machine-readable build state, generated patches, and an abductive-triage skill for build failures.
Windows NVIDIA-only Triton 3.7.0 build pipeline for RTX 5090 / Blackwell sm_120a, with FP8 tl.dot validation and peak benchmark results.
Hunyuan3D-2 fork — image→textured 3D→sliced STL + part segmentation. RTX 50-series (Blackwell/sm_120), CUDA 13.0, Python 3.12, PyTorch 2.11+cu130.
Community FlashAttention fork for CUDA 13 and NVIDIA Blackwell SM120 inference workloads
Automated Docker build and serving stack for vLLM with native NVFP4 KV cache quantization on NVIDIA Blackwell (RTX 5090 / SM120)
Build paddlepaddle-gpu from source for NVIDIA DGX Spark and GB10 (Grace Blackwell, aarch64, CUDA 13, sm_121). Prebuilt wheel plus the exact patches upstream is missing.
Accelerate Kimi Delta Attention computations with high-performance CUTLASS kernels designed for NVIDIA SM90 architectures and beyond.
Add a description, image, and links to the cuda-13 topic page so that developers can more easily learn about it.
To associate your repository with the cuda-13 topic, visit your repo's landing page and select "manage topics."