Windows prebuilt of llama.cpp combining Multi-Token Prediction (MTP) + TurboQuant KV cache compression + native sm_120 (Blackwell consumer GPU, FP4 tensor cores). For RTX 5060 Ti / 5070 / 5080 / 5090.
-
Updated
Sep 17, 2026
Windows prebuilt of llama.cpp combining Multi-Token Prediction (MTP) + TurboQuant KV cache compression + native sm_120 (Blackwell consumer GPU, FP4 tensor cores). For RTX 5060 Ti / 5070 / 5080 / 5090.
Blackwell-optimized llama.cpp Docker image – works on all NVIDIA GPUs, but tuned for RTX 50 series. Built from scratch with CUDA 12.8, sm_120, NVFP4-ready. 250+ tok/s on 4B F16. Includes llama-chat script.
Building PaddlePaddle Inference C++ from source for NVIDIA Blackwell (RTX 50-series, sm_120) with CUDA 13 and TensorRT 10
Optimized CSM-1B TTS pipeline for RTX 5090 (Blackwell sm_120). CUDA graph replay via patched HF Transformers. ~0.46x RTF. Topics (tags): csm text-to-speech rtx-5090 blackwell cuda-graphs torch-compile sesame streaming pytorch
Run llama.cpp with Multi-Token Prediction and TurboQuant on Windows using native sm_120 Blackwell support for RTX 50-series GPUs.
OpenAlchemy fork of whisper.cpp — RTX 50-series Blackwell (sm_120) NULL-slot guards on top of upstream. Powers /v1/audio/transcriptions
GEN3C: Generative Novel 3D Captions - Adapted for NVIDIA Blackwell GPU architecture (sm_120). Includes automatic GPU detection, CPU-based T5 text encoding for Blackwell compatibility, and full backward compatibility with older GPUs.
To associate your repository with the sm-120 topic, visit your repo's landing page and select "manage topics."