Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[SYCL] Support limit max alloc memory within 2GB for host-pinned memory documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#27559 opened Aug 22, 2026 by arthw Contributor Draft
HIP: Expand Q5_K and Q6_K tile widths for RDNA2 CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27558 opened Aug 22, 2026 by draetheus Loading…
vulkan: double the K-quant int-mmq large tile on RDNA3 iGPUs ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#27554 opened Aug 22, 2026 by aic0d3r Contributor Loading…
[SYCL] enhance the api to support peer-to-peer copy ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#27550 opened Aug 22, 2026 by arthw Contributor Draft
test: move tools/parser to tests documentation Improvements or additions to documentation examples testing Everything test related
#27548 opened Aug 22, 2026 by ngxson Collaborator Loading…
webgpu : reorder fattn includes to fix unresolved V error ggml changes relating to the ggml tensor library for machine learning WebGPU
#27545 opened Aug 22, 2026 by fairydreaming Contributor Loading…
webgpu : fix handling of infinity values during ARGSORT and TOP_K ggml changes relating to the ggml tensor library for machine learning WebGPU
#27538 opened Aug 22, 2026 by fairydreaming Contributor Loading…
hexagon: add alternative Hexagon NPU backend implementation ggml changes relating to the ggml tensor library for machine learning Hexagon
#27535 opened Aug 22, 2026 by zhouwg-jeffzhou Loading…
llama : fix K/V and recurrent state cleanup after failed restores testing Everything test related
#27530 opened Aug 22, 2026 by CHIPMUNK-T0T Contributor Loading…
Thread swizzling in kernel_mul_mm (Metal) for better cache locality. Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#27529 opened Aug 22, 2026 by skoulik Draft
vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#27526 opened Aug 22, 2026 by jeffbolznv Contributor Loading…
cuda : fuse RWKV7 recurrent input and state paths CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#27523 opened Aug 21, 2026 by 123123213weqw Contributor Draft
Quant: OCP FP8 E4M3 support conversion ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#27512 opened Aug 21, 2026 by ORippler Collaborator Draft
Add Q2_K reordered MMVQ and ESIMD kernels ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#27509 opened Aug 21, 2026 by malsbat Contributor Loading…
hexagon: the ggml_hexagon_supported_mul_mat function has incorrect logic when checking quantization types ggml changes relating to the ggml tensor library for machine learning Hexagon
#27502 opened Aug 21, 2026 by zhouwg-jeffzhou Draft
mtmd: fix hanging with specific video-vision mtmd in Windows documentation Improvements or additions to documentation mtmd Related to multimodal functionality (video/image/audio)
#27500 opened Aug 21, 2026 by craftingmod Loading…
Optimize Krea Vulkan - fusion and misc. changes CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related Vulkan Issues specific to the Vulkan backend
#27495 opened Aug 21, 2026 by pwilkin Member Loading…
Optimize Krea Vulkan - Flash Attention CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related Vulkan Issues specific to the Vulkan backend
#27494 opened Aug 21, 2026 by pwilkin Member Loading…
Optimize Krea Vulkan - MUL_MAT CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related Vulkan Issues specific to the Vulkan backend
#27493 opened Aug 21, 2026 by pwilkin Member Loading…
SVE 128 bit Implementation of gemm_q4_k_8x8_q8_k kernel ggml changes relating to the ggml tensor library for machine learning
#27491 opened Aug 21, 2026 by anubhavfujitsu Loading…
ggml : reuse compute buffers for MTP (#27282) ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#27489 opened Aug 21, 2026 by mushang0 Loading…
ProTip! Updated in the last three days: updated:>2026-08-19.