Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
ggerganov/llama.cpp
Pull Requests
Commits
Open
Closed
rpc : read cached tensor files with fread instead of std::ifstream
ggml
#28344 opened 2026-09-03 19:14 by
ikari-pl
vulkan: use shared-memory reduction for dequant mul_mat_vec on Imagination PowerVR
Vulkan
ggml
#28341 opened 2026-09-03 18:44 by
JulianPscheid
CUDA/HIP: derive MMVQ nwarps from kernel launch bounds
ggml
CUDA
#28339 opened 2026-09-03 17:33 by
ACNoonan
model, mtmd: fix gemma4 vision handling
model
mtmd
#28335 opened 2026-09-03 17:22 by
ngxson
args: remove mmap/mlock/dio flags from arg parser
documentation
server
#28334 opened 2026-09-03 16:14 by
taronaeo
fix: speculative: zero MTP carrier at sequence start
#28333 opened 2026-09-03 16:03 by
mxxm-t
ggml-cpu(s390x): fix q5_1 uninitialized v_acc
ggml
merge ready
#28332 opened 2026-09-03 15:23 by
taronaeo
respect -fitc from llama-bench instead of using the benchmark context size
#28331 opened 2026-09-03 15:02 by
MartinEmrich
memory : avoid allocating V cache for indexer (it's not used) in Qwen3.8-Flash-Next (qwen4exp)
#28330 opened 2026-09-03 14:36 by
fairydreaming
ggml-cpu : add OS XSAVE checks to x86 CPU feature scoring.
ggml
#28329 opened 2026-09-03 14:26 by
jesuscbm
fix(tool-call): capture non-string args raw to fix array<object> parsing
documentation
model
build
testing
Vulkan
devops
server
ggml
SYCL
Apple Metal
Ascend NPU
OpenCL
CUDA
OpenVINO
conversion
#28327 opened 2026-09-03 13:55 by
cyberluke
llama: refactor lazy mode auto, fix iGPU regression
examples
#28326 opened 2026-09-03 13:29 by
0cc4m
tests : use 1 thread for data initialization
testing
#28325 opened 2026-09-03 13:02 by
ggerganov
convert: fix LoRA conversion crash for Qwen3.5 V-head reorder
conversion
#28324 opened 2026-09-03 12:59 by
Swigler
src : add n_expert_used_max function
#28323 opened 2026-09-03 12:49 by
danbev
common : make build info output stream configurable
examples
#28322 opened 2026-09-03 12:42 by
angt
server : encode images on the GPU by borrowing the idle compute buffers
server
mtmd
#28320 opened 2026-09-03 12:27 by
efschu
api : mark `llama_context_params.ctx_other` as `const`
#28316 opened 2026-09-03 11:38 by
madsmtm
ROCm: resolve TOP_K kernels
ggml
CUDA
#28313 opened 2026-09-03 10:52 by
pwilkin
fix edge case for repetition count threshold
testing
#28311 opened 2026-09-03 10:45 by
SergeySklyarov
graph : keep the backend sampling subgraph static across ubatches also helps fix CI issue
#28305 opened 2026-09-03 07:59 by
ynankani
HIP : add shared-memory tiled transpose fast path for F32 concat on RDNA3.5
testing
ggml
CUDA
draft
#28303 opened 2026-09-03 07:14 by
Casten-Wang
server : fix context checkpoint eviction on prompts shorter than checkpoint_min_step
server
#28302 opened 2026-09-03 07:12 by
masterFoad
metal : skip the empty half of the mul_mm_id token tile, load iq2/iq3 codebooks as uint32
testing
ggml
Apple Metal
#28301 opened 2026-09-03 07:09 by
masterFoad
compilation: detect and enable AVX-VNNI compilation on MSVC
ggml
#28297 opened 2026-09-03 05:55 by
mmpataki
common : add defer_loading to tool definitions
documentation
testing
server
#28292 opened 2026-09-03 04:45 by
prskid1000
add gt-asr-1.7b
server
mtmd
conversion
#28289 opened 2026-09-03 03:46 by
ypw757
fix the msvc clang compile test backend ops mismatch
ggml
#28284 opened 2026-09-03 00:20 by
sarahwu185
imatrix : parallelize with OpenMP
examples
#28283 opened 2026-09-02 23:51 by
sanmai
Fixed GBNF grammar generation for empty object schemas
testing
examples
#28279 opened 2026-09-02 23:17 by
SergeySklyarov
Older