Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
ggml-org/llama.cpp
Pull Requests
Commits
Open
Closed
ggml-hexagon: add HTP unary ops for ABS and LOG
ggml
Hexagon
#27786 opened 2026-08-27 05:56 by
cqderek
hexagon: support for device discovery and create sessions on demand
ggml
Hexagon
#27785 opened 2026-08-27 04:38 by
max-krasnyansky
tests: add ABS and LOG coverage to test-backend-ops
testing
#27782 opened 2026-08-27 03:13 by
cqderek
speculative : fix shared-KV detection for draft-mtp contexts
#27781 opened 2026-08-27 02:29 by
Gwonk1
ggml-cpu: add F16 input to the FWHT
testing
ggml
Apple Metal
#27779 opened 2026-08-27 01:24 by
bri-prism
Revise function calling documentation
documentation
#27778 opened 2026-08-27 01:08 by
apollo-2006
[Metal/Apple Silicon Performance] Disable K-quant mul_mv_ext for M3 Pro in lieu of regular mul_mv
ggml
Apple Metal
#27776 opened 2026-08-27 00:49 by
DavidLandup0
opencl: perf optimization for mamba2 ssm_scan by folding 4 dim rows into one workgroup and add more coverage
ggml
OpenCL
#27775 opened 2026-08-26 23:48 by
wanghqc
add GLM-5.3-Flash (GLM5-Next) support
model
mtmd
conversion
#27773 opened 2026-08-26 23:27 by
timkhronos
llama-quantize: Add MoE chunk queue for faster multi-threaded quant creation
#27770 opened 2026-08-26 22:37 by
bartowski1182
opencl: add bin kernels `kernel_gemm_moe_q4_0_q8_1_dp4a_bin`, `kernel_gemm_moe_mxfp4_q8_1_dp4a_bin`
ggml
OpenCL
#27768 opened 2026-08-26 22:04 by
shawngu-quic
vulkan: add TQ1_0 support (mm, mat-vec, mat-vec-id, dequant, get_rows)
testing
Vulkan
ggml
#27765 opened 2026-08-26 20:56 by
Anjielon
chat : split specialized parsers into common/parsers
#27764 opened 2026-08-26 20:17 by
pwilkin
server : report live generation throughput
server
#27760 opened 2026-08-26 19:08 by
xternet
metal : fix memory leaks due to missing autoreleasepools
ggml
merge ready
Apple Metal
#27758 opened 2026-08-26 18:32 by
nikwen
tests : run test-save-load-state across all architectures
model
testing
#27755 opened 2026-08-26 17:37 by
ggerganov
model: add GLM-5-Next (GLM-5.3-Flash)
model
testing
mtmd
conversion
#27754 opened 2026-08-26 16:53 by
danielhanchen
ci : build only the ggml-hip backend for windows-rocm release
devops
#27753 opened 2026-08-26 16:12 by
harkgill-amd
model : add GLM-5.3-Flash (glm5next)
model
testing
conversion
#27752 opened 2026-08-26 16:07 by
eauchs
server : fix Responses API item parsing for reasoning and Codex compatibility
draft
#27751 opened 2026-08-26 15:04 by
shengbox
ui: Improve Chat Form Actions UI/UX (models selector, add panel)
server/ui
#27746 opened 2026-08-26 13:51 by
allozaur
ui: Replace per-conversation MCP overrides with per-conversation tool policy
server/ui
#27745 opened 2026-08-26 13:48 by
allozaur
model: add Qwen3.8-Flash-Next (qwen4exp)
model
testing
conversion
#27742 opened 2026-08-26 13:43 by
danielhanchen
fix: model loader tensor count mismatch when MTP block contains NVFP4 weights with scales / input scales.
#27736 opened 2026-08-26 12:18 by
stevelikesrhino
server : accept data: URLs for input_video and input_audio
server
#27735 opened 2026-08-26 11:48 by
geckguy
ggml-vulkan: allow fp32-only devices (Haswell hasvk)
Vulkan
ggml
#27723 opened 2026-08-25 23:59 by
scottjoyner
CUDA graphs on Pascal (sm_61): +40% on MoE, +7% on dense, no regression
ggml
CUDA
#27721 opened 2026-08-25 22:49 by
ArtyomITA
server : make /v1/models created timestamp stable
server
#27719 opened 2026-08-25 20:16 by
theycallmeloki
fix: Prevent Qwen3.6 reranker from allocating exponential memory as physical batch size increases
model
#27715 opened 2026-08-25 18:21 by
sredman
spec: Add benchmark-only synthetic speculative acceptance options
documentation
testing
server
#27711 opened 2026-08-25 16:57 by
gaugarg-nv
Older