ggml : fix ggml_gallocr_ptr type (ggml/1205)
fc0c1877
rpc : do not wait for response when sending RPC_CMD_SET_TENSOR (llama…
fd0ca813
change the reorder tensor from init to execute OP (llama/13003)
abb24a63
ggml: move fp16/bf16 conversion optimizations to CPU backend + export…
123bf286
musa: fix build warning (llama/13129)
3c7e3afc
CUDA: fix q_nope_absorbed prec for DS 2 Lite f16 (llama/13137)
32b93f59
musa: fix typo in cc control (llama/13144)
74f09c41
SYCL: Add all missing unary kernels (llama/13074)
b89f84f8
fix(rpc): Improve input validation and error handling (llama/13069)
fa4d6b45
CUDA: fix non-cont. inputs for batched mat mul (llama/13155)
16121dc3
feat(ggml-cpu): enable z17 compile (llama/13182)
cd8054fc
ggml : fix ppc64le build (llama/13176)
80dd958f
vulkan: use uint array index to avoid glslang bug (llama/13193)
4d38d023
CUDA: batched+noncont MMQ, refactor bs>1 MoE code (llama/13199)
cd618ce0
sync : ggml
16b50222
talk-llama : sync llama.cpp
842678bf
danbev
approved these changes
on 2025-05-01
ggerganov
merged
0778b6ff
into master 1 year ago
ggerganov
deleted the sync-ggml-25-05-01 branch 1 year ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub