llama.cpp
opencl: Q6_K GEMM/GEMV fix for ne01 of weights that are not multiples of 128.
#25464
Merged

Commits
  • opencl: fix garbled output for Q6_K weights with ne01 % 128 != 0 on Adreno
    lhez committed 55 days ago
  • opencl: reserve alignment slack for the SOA subbuffer carve in alloc size
    lhez committed 55 days ago
  • opencl: use lm based q6_k mm when ne1 is not multiple of 128
    lhez committed 55 days ago
Loading