llama.cpp
opencl: perf optimization for mamba2 ssm_scan by folding 4 dim rows into one workgroup and add more coverage
#27775
Open

Loading