llama.cpp
4a84b0ad - metal : add TQ2_0 support (#26980)

Commit
1 day ago
metal : add TQ2_0 support (#26980) * metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : optimize mul_mv kernel - float ops over integer ops - precalculate sums - hoist coef out of the inner loop - contiguous y loads llama.cpp:DeepSeek-v4-Flash-0731
Author
Parents
Loading