llama.cpp
035e2273 - llama: hold the qwen4exp indexer cache in a new llama_memory_hybrid_idx

Commit
7 days ago
llama: hold the qwen4exp indexer cache in a new llama_memory_hybrid_idx The indexer key cache was added by extending llama_memory_hybrid with an optional third cache, and the host-side cell/block mapping that drives QSA was added as set_input_qsa on llama_kv_cache. Both are shared classes that every hybrid and every attention model goes through. Move both into a new memory type, llama_memory_hybrid_idx, following llama_kv_cache_msa: the indexer cache and the pos<->cell translation live with the sparse-attention memory rather than in the classes that serve every other architecture. llama-kv-cache.{h,cpp} and llama-memory-hybrid.{h,cpp} are restored to their unmodified state. init_batch is repeated from llama_memory_hybrid because the indexer cache has to be handed the attention cache's slot infos, and those are not reachable through the context the base returns. Allocating them separately lets the two caches drift, which is what pointed QSA's top-k at the wrong cells before. The context derives from llama_memory_hybrid_context so build_inp_mem_hybrid keeps working unchanged, and get_n_stream is computed from the slot infos exactly as llama_kv_cache_context did. Behaviour is unchanged: logits over an 8192-token sequence are bit-identical to the previous implementation, sparse and dense alike.
Author
Committer
Parents
Loading