llama.cpp
c173a53b - llama : fix unexpected graph reallocation in the k-pool models (#29958)

Commit
5 days ago
llama : fix unexpected graph reallocation in the k-pool models (#29958) * llama : fix unexpected graph reallocation in the k-pool models Both k-pool models built a graph shape that depends on state the full-context reserve cannot know: - qwen4exp branched on inp->cache_safe, which turns false as soon as llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA layers swapped scatter+gather for fill+concat and dropped the new_pool_rep leaf, so the decode graph had 12 fewer nodes than the reserved one - glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the TG decode built the gather shape (7564 nodes) while the last reserve, the PP one, had the dense shape (7762 nodes) Either mismatch forces a decode-time re-reserve that drops the worst-case sizing and bakes in the current state, so the next state growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.: GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench \ -hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2 \ -c 32768 -pps -kvu Always scatter+gather the pooled keys, and pick gather from context constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds n_sel. Every graph of a context then shares one shape, which the reserve covers, and the dense path measured faster than the gather path at 2.5k and 16k context. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * llama : drop the unused k-pool cache_safe graph API The k-pool graphs no longer branch on cache_safe, so nothing reads get_kpool_cache_safe() or the conditional new_pool_rep any more: both models always pass the scatter target, which set_input_kpool now requires instead of merely preferring. Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only helper. The layout and state flag itself stays, it still decides which pools a layout with shared cells must re-pool. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * tests : add a shared-seq graph reserve regression test Decode a prompt into seq 0, share its cells with seq 1 via llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep decoding both sequences. For the k-pool models sharing clears cache_safe, which changes the graph topology while the pools keep growing, so a scheduler that re-reserves with the current state instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1. The test registration sets that flag, and the test aborts on both k-pool models before 2220411ec1. kimi-linear and minimax-01 are skipped: they reserve the final pp graph with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so every multi-seq graph has a different layout and re-reserves by design. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * cont : add TODOs * cont : fix comment * cuda: match the moe weighted reduction on empty ubatches ggml_cuda_match_moe_weighted_reduction rejected tensors with zero rows. A ubatch without outputs shrinks the last layer to zero rows through inp_out_ids, so graph_optimize dropped its alloc dep there and the scheduler graph lost one node compared to the reserved one. The scheduler then re-reserved at the size of that ubatch, and the next ubatch with the same node count but larger tensors aborted under GGML_SCHED_DEBUG_REALLOC=1. The compute loop already skips empty nodes before trying any fusion, so the guard only made the alloc deps depend on the row count. * tests: build the rollback test only where internal symbols link The shared-seq case calls llm_arch_from_string, which libllama does not export through LLAMA_API, so linking test-recurrent-state-rollback fails on Windows with shared libraries. Its build now sits in the NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and the test registration it already lives under. * tests: skip archs by name in the shared-seq reserve test The skip of kimi-linear and minimax-01 went through llm_arch_from_string, which libllama does not export through LLAMA_API, so the test could not link on Windows with shared libraries. It now compares the general.architecture string directly, and the test builds on every platform again. --------- Co-authored-by: Pascal <admin@serveurperso.com>
Author
Parents
Loading