llama.cpp
d22d2be2 - llama: make the qwen4exp PLE n-gram history per context and serialise it

Commit
6 days ago
llama: make the qwen4exp PLE n-gram history per context and serialise it The PLE hash of a token mixes in the ple_ngram_size - 1 tokens before it, which a decode ubatch does not carry, so they were remembered in a map on llama_model_qwen4exp. That is the wrong owner twice over. A llama_model is shared by every context that loads it, and the map was keyed only by llama_seq_id, so two contexts running the same sequence id - two server instances on one model, or a draft/target pair - overwrote each other's window. The next_pos guard turned that into EOS padding instead of a crash, so it degraded quality silently. The map was also in no state blob: grep found ple_hist in neither llama-kv-cache.cpp nor llama-memory-*.cpp nor llama-context.cpp. A restored context therefore failed the next_pos check on its first ubatch and hashed the first tokens after the restore against EOS padding. This is why a session blob round-tripped byte for byte while the restored context computed different logits: the state was never in the bytes. It moves to llama_memory_hybrid_idx, which is per context, is the memory type qwen4exp always builds, and already does the per-sequence bookkeeping this needs. Every sequence operation now carries the window with it: seq_rm a rewind (p1 < 0) truncates the window to the surviving prefix and moves next_pos to p0, so a rollback keeps exact context; a hole punched in the middle leaves the window non-contiguous, so it is dropped seq_cp the destination inherits the source's window, truncated to the copied position range - a copied sequence continues with the same n-grams the source would have used seq_keep every other sequence's window is dropped, like its cells seq_add a shift that moves the whole window keeps it and moves next_pos with it, which is the context-shift case; one that cuts through it drops it seq_div positions stop being consecutive, so an overlapping window is dropped clear everything is dropped Dropping means next_pos = -1, which set_input turns into full EOS padding: the same thing a fresh sequence gets, and the same thing this code did before it followed the sequence operations at all, so no case is worse than before. The state payload is a self-delimiting list, u32 count then per entry { i32 seq_id, i32 next_pos, u32 n_toks, i32 toks[n_toks] }, so a whole-context save and a single-sequence save share one format and a single-sequence restore can retarget the window at its destination seq_id. It is written after the indexer section, last, for the same reason that one is: as a pure suffix an older reader stops early instead of parsing these bytes as something else. Unlike the indexer section it is not under LLAMA_STATE_SEQ_FLAGS_PARTIAL_ONLY. The window is recurrent state - it is the input the PLE convolution's own recurrent state is derived from - and the recurrent cache beside it is written for partial checkpoints too. Gating it would leave the server's speculative decoding checkpoints restoring the conv state without the window that produced it. No further version bump: LLAMA_SESSION_VERSION 10 and LLAMA_STATE_SEQ_VERSION 3 were introduced for the indexer section in the same unreleased series, and both changes are qwen4exp-only additions to the same blob layout. Also fixes the padding of a short window. set_input pads a window shorter than ngram_size - 1 up to that length, but prev() indexes the snapshot with the most recent token last, and resize() pads at the back, so the filler EOS landed where the immediately preceding token belongs. It now pads at the front. A window is short at a sequence start after a one-token prefill, and after a seq_rm rewind, which the new bookkeeping makes common. Every architecture other than qwen4exp builds llama_memory_hybrid rather than llama_memory_hybrid_idx, has no PLE table and never asks for a history, so nothing about its graph, its sequence operations or its state bytes changes. (cherry picked from commit de170364c052c68fcf63285cc0028095edb9f23c)
Author
Committer
Parents
Loading