llama.cpp
6c487e2f - server: enforce prompt cache RAM limit (#25070)

Commit
43 days ago
server: enforce prompt cache RAM limit (#25070) Before this commit, --cache-ram was not a hard limit: - The cache always kept at least one entry, even if that entry exceeded the RAM/token limits. - Old entries were only evicted for the RAM/token limits after saving the new one, which could cause the cache to temporarily exceed the RAM/token limits even if individual entries were below the limit. Now, ensure that the RAM limit is strict with these changes: - Skip saving state to cache if by itself it exceeds the RAM limit. - Evict old entries as necessary to make the new entry fit. Additionally, token-limit cleanup may now evict the last remaining cache entry instead of always preserving one.
Author
Parents
Loading