llama.cpp
llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4
#25370
Merged

llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 #25370

fairydreaming
sszymczy llama : make all KQ masks (except the lightning indexer one) f16 if F…
17252233
github-actions github-actions added model
am17an
am17an approved these changes on 2026-07-07
fairydreaming
ggerganov
ggerganov
fairydreaming
sszymczy llama : remove dead code that repeats unified raw_k cache for each st…
a0e9766f
fairydreaming fairydreaming changed the title llama : make all KQ masks (except the lightning indexer one) f16 if FA is used and remove zero attention bias in DeepSeek V4 llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 40 days ago
sszymczy Merge remote-tracking branch 'upstream/master' into dsv4-remove-kq-bias
2b0d9804
am17an
am17an approved these changes on 2026-07-08
fairydreaming fairydreaming marked this pull request as ready for review 39 days ago
fairydreaming fairydreaming requested a review from CISC CISC 39 days ago
am17an am17an added merge ready
ggerganov
ggerganov approved these changes on 2026-07-10
fairydreaming fairydreaming merged 2ed3c1ab into master 38 days ago
arch-btw
arch-btw
fairydreaming

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone