llama.cpp
Fix granite speech model inference by applying embedding scale when deepstack is not used
#24357
Merged

Fix granite speech model inference by applying embedding scale when deepstack is not used #24357

ngxson merged 3 commits into ggml-org:master from arnu515:master
arnu515
arnu515 llama-graph : apply embedding scale when deepstack is not used
c4b22de6
arnu515 arnu515 requested a review from CISC CISC 101 days ago
ngxson
gabe-l-hart
gabe-l-hart approved these changes on 2026-06-09
CISC
CISC approved these changes on 2026-06-09
gabe-l-hart
gabe-l-hart
ngxson
ngxson nits: remove non-existant hunyuan-vl from the tests
4584f943
ngxson ngxson requested a review 100 days ago
ngxson ngxson requested a review from CISC CISC 100 days ago
github-actions github-actions added examples
ngxson
ngxson apply suggestion from @gabe-l-hart
c28b9386
ServeurpersoCom
ServeurpersoCom approved these changes on 2026-06-09
ngxson ngxson merged d73cd076 into master 100 days ago

Login to write a write a comment.

Login via GitHub

Assignees
No one assigned
Labels
Milestone