transformers
Update Gemma4 weight conversion script
#45328
Merged

Update Gemma4 weight conversion script #45328

RyanMullins
RyanMullins RyanMullins force pushed from 48be161d to 23b00720 125 days ago
lucianommartins
Cyrilvallez
Cyrilvallez
Cyrilvallez
RyanMullins
RyanMullins RyanMullins force pushed from 23b00720 to c3635231 124 days ago
RyanMullins RyanMullins changed the title fix: KV cache sharing Drop unused submodules from Gemma4TextAttention layers that share KV Cache 124 days ago
RyanMullins
Cyrilvallez
RyanMullins RyanMullins closed this 124 days ago
RyanMullins RyanMullins reopened this 124 days ago
github-actions
RyanMullins RyanMullins force pushed from a187bcc8 to 09156ad1 124 days ago
RyanMullins fix: KV cache sharing
bf50b4ae
RyanMullins fix: duplicate shared_kv_states in function params
a09b0706
RyanMullins fix: module style
7d863e52
RyanMullins revert: modeling changes
2685bc70
RyanMullins RyanMullins force pushed from 09156ad1 to 2685bc70 124 days ago
github-actions
RyanMullins RyanMullins changed the title Drop unused submodules from Gemma4TextAttention layers that share KV Cache Drop unused Gemma4TextAttention weights when sharing KV Cache 124 days ago
RyanMullins
Cyrilvallez Cyrilvallez changed the title Drop unused Gemma4TextAttention weights when sharing KV Cache Update Gemma4 weight conversion script 112 days ago
Cyrilvallez
Cyrilvallez approved these changes on 2026-04-22
Cyrilvallez Cyrilvallez merged 2da596b0 into main 112 days ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone