auto-round
a1480af3
- perf: hoist per-group scale and split accumulators in int4 MoE decode GEMV
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
13 days ago
perf: hoist per-group scale and split accumulators in int4 MoE decode GEMV Co-authored-by: a32543254 <53296245+a32543254@users.noreply.github.com>
References
#2111 - Optimize INT4 MoE prefill/decode (sym) XPU kernels with dedicated w4a16 tiling
#2143 - feat: W4A8 ARK XPU MoE kernel (int4 weight / int8 compute) with prefill + decode
Author
Copilot
Parents
23ba0dbf
Loading