auto-round
1f7b6313
- perf(ark): route decode-sized int4-sym moe through the asym scalar GEMV
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
12 days ago
perf(ark): route decode-sized int4-sym moe through the asym scalar GEMV Co-authored-by: a32543254 <53296245+a32543254@users.noreply.github.com>
References
#2111 - Optimize INT4 MoE prefill/decode (sym) XPU kernels with dedicated w4a16 tiling
#2143 - feat: W4A8 ARK XPU MoE kernel (int4 weight / int8 compute) with prefill + decode
Author
Copilot
Parents
5b88c2ec
Loading