onnxruntime
[CUDA] Speed up the NVFP4 QMoE decode GEMV and enable it for MTP verify
#31159
Merged
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Overview
Commits
8
Changes
View On
GitHub
[CUDA] Speed up the NVFP4 QMoE decode GEMV and enable it for MTP verify
#31159
tianleiwu
merged 8 commits into
main
from
tlwu/20260730/nvfp4_moe_gemv
tianleiwu
force pushed
from
256c3155
to
d73cf351
6 days ago
tianleiwu
requested a review
from
copilot-pull-request-reviewer
6 days ago
copilot-pull-request-reviewer
commented on 2026-07-30
Base automatically changed from
tlwu/20260730/fp4_qmoe_no_weight_copy
to
main
6 days ago
tianleiwu
force pushed
from
d73cf351
to
4a22d73b
6 days ago
tianleiwu
marked this pull request as ready for review
6 days ago
github-actions
commented on 2026-08-01
tianleiwu
requested a review
from
hariharans29
4 days ago
tianleiwu
requested a review
from
kunal-vaishnavi
4 days ago
tianleiwu
requested a review
from
titaiwangms
4 days ago
titaiwangms
approved these changes on 2026-08-04
tianleiwu
force pushed
from
87b93927
to
e25d639b
11 hours ago
titaiwangms
approved these changes on 2026-08-05
NVFP4 GEMV Packed E2M1 Dequantize
1de0bd1e
ORT_FP4_GEMV_DEFAULT_TILING
0fff6146
Cut memory and ALU traffic in the NVFP4 MoE decode GEMV
bc37a129
[CUDA] QMoE NVFP4 GEMV: widen the profiled expanded-rows window for M…
62d644ba
fix(cuda): align GEMV tile arrays to their widest access
07cd3ffe
fix(cuda): preserve QMoE GEMV correctness
3b1d0463
lintrunner
80c89038
docs(cuda): correct GEMV bit-exactness claims and widen GEMV test cov…
43e7ccb4
tianleiwu
force pushed
from
e25d639b
to
43e7ccb4
11 hours ago
tianleiwu
merged
e56b1073
into main
6 hours ago
tianleiwu
deleted the tlwu/20260730/nvfp4_moe_gemv branch
6 hours ago
Login to write a write a comment.
Login via GitHub
Reviewers
titaiwangms
github-actions
copilot-pull-request-reviewer
hariharans29
kunal-vaishnavi
Assignees
No one assigned
Labels
None yet
Milestone
No milestone
Login to write a write a comment.
Login via GitHub