llama.cpp
CUDA: Improve NVFP4 W4A4 activation quantization
#25730
Merged

CUDA: Improve NVFP4 W4A4 activation quantization #25730

ORippler
ORippler ORippler requested a review 25 days ago
github-actions github-actions added ggml
github-actions github-actions added CUDA
ORippler Squash history before conflict-resolution during rebase on master
46caff64
ORippler ORippler force pushed from 5d70e935 to 46caff64 25 days ago
ORippler
ORippler compiler massaging to avoid unnecessary LDCs
8a6ca53c
ORippler
am17an
am17an approved these changes on 2026-07-21
ORippler kvalues_mxfp4 -> kvalues_nvfp4 in quantize_mmq_nvfp4
b1ea6058
ORippler Always pass in src1_scale.ptr
598f69b6
ORippler Extract ggml_cuda_is_aligned helper
1394d6ee
ORippler
taronaeo
taronaeo approved these changes on 2026-07-22
ORippler ORippler merged 1a064ab0 into master 18 days ago
ORippler ORippler deleted the osimons/nvfp4_amax_quant branch 18 days ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone