CUDA: Improve NVFP4 W4A4 activation quantization #25730
Squash history before conflict-resolution during rebase on master
46caff64
ORippler
force pushed
from
5d70e935
to
46caff64
25 days ago
compiler massaging to avoid unnecessary LDCs
8a6ca53c
am17an
approved these changes
on 2026-07-21
kvalues_mxfp4 -> kvalues_nvfp4 in quantize_mmq_nvfp4
b1ea6058
Always pass in src1_scale.ptr
598f69b6
Extract ggml_cuda_is_aligned helper
1394d6ee
taronaeo
approved these changes
on 2026-07-22
ORippler
merged
1a064ab0
into master 18 days ago
ORippler
deleted the osimons/nvfp4_amax_quant branch 18 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub