[Experimental] Add NVFP4 E5M3 support and related components for model-free quantization #2123
Add NVFP4 E5M3 support and related components for model-free quantiza…
cfc5eadb
xin3he
force pushed
from
9acf8168
to
cfc5eadb
52 days ago
[llm_compressor format] Add NVFP4 E5M3 quantization support and relat…
cc0e5feb
xin3he
force pushed
from
67b2362a
to
cc0e5feb
52 days ago
[pre-commit.ci] auto fixes from pre-commit.com hooks
47bfc852
fix CI
89dbe07e
Add NVFP4 E5M3 support with CuTe QDQ integration and related tests
814659b1
[auto_round format] Add NVFP4 E5M3 support with CuTe integration and …
fa58ab62
Enhance model conversion and CUDA integration for NVFP4 E5M3 support
e7e2f015
Add weight dequantization support and caching mechanism for NVFP4 E5M3
27cc2654
Merge remote-tracking branch 'origin/main' into xinhe/8-4
fa73368b
Implement hydration of scale tensors from sibling shards in FP8 dequa…
1f16c07e
Add logging for fp32 to UE8M0 conversion in scale tensor processing
52adf875
Enhance E8M0 block scale expansion for DeepSeek variants and improve …
5c0e8480
Normalize scale layout in _pack_weight_nvfp4_e5m3 for consistent 2D s…
8d1b51f8
Merge branch 'main' into xinhe/8-4
15b99fae
Add support for mixed precision quantization with MXFP and NVFP4_E5M3…
3cc36121
Add AR_NVFP4_E5M3_CACHE_HP_WEIGHT environment variable for caching hi…
6261b986
xin3he
added this to the 0.15.0 milestone 50 days ago
Implement streaming pipeline for shard processing with quantization a…
c5f66a50
Add NVFP4 support with normalization and passthrough handling for leg…
d11676f6
Fix tensor reshaping for NVFP4 source tensors in normalization and ha…
639b54d0
Merge branch 'main' into xinhe/8-4
2203adf7
Enhance quantization handling by adding NVFP4 support in _get_mxfp_gr…
979796f9
Normalize NVFP4 global scales to reciprocal form in _normalize_nvfp4_…
6942d535
Add NVFP4 quantization support with layer configuration checks in mod…
2b9947ea
Merge branch 'main' into xinhe/8-4
c5873505
xin3he
removed this from to the 0.15.0 milestone 46 days ago
xin3he
added this to the 0.16.0 milestone 46 days ago
Refactor NVFP4 E5M3 quantization to use "nvfp4_v2" data type
aa967f87
update per comments
dfe6c605
yiliu30
approved these changes
on 2026-08-11
Enhance model caching and warnings for NVFP4 E5M3 scheme in model_fre…
fe5e5976
Merge remote-tracking branch 'origin/main' into xinhe/8-4
caccbeae
Rename fp4_v2_with_global_scale to nvfp4_v2_with_global_scale in nvfp…
26a36021
Refactor code structure for improved readability and maintainability
b2926220
[pre-commit.ci] auto fixes from pre-commit.com hooks
1ce1f8eb
Update low precision tensor handling and adjust test timeout for auto…
69ae8d04
fix code scan
6c6276b7
Merge branch 'main' into xinhe/8-4
94879f4b
fix CI
81082135
Implement cross-shard FP8 scale handling and add unit tests for weigh…
d5a03327
Enhance cross-shard dependency management and logging for FP8 scale h…
fb3da9a4
Refactor scale name generation for consistency in FP8 tensor handling
1382f32f
Refine tensor handling by adding layer ignore checks and preserving o…
a42fd79d
Merge branch 'main' into xinhe/8-4
113a4cb6
Remove test for ignored layer preserving original FP8 tensors
67ecd8f1
xin3he
merged
0c0511b5
into main 38 days ago
xin3he
deleted the xinhe/8-4 branch 38 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub