Optimize CPU RAM peak memory during quantization #1386
Optimize CPU RAM peak memeory during quantization
dee1db76
Merge branch 'main' into lvl/ram_usage_optimization
459ee8ac
rm duplicate args of the quantization extra config (#1334)
2a78a189
fix --device_map cuda and xpu issue (#1383)
e00c176c
[pre-commit.ci] auto fixes from pre-commit.com hooks
d6d9f775
refine test case
ca55ae84
Disable replace `FP8Expert` (#1379)
7a3dcacc
Support general MOE replacement for MOE models (Transformers 5.0 comp…
082bf4ca
fix cuda ut fail (#1370)
dd45c310
[Regression] Detach scale tensor to prevent holding computation graph…
10028e83
fix layer config (#1373)
b2dff810
Merge branch 'main' into lvl/ram_usage_optimization
56148940
update code for comments
a041da86
[pre-commit.ci] auto fixes from pre-commit.com hooks
e62a7088
support AutoScheme cpu ram optimization
5dcd064c
Merge branch 'main' into lvl/ram_usage_optimization
10f0a4a4
Refactor evaluation in tests to use evaluate_accuracy function (#1402)
13c140b8
fix bug in PR-1244 (#1422)
4daa3253
remove require_intel_extension_for_pytorch and fix TypeError: unhasha…
b32e9a13
Fix cuda model ut [glm4][Molmo] (#1428)
1e41b94c
Update BackendInfos for AutoGPTQ based on transformer version and sav…
accef777
fix CUDA CI (#1431)
b7a7a307
refine global scale calculation to blockwise (#1421)
b9e925f5
support GPTQ_FORMAT for "gptqmodel:exllamav2" backend (#1434)
34647113
Update diffusion README (#1439)
564f9a1c
Enable new cpu test pool (#1435)
8a45c86e
Update auto-round-lib README.md (#1437)
5b98860a
update logic for compatibility with gptqmodel:exllamav2 backend (#1438)
b2a7b7f0
update readme for Intel xpu usage (#1441)
60fe3d4c
fix bug of evaluate_accuracy (#1430)
48294a2e
bump version (#1445)
4f8f3555
Attach `act_max_hook` for FP8 model (#1447)
32d3a30e
[Regression] fix FP8_STATIC loading (#1452)
359a4a76
Fix requirements packaging for source distribution (#1455)
98351f9e
fix: preserve classmethod descriptor in from_pretrained monkey patch …
b6193867
Support load FP8 model on HPU (#1449)
a04a665c
Update compatibility test and version to 0.10.2 (#1463)
5baa6e7f
support multiple device evaluation for activation quantized model (#1…
36944704
Fix KeyError when GPU is missing from accelerate max_memory (#1457)
2203d36d
support glm5 (#1466)
d6ef4cec
Disable replace `FP8Expert` (#1379)
27a64224
fix layer config (#1373)
4a133f20
update code for comments
e497f4d9
fix issue 1368
6066d838
Merge branch 'main' into lvl/ram_usage_optimization
581b4aea
[pre-commit.ci] auto fixes from pre-commit.com hooks
88c60ec2
refactor code according to comments
471931f2
[pre-commit.ci] auto fixes from pre-commit.com hooks
05d06e10
Merge branch 'main' into lvl/ram_usage_optimization
341240cb
refactor code according to comments
5c71a836
[pre-commit.ci] auto fixes from pre-commit.com hooks
137b759c
Merge branch 'main' into lvl/ram_usage_optimization
3b818099
Grep core dumped issue in UT test (#1484)
1a4a9c68
gguf better support for transformers5.0 and fix bug of Qwen3Next (#1474)
2de46ac2
Optimize CPU RAM peak memeory during quantization
57664fb6
Support general MOE replacement for MOE models (Transformers 5.0 comp…
454933c8
fix cuda ut fail (#1370)
30593f1f
update code for comments
484d7a9a
Update BackendInfos for AutoGPTQ based on transformer version and sav…
1d4d6ec6
refactor code according to comments
cfdce02d
refine code
337965fb
[pre-commit.ci] auto fixes from pre-commit.com hooks
1fd7a1e0
Merge branch 'main' into lvl/ram_usage_optimization
7e202d74
fix merge issue
0928e8e6
fix merge issue
06f5c85c
[pre-commit.ci] auto fixes from pre-commit.com hooks
a199843a
Merge branch 'lvl/ram_usage_optimization' of https://github.com/intel…
89a44e8f
lvliang-intel
force pushed
from
9718d683
to
89a44e8f
148 days ago
fix ci issues
dd2e35a0
refactor offload manager
a088346a
[pre-commit.ci] auto fixes from pre-commit.com hooks
6c92c30e
Merge branch 'main' into lvl/ram_usage_optimization
fa186ff1
Update base.py
31c4176e
update code for comments
e7c606bd
Merge branch 'main' into lvl/ram_usage_optimization
11909355
[pre-commit.ci] auto fixes from pre-commit.com hooks
1c0911e1
fix ci issues
cce2dcdd
Merge branch 'main' into lvl/ram_usage_optimization
72f84334
Merge branch 'main' of https://github.com/intel/auto-round into lvl/r…
876fa663
n1ck-guo
approved these changes
on 2026-03-10
chensuyue
merged
0d372151
into main 142 days ago
chensuyue
deleted the lvl/ram_usage_optimization branch 142 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub