auto-round
Optimize CPU RAM peak memory during quantization
#1386
Merged

Optimize CPU RAM peak memory during quantization #1386

chensuyue merged 78 commits into main from lvl/ram_usage_optimization
lvliang-intel
lvliang-intel Optimize CPU RAM peak memeory during quantization
dee1db76
lvliang-intel lvliang-intel requested a review from copilot-pull-request-reviewer copilot-pull-request-reviewer 177 days ago
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
459ee8ac
copilot-pull-request-reviewer
copilot-pull-request-reviewer commented on 2026-02-03
WeiweiZhang1 rm duplicate args of the quantization extra config (#1334)
2a78a189
wenhuach21 fix --device_map cuda and xpu issue (#1383)
e00c176c
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
d6d9f775
lvliang-intel refine test case
ca55ae84
n1ck-guo
n1ck-guo commented on 2026-02-04
yiliu30 Disable replace `FP8Expert` (#1379)
7a3dcacc
lvliang-intel Support general MOE replacement for MOE models (Transformers 5.0 comp…
082bf4ca
n1ck-guo fix cuda ut fail (#1370)
dd45c310
xin3he [Regression] Detach scale tensor to prevent holding computation graph…
10028e83
wenhuach21 fix layer config (#1373)
b2dff810
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
56148940
lvliang-intel update code for comments
a041da86
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
e62a7088
lvliang-intel
lvliang-intel support AutoScheme cpu ram optimization
5dcd064c
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
10f0a4a4
xin3he Refactor evaluation in tests to use evaluate_accuracy function (#1402)
13c140b8
xin3he fix bug in PR-1244 (#1422)
4daa3253
xin3he remove require_intel_extension_for_pytorch and fix TypeError: unhasha…
b32e9a13
Kaihui-intel Fix cuda model ut [glm4][Molmo] (#1428)
1e41b94c
xin3he Update BackendInfos for AutoGPTQ based on transformer version and sav…
accef777
xin3he fix CUDA CI (#1431)
b7a7a307
WeiweiZhang1 refine global scale calculation to blockwise (#1421)
b9e925f5
xin3he support GPTQ_FORMAT for "gptqmodel:exllamav2" backend (#1434)
34647113
mengniwang95 Update diffusion README (#1439)
564f9a1c
chensuyue Enable new cpu test pool (#1435)
8a45c86e
chensuyue Update auto-round-lib README.md (#1437)
5b98860a
xin3he update logic for compatibility with gptqmodel:exllamav2 backend (#1438)
b2a7b7f0
chensuyue update readme for Intel xpu usage (#1441)
60fe3d4c
xin3he fix bug of evaluate_accuracy (#1430)
48294a2e
chensuyue bump version (#1445)
4f8f3555
yiliu30 Attach `act_max_hook` for FP8 model (#1447)
32d3a30e
xin3he [Regression] fix FP8_STATIC loading (#1452)
359a4a76
timkpaine Fix requirements packaging for source distribution (#1455)
98351f9e
yiliu30 fix: preserve classmethod descriptor in from_pretrained monkey patch …
b6193867
yiliu30 Support load FP8 model on HPU (#1449)
a04a665c
XuehaoSun Update compatibility test and version to 0.10.2 (#1463)
5baa6e7f
wenhuach21 support multiple device evaluation for activation quantized model (#1…
36944704
lvliang-intel Fix KeyError when GPU is missing from accelerate max_memory (#1457)
2203d36d
wenhuach21 support glm5 (#1466)
d6ef4cec
yiliu30 Disable replace `FP8Expert` (#1379)
27a64224
lvliang-intel fix layer config (#1373)
4a133f20
lvliang-intel update code for comments
e497f4d9
lvliang-intel fix issue 1368
6066d838
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
581b4aea
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
88c60ec2
wenhuach21
wenhuach21 commented on 2026-02-28
wenhuach21
wenhuach21
wenhuach21
wenhuach21 commented on 2026-02-28
lvliang-intel refactor code according to comments
471931f2
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
05d06e10
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
341240cb
lvliang-intel refactor code according to comments
5c71a836
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
137b759c
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
3b818099
wenhuach21
wenhuach21 commented on 2026-03-02
wenhuach21
wenhuach21 commented on 2026-03-02
wenhuach21
wenhuach21 commented on 2026-03-02
lvliang-intel
wenhuach21
wenhuach21 commented on 2026-03-02
wenhuach21
wenhuach21 commented on 2026-03-02
chensuyue Grep core dumped issue in UT test (#1484)
1a4a9c68
n1ck-guo gguf better support for transformers5.0 and fix bug of Qwen3Next (#1474)
2de46ac2
lvliang-intel Optimize CPU RAM peak memeory during quantization
57664fb6
lvliang-intel Support general MOE replacement for MOE models (Transformers 5.0 comp…
454933c8
n1ck-guo fix cuda ut fail (#1370)
30593f1f
lvliang-intel update code for comments
484d7a9a
xin3he Update BackendInfos for AutoGPTQ based on transformer version and sav…
1d4d6ec6
lvliang-intel refactor code according to comments
cfdce02d
lvliang-intel refine code
337965fb
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
1fd7a1e0
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
7e202d74
lvliang-intel fix merge issue
0928e8e6
lvliang-intel fix merge issue
06f5c85c
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
a199843a
lvliang-intel Merge branch 'lvl/ram_usage_optimization' of https://github.com/intel…
89a44e8f
lvliang-intel
lvliang-intel lvliang-intel force pushed from 9718d683 to 89a44e8f 148 days ago
wenhuach21
wenhuach21 commented on 2026-03-04
wenhuach21
wenhuach21 commented on 2026-03-04
wenhuach21
wenhuach21 commented on 2026-03-04
wenhuach21
wenhuach21 commented on 2026-03-04
lvliang-intel fix ci issues
dd2e35a0
lvliang-intel refactor offload manager
a088346a
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
6c92c30e
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
fa186ff1
lvliang-intel Update base.py
31c4176e
wenhuach21
wenhuach21 commented on 2026-03-04
wenhuach21
wenhuach21 commented on 2026-03-04
wenhuach21
wenhuach21 commented on 2026-03-04
wenhuach21
wenhuach21 commented on 2026-03-04
wenhuach21
wenhuach21 commented on 2026-03-04
lvliang-intel update code for comments
e7c606bd
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
11909355
pre-commit-ci[bot] [pre-commit.ci] auto fixes from pre-commit.com hooks
1c0911e1
wenhuach21
wenhuach21 commented on 2026-03-04
lvliang-intel fix ci issues
cce2dcdd
lvliang-intel Merge branch 'main' into lvl/ram_usage_optimization
72f84334
lvliang-intel
lvliang-intel Merge branch 'main' of https://github.com/intel/auto-round into lvl/r…
876fa663
n1ck-guo
n1ck-guo approved these changes on 2026-03-10
n1ck-guo
chensuyue chensuyue merged 0d372151 into main 142 days ago
chensuyue chensuyue deleted the lvl/ram_usage_optimization branch 142 days ago

Login to write a write a comment.

Login via GitHub

Assignees
No one assigned
Labels
Milestone