Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
intel/auto-round
Pull Requests
Commits
fix_bug_1105
AutoAdamRound_bugfix
Chinesization
ZaneMark-patch-1
ZaneMark-patch-3
acp
actvation_quant
add_task_args_for_lmeval
ar_agent
ark_zp
autoround_support_qbits_backend
bf16_scale
chore/claude-init
copilot/add-xpu-moe-decode-implementation
copilot/convert-script-improvements
copilot/fix-corner-case-in-auto-round
copilot/fix-deprecated-fp-layers-handling
copilot/fix-docstrings-in-python-files
copilot/fix-issue-with-auto-rounding
copilot/fix-llm-type-70b-bits-setting
copilot/fix-typeerror-wrapped-fn
copilot/fix-vllm-model-inference-issue
copilot/improve-pr-template-type-of-change
copilot/investigate-quantization-group-and-ffn
copilot/replace-getset-module-torch-api
copilot/sageattention
copilot/speedup-fp8-linear-convert
copilot/speedup-fp8-linear-convert-again
copilot/speedup-fp8-linear-convert-another-one
copilot/sub-pr-1237-again
copilot/sub-pr-1237
copilot/sub-pr-1324
copilot/sub-pr-1522-again
copilot/sub-pr-1532
copilot/update-user-settings-page
copilot/vscode-mo3shmf8-8qa6
ddp
debug_time_cost
debug/usable_rotation
debug-hang
debug-nvfp4
deepseekv3
ds-qwen
ds-v5
ds-v32
dsv4
enable_glm4_moe_lite_quantization
enable_llama4_int8_baseline
enable_llama4_quant
enable_mxfp_exporting
feat/activation-checkpointing
feat/ark-xpu-int3-woq-gemv
feat/autoround-quarot
feature/overlap_for_nblocks
fix_bug0627
fix_bug_0722
fix_bug_1105
fix_compile
fix_disable_act_dynamic_usage_in_mxfp.py
fix_dq
fix/fp-layers-deprecation-mapping
fix_gemma3_issue
fix_gguf_fp8
fix_gptqmodel
fix_low_cpu
fix_rotation
fix_save_quantized_func_nvfp_checker
fix_0107
fix_0109
fix_0113
fix-attn-mask-b60
fix-ds
fix-flashinfer
fix-gpt-oss
fix-hpu
fix-to-meta-assertion-error-1499
fixbug_0717
fp4_v2
fp4_v3
fp8-cache
fp8-cache-based-export
fp8-static-quant-patch
fp8_export_backup_stable
fp8_export_for_test
good-flux
hengguo/fix_cuda_ut
hengguo/fix_gguf_ds
hengguo/quantizers
hengguo/refactor_algs
hengguo/refactor_init
hengguo/refactor_quant_step1
hengguo/smoothquant
hengguo/w4afp8_sim
henguo/update_so
hpu_only_kg
hpu_only_pkg
hpu/only/v1
hpu-limit-tran
kaihui/torch_dtype
lazy-model-replace
leq_opub
lib/pre-4.4.0
llama/new/9-610
llama/new/9
llm-main
llmc
llmc-backup
llmc-test
lm-head-quant
load-kv
load-w8a8-replace-mod
load-w8a8
lvl/autoscheme_ram_opt
lvl/cpu_ram_optimization
lvl/fix_no_init_weights
lvl/general_moe_replacement
lvl/support_fp8_with_ark
lvl/support_hunyuan_image
lvl/support_turbo_quant
lyt/numpy_fix
lyt/omni
main
marlin_modify
mengni/bug_fix
mengni/expert
mengni/mengni/block_wise
mengni/new_vllm
mengni/vllm
mengni/vlm
mengniwang95-patch-1
more-ar-ext
mxfp8
origin/block_wise
patch/for/ao/581/stable
patch-for-ao-2
pr1775-followup
pre-release/internal-inc/w4a8
quant-attn-hpu
quant-attn-hpu-o-scale
quant-attn-hpu-pr
quant-llama
quarot-llama
qwen3-vl
qwen3_vl_moe
qwen-split
qwen-v5
refine_device
refine_device_1
refine-doc-table
replace-lm-head
revert_order
revert-318-fix/hpu/check
revert-1231-set_disable_opt_rtn_default_2_none
revert-1562-suyue/ut
save_memory
set_disable_opt_rtn_default_2_none
static_quant
support_gemma4
test-git
try_new_optimizer
try_to_fix_hadamard_regression
update_fp_compile
update_0522
update_0819
upstream-ao
use-ep
ut-time
v0.7.0rc
v0.7.1rc
v0.8.0rc
v0.8.0rc2
v0.9.1rc
v0.9.2-release
v0.9.2rc
v0.9.3rc
v0.9.4rc
v0.9.5rc
v0.9.6rc
v0.9.7rc
v0.10.0rc
v0.10.1rc
v0.10.2rc
v0.10.3rc
v0.12.0rc
v0.12.1rc
v0.12.2rc
v0.12.3rc
v0.13.0rc
w4a4_int_quaro
w4int8dynamic
wangchang/fix_oom
wangchang/vllm
wfp8-afp8-bk
xin3he-patch-1
xinhe/3-20c
xinhe/3-27c
xinhe/3-27d
xinhe/4-7
xinhe/4-15
xinhe/5-11
xinhe/5-28
xuehao/test_gptq
zhenzhong/toolkit_release
fix bugs
wenhuach21
committed
1 year ago
c1d5dafa
use torch.compile by default for PyTorch versions 2.6 and above (#295)
wenhuach21
committed
1 year ago
Verified
c922f5b3
[Experimental Feature]support for common hf multimodel (#276)
n1ck-guo
committed
1 year ago
Verified
e6432125
fix bug of backend (#294)
wenhuach21
committed
1 year ago
Verified
4f228717
fix ipex tqdm mismatch issue (#293)
wenhuach21
committed
1 year ago
Verified
487abd6f
Add ipex support for intel cpu (#292)
wenhuach21
committed
1 year ago
Verified
168a1f69
Refine code (#291)
wenhuach21
committed
1 year ago
Verified
f41094a9
update readme (#287)
wenhuach21
committed
1 year ago
Verified
8efff6f4
fix mx_fp issues (#286)
wenhuach21
committed
1 year ago
Verified
99cff1fb
avoid deterministic algorithm warning in inference (#285)
wenhuach21
committed
1 year ago
Verified
ba5be40a
update readme for cpu inference
wenhuach21
committed
1 year ago
Verified
141c149f
update readme for v0.3.1 release (#283)
wenhuach21
committed
1 year ago
Verified
a3592220
refine AuoRound format and support marlin repacking (#280)
wenhuach21
committed
1 year ago
Verified
68138e82
qwen2_bugfix, add adamround vision UT (#281)
WeiweiZhang1
committed
1 year ago
Verified
7cfff967
refine eval (#282)
wenhuach21
committed
1 year ago
Verified
afa9e262
[Important update]set full range sym as the default (#278)
wenhuach21
committed
1 year ago
Verified
00122bc6
adamround bugfix, refine import (#275)
WeiweiZhang1
committed
1 year ago
Verified
a633aa70
change to even rounding for mantissa of mx_fp (#277)
wenhuach21
committed
1 year ago
Verified
98a9c755
fix mutable default value (#272)
wenhuach21
committed
1 year ago
Verified
8bf63f39
enable llama3.2-vision model quantization (#269)
WeiweiZhang1
committed
1 year ago
Verified
6b99d10a
keep the dtype after qdq (#268)
wenhuach21
committed
1 year ago
Verified
3a70be84
remove g_idx in gptq format (#267)
wenhuach21
committed
1 year ago
Verified
fdfd9711
Update readme for VLM support and integration (#266)
wenhuach21
committed
1 year ago
Verified
af3db170
Add a warning for improper export formats. (#265)
wenhuach21
committed
1 year ago
Verified
be32686b
Fix 3bit packing for auto-gptq format (#264)
wenhuach21
committed
1 year ago
Verified
6ee91a9f
better support quant_lm_head for larger models (#263)
wenhuach21
committed
1 year ago
Verified
82322ac9
refine autoawq exporting code (#261)
wenhuach21
committed
1 year ago
Verified
6539d506
update eval and fix example (#260)
n1ck-guo
committed
1 year ago
Verified
7816eea5
enable_qwen2-vl_quantization (#248)
WeiweiZhang1
committed
1 year ago
Verified
3275df93
fix preci (#258)
n1ck-guo
committed
1 year ago
Verified
200fcddc
Older