Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
intel/auto-round
Pull Requests
Commits
yi/random-init-diffusion-handoff
AutoAdamRound_bugfix
Chinesization
ZaneMark-patch-1
ZaneMark-patch-3
acp
actvation_quant
add_task_args_for_lmeval
agent/fp8-per-head-kv-attn-merge
ar_agent
ark_v0.13.4
ark_zp
autoround_support_qbits_backend
bf16_scale
chore/claude-init
copilot/adapt-to-v5-chat-template
copilot/add-bandwidth-metrics
copilot/convert-script-improvements
copilot/copilotoptimize-int4-moe-performance
copilot/fix-corner-case-in-auto-round
copilot/fix-deprecated-fp-layers-handling
copilot/fix-docstrings-in-python-files
copilot/fix-issue-with-auto-rounding
copilot/fix-layer-config-priority
copilot/fix-llm-type-70b-bits-setting
copilot/fix-occasional-test-failure
copilot/fix-typeerror-wrapped-fn
copilot/fix-vllm-model-inference-issue
copilot/improve-pr-template-type-of-change
copilot/investigate-quantization-group-and-ffn
copilot/optimize-int4-moe-performance
copilot/optimize-xpu-moe-bf16-fp16
copilot/replace-getset-module-torch-api
copilot/sageattention
copilot/speedup-fp8-linear-convert
copilot/speedup-fp8-linear-convert-again
copilot/speedup-fp8-linear-convert-another-one
copilot/sub-pr-1237-again
copilot/sub-pr-1237
copilot/sub-pr-1324
copilot/sub-pr-1522-again
copilot/sub-pr-1532
copilot/update-phase-auto-dispatch-logic
copilot/update-user-settings-page
copilot/vscode-mo3shmf8-8qa6
cosmos
ddp
debug_time_cost
debug/usable_rotation
debug-hang
debug-nvfp4
deepseekv3
ds-qwen
ds-v5
ds-v32
dsv4
enable_glm4_moe_lite_quantization
enable_llama4_int8_baseline
enable_llama4_quant
enable_mxfp_exporting
feat/activation-checkpointing
feat/ark-xpu-int3-woq-gemm
feat/ark-xpu-int3-woq-gemv
feat/auto-code-calibration-dataset-v2
feat/autoround-quarot
feat/cosmos3-nano-quant
feat/fp8-per-head-kv-attn
feat/sparse-bf16-prefill-v2-bk-818
feat/sparse-bf16-prefill-v2
feature/overlap_for_nblocks
fix_bug0627
fix_bug_0722
fix_bug_1105
fix_compile
fix_disable_act_dynamic_usage_in_mxfp.py
fix_dq
fix/fp-layers-deprecation-mapping
fix_gemma3_issue
fix_gguf_fp8
fix_gptqmodel
fix_int4_acc
fix/llm-compressor-mixed-targets
fix/llm-compressor-targets-main
fix/llm-compressor-targets-only
fix/llmc-ut-skips
fix_low_cpu
fix_rotation
fix_save_quantized_func_nvfp_checker
fix/sycl-tla-memory-budget
fix_0107
fix_0109
fix_0113
fix-attn-mask-b60
fix-ds
fix-flashinfer
fix-gpt-oss
fix-hpu
fix-to-meta-assertion-error-1499
fixbug_0717
fp4_v2
fp4_v3
fp8-cache
fp8-cache-based-export
fp8-static-quant-patch
fp8_export_backup_stable
fp8_export_for_test
good-flux
hengguo/fix_cuda_ut
hengguo/fix_gguf_ds
hengguo/quantizers
hengguo/refactor_algs
hengguo/refactor_cli
hengguo/refactor_init
hengguo/refactor_quant_step1
hengguo/refactor_uts
hengguo/smoothquant
hengguo/update_doc_706
hengguo/w4afp8_sim
henguo/update_so
hpu_only_kg
hpu_only_pkg
hpu/only/v1
hpu-limit-tran
kaihui/torch_dtype
lazy-model-replace
leq_opub
lib/pre-4.4.0
llama/new/9-610
llama/new/9
llm-main
llmc
llmc-backup
llmc-test
lm-head-quant
load-kv
load-w8a8-replace-mod
load-w8a8
lvl/autoscheme_ram_opt
lvl/cpu_ram_optimization
lvl/fix_no_init_weights
lvl/general_moe_replacement
lvl/sa_fp8
lvl/support_diffusiongemma
lvl/support_fp8_with_ark
lvl/support_hunyuan_image
lvl/support_turbo_quant
lyt/numpy_fix
lyt/omni
main
marlin_modify
mengni/bug_fix
mengni/expert
mengni/mengni/block_wise
mengni/vlm
mengniwang95-patch-1
minimax-h3-w4a16-clean
more-ar-ext
mxfp8
opt_moe_kernel
origin/block_wise
patch/for/ao/581/stable
patch-for-ao-2
pr1775-followup
pre-release/internal-inc/w4a8
quant-attn-hpu
quant-attn-hpu-o-scale
quant-attn-hpu-pr
quant-llama
quarot-llama
qwen3-vl
qwen3_vl_moe
qwen-split
qwen-v5
refine_calib
refine_device_tmp
refine_device
refine_dl
refine_quantizer2
refine-doc-table
replace-lm-head
revert_order
revert-318-fix/hpu/check
revert-1231-set_disable_opt_rtn_default_2_none
revert-1562-suyue/ut
revert-2113-suyue/xpu
save_memory
scheme_awq_bk
set_disable_opt_rtn_default_2_none
sparse-attn
sparse-attn-clean
sparse-attn-prefill-clean
sparse-attn-v0
static_quant
support_gemma4
suyue/ai4ci
suyue/ci
suyue/ci-bak
suyue/hpu
suyue/xpu
test-git
try_new_optimizer
try_to_fix_hadamard_regression
update_fp_compile
update_0522
update_0819
upstream-ao
use-ep
ut-time
v0.7.0rc
v0.7.1rc
v0.8.0rc
v0.8.0rc2
v0.9.1rc
v0.9.2-release
v0.9.2rc
v0.9.3rc
v0.9.4rc
v0.9.5rc
v0.9.6rc
v0.9.7rc
v0.10.0rc
v0.10.1rc
v0.10.2rc
v0.10.3rc
v0.12.0rc
v0.12.1rc
v0.12.2rc
v0.12.3rc
v0.13.0rc
v0.13.1rc
v0.14.0rc
v0.14.1rc
v0.14.2rc
v0.15.0rc
vllm-sharing-deck-2026
w4a4_int_quaro
w4int8dynamic
wangchang/agent
wangchang/fix_oom
wangchang/vllm_loading
wangchang/vllm
wangchang/vllmmixin
wfp8-afp8-bk
xinhe/8-14a
xinhe/8-19
yi/random-init-diffusion-handoff
zhenzhong/toolkit_release
zhenzhong/woqgemm_s8_update
package native MiniMax H3 W4A16 FL2VA
yiliu30
committed
2 days ago
9e702c3c
test: add MiniMax-H3 packed MXFP8 Diffusers smoke
yiliu30
committed
7 days ago
2f747bca
feat: support MiniMax-H3 modular diffusion handoff
yiliu30
committed
7 days ago
789721d9
Add random-init diffusion quantization path
Yi Liu
committed
13 days ago
9b02b5c3
Enable torch.compile for selected tests and update related configurations (#2124)
xin3he
committed
14 days ago
Verified
a2f6b316
support fused MoE quantization and export for GGUF format (#2072)
n1ck-guo
committed
14 days ago
Verified
1fca0c74
refactor scheme,layer_config and formats (#2013)
n1ck-guo
committed
14 days ago
Verified
77e8ea59
Automatically select calibration datasets for code models (#2107)
changwangss
committed
15 days ago
Verified
f8de00c3
disable parallel for non low_gpu_mem_usage (#2116)
xin3he
committed
16 days ago
Verified
417d87de
add opt_rtn support for mxfp4 and nvfp4 (#2100)
wenhuach21
committed
16 days ago
Verified
41cca72f
Fix: temporarily skip tests due to known issues with xfail annotations (#2113)
chensuyue
committed
19 days ago
Verified
60b813cb
Enable AutoScheme parallel and disk stream by default (#2105)
xin3he
committed
19 days ago
Verified
8a9a1d13
Fix: skip exporting fp8 attention q_max (#2103)
yiliu30
committed
19 days ago
Verified
2a267e72
fix MXFP model free auto-round format bug (#2106)
xin3he
committed
19 days ago
Verified
378187c7
[Feature] Register_block_output to support GLM5.2 tuning (#2104)
xin3he
committed
19 days ago
Verified
d05fb638
feat: add OpenCodeInstruct calibration dataset (#2096)
changwangss
committed
20 days ago
Verified
ca9524ae
Enable torch.compile by default and update related documentation (#2070)
xin3he
committed
20 days ago
Verified
be862d37
Stream AutoScheme's per-layer sensitivity scoring block-by-block (#2063)
aquilarubra
committed
20 days ago
Verified
73661edb
Fix dataset and batch_size not taking effect in CLI (#2102)
changwangss
committed
20 days ago
Verified
212c8bda
Remove RoutedExperts since the bug is fixed. (#2098)
xin3he
committed
21 days ago
Verified
8b0923c8
refine alg and support lfq (#2082)
wenhuach21
committed
21 days ago
Verified
f1af646b
ark: add SpargeAttn support (#2027)
yiliu30
committed
21 days ago
Verified
5ce9733c
Decoupling oneDNN (#2049)
Zhenzhong1
committed
22 days ago
Verified
54b4b84c
[Feature]: speedup autoscheme with multi-process and cache files (#2083)
xin3he
committed
22 days ago
Verified
288f848b
fix gguf format gemma4 High RAM and glm-4.7 error (#2084)
n1ck-guo
committed
22 days ago
Verified
0212c163
Temporarily skip the issue in the CT (#2093)
xin3he
committed
23 days ago
Verified
33ec4c75
Fix Flux diffusion state propagation during block replay (#2086)
changwangss
committed
26 days ago
Verified
3cae4b68
[high-risk][Step 1]Refactor compressor and quantizer (#2039)
wenhuach21
committed
27 days ago
Verified
8f8c8e9d
Add CUDA CI coverage (#2074)
XuehaoSun
committed
27 days ago
Verified
6ff414b1
fix: AutoScheme embedding budget and gguf mixed-bits config for multimodal models (#2055)
n1ck-guo
committed
30 days ago
Verified
06c33eae
Older