optimum-habana-fork
Add llamafactory subtree
#342
Closed

Add llamafactory subtree #342

wenbinc-Bin wants to merge 2948 commits into HabanaAI:aice/v1.22.0 from wenbinc-Bin:llamafactory
wenbinc-Bin
hiyouga [ci] update workflow (#7255)
cdafa8a1
hiyouga [misc] upgrade format to py39 (#7256)
7c1640ed
hiyouga [misc] upgrade deps (#7257)
142fd7e7
emascarenhas This change adds support for Intel Gaudi HPUs.
58bec79c
hiyouga [model] support gemma3 (#7273)
165d3ed0
hiyouga [misc] update format (#7277)
9ccfb97a
BlackWingedKing [data] efficient 4d_attention_mask creation in neat_packing (#7272)
d7d79f7e
emascarenhas Delete examples of Gaudi yaml files
3074690a
emascarenhas Revert to original requirements.txt and capture transformer version f…
a0f16616
emascarenhas Revert transformers and other versioning
a16e3d47
hiyouga [assets] update video (#7287)
e9b427d5
hiyouga [assets] update wechat (#7288)
0dbce72f
felladrin [dataset] fix ultrachat_200k dataset (#7259)
3dff4ecc
hiyouga [data] gemma3 plugin pan and scan (#7294)
ef5f1c1d
Qiaolin-Yu [inference] support sglang backend (#7278)
30038d9c
isLinXu [model] support hunyuan 7b (#7317)
a71e6850
hiyouga [assets] update videos (#7340)
48a6584f
hiyouga [data] fix template (#7349)
1d2131e5
hiyouga [misc] set dev version (#7351)
c518146e
hiyouga [assets] update wechat (#7361)
4a5d0f0b
hiyouga [version] fix minicpmo (#7378)
555b71a1
erictang000 [3rdparty] fix redundant process group destroy for ray (#7395)
d8a5571b
guoquan [misc] fix sglang deps (#7432)
ebc989ad
emascarenhas Update docker loaded version to 1.20.0
c158c20b
hiyouga [deps] upgrade vllm to 0.8 (#7436)
dfbe1391
hiyouga [deps] upgrade transformers to 4.50.0 (#7437)
b1b78daf
SnowFox4004 [scripts] support compute score on vllm's predictions (#7419)
7d4dc25c
hiyouga [misc] fix license (#7440)
fbf49e25
hiyouga [misc] fix ci (#7441)
c841e921
rumichi2210 [docker] upgrade to torch 2.6 (#7442)
747e02d6
hiyouga [trainer] fix vlm loss for transformers 4.49 (#7448)
42e090d3
hiyouga [assets] fix gemma3 readme (#7449)
833edc7c
hiyouga [assets] update wechat (#7455)
9abee9cd
kennylam777 [misc] enable liger kernel for gemma3 (#7462)
59a56f72
eljandoubi [misc] enable liger kernel for gemma3 text and paligemma (#7466)
b6dc7e01
xiaosu-zhu [misc] update liger-kernel's monkey patch (#7453)
bc9ada9d
GuoCoder [model] fix lora on quant models (#7456)
b6d8749b
hiyouga [model] add qwen2vl 32b & upgrade peft (#7469)
59e12bff
x22x22 [trainer] fix wsd scheduler (#7304)
01166841
Xu-pixel [3rdparty] support swanlab lark notification (#7481)
f5473346
Kuangdd01 [data] fix pixtral plugin (#7505)
b00cb2ed
hiyouga [assets] update wechat (#7523)
49436e93
hiyouga [deps] pin pydantic to 2.10.6 (#7546)
468eea6f
Kuangdd01 [model] add Qwen2.5-Omni model (#7537)
185c76f6
hiyouga [data] fix qwen2.5 omni collator (#7553)
2d421c57
himalalps [trainer] new kto mismatch pair creation strategy (#7509)
6d6e0f44
aliencaocao [data] shard the dataset to allow multiprocessing when streaming is e…
5d1cc863
taoharry [webui] fix launch with proxy (#7332)
6faa6fb5
BlackWingedKing [data] specify position_ids in PackedSupervisedDatasetProcessor for n…
f06a74ad
ysjprojects [model] fix use_cache patching for gemma3 multimodal (#7500)
9deece1d
hiyouga [model] fix kv cache (#7564)
aaf2e6ba
hiyouga [infer] vllm video/audio inference (#7566)
903db098
gechengze [trainer] fix batch processing in PPO trainer (#7576)
11997593
Kuangdd01 [data] fix qwen2.5 omni plugin (#7573)
80f8d037
Kuangdd01 [data] fix qwen2.5 omni plugin (#7578)
32cb086b
hiyouga [assets] update wechat (#7594)
61b24c38
hiyouga [model] add llama4 (#7611)
6c200fd2
hiyouga [assets] update readme (#7612)
7e0cdb1a
hiyouga [misc] fix packing and eval plot (#7623)
5817cda3
adarshxs [sglang] support transformers 4.51.0 (#7639)
8ee26642
Shawn-Tao [trainer] fix key error (#7635)
8f5f4cc5
Kuangdd01 [data] Fix bugs of `use_audio_in_video` in Qwen2.5 Omni (#7638)
7d8bee96
hiyouga [assets] update readme (#7644)
39876b85
hiyouga [assets] update readme (#7654)
24cb8904
hiyouga [data] add coig-p dataset (#7657)
34fdabe0
jilongW [misc] fix cuda warn on intel GPU (#7655)
3bdc7e1e
danny980521 [bugfix] enable_gemma_liger_kernel (#7660)
ee840b4e
erictang000 [ray] allow for specifying ray.init kwargs (i.e. runtime_env) (#7647)
39c1e29e
erictang000 [data] support for specifying a dataset in cloud storage (#7567)
6c53471d
hiyouga [assets] update wechat (#7674)
11bcafd0
hiyouga [deps] fix uv conflicts (#7686)
60a84f66
zRzRzRzRzRzRzR [model] add GLM-4-0414 (#7695)
481ecbf9
hiyouga [deps] upgrade transformers (#7704)
1fd4d14f
hiyouga [misc] upgrade cli (#7714)
3ef36d00
hiyouga [misc] fix env vars (#7715)
3a13d2cd
Kuangdd01 [model] Support Kimi_VL thinking/instruct (#7719)
df8752e8
hiyouga [assets] update model readme (#7724)
ac8c6fdd
fluidnumericsJoe [docker] patch docker-rocm (#7725)
b5d667ce
hiyouga [deps] upgrade vllm (#7728)
0fe5631f
hiyouga [api] fix chat messages (#7732)
9b942110
codemayq [assets] wechat (#7740)
d07983dc
leo-pony [infer] support vllm-ascend (#7739)
e1fdd6e2
ENg-122 [misc] improve entrypoint (#7345)
85434005
Kuangdd01 [model] support intern-VL 2.5-3 series (#7258)
125513fa
hiyouga [infer] set env for vllm ascend (#7745)
48315528
hiyouga [breaking] bump transformers to 4.45.0 & improve ci (#7746)
0a0cfeb7
hiyouga [trainer] fix pt loss (#7748)
610f164c
emascarenhas Update transformers to 4.49.0
0ab20a87
emascarenhas Set compatible datasets version
46cf93fc
emascarenhas Merge branch 'main' into feature/hpu
512ac7f4
emascarenhas Allow peft to be version 0.15.1
c78a129b
emascarenhas Allow peft to be version 0.15.1
8ddd5274
emascarenhas new tests and fixes
38f46def
hiyouga [assets] update wechat (#7792)
a4455e30
AlphaBladez [misc] fix bug in constant (#7765)
ec7257e7
ddddng [model] fix gemma3 export (#7786)
b8cddbc7
flashJd [misc] fix new tokens adding (#7253)
1302ca39
GeoffreyChen777 [data] Fix wrong position ids with packed attention masks (#7754)
81768df0
hiyouga [parser] support omegaconf (#7793)
278df430
tianshijing [trainer] Add Muon Optimizer (#7749)
d128382d
hiyouga [example] add bash usage (#7794)
a62cba3d
hiyouga [data] improve mmplugin (#7795)
92101f34
hiyouga [trainer] support early stop (#7797)
7f3c31f6
Kuangdd01 [misc] update internvl constants (#7801)
c91165a5
Kuangdd01 [model] add arch check for InternVL (#7803)
d43013f1
hiyouga [assets] update model readme (#7804)
2b7d564e
Kuangdd01 [data] fix internvl plugin (#7817)
1dd67eb0
hiyouga [model] fix moe zero3 (#7826)
13444163
hiyouga Merge commit from fork
2989d392
hiyouga [model] fix vit gradient checkpointing (#7830)
4fbdc65f
hiyouga [assets] update wechat (#7840)
036a76e9
erictang000 [ray] add storage filesystem to ray config (#7854)
b4407e4b
Kuangdd01 fix attn patch for kimivl (#7867)
035e9803
hiyouga [data] fix minicpmo vllm infer (#7870)
fcca3b0b
zhaop-l [trainer] make projector trainable in freeze training (#7872)
1bd319d1
hiyouga [data] fix qwen2 omni plugin (#7875)
00b5c059
hiyouga [model] fix dsv3 leaf node (#7879)
1f338deb
Kuangdd01 [data] fix qwen2.5 omni template (#7883)
36947445
hiyouga [model] add qwen3 (#7885)
ae392e05
hiyouga [data] improve mm plugin (#7910)
77c569e0
hiyouga [data] replace eos token for base models (#7911)
c5b1d07e
hiyouga [data] add eval_on_each_dataset arg (#7912)
072bfe29
hiyouga [misc] fix uv (#7913)
a8430f42
hiyouga [data] optimize qwen3 loss computation (#7923)
d8295cd6
hiyouga [data] fix base plugin (#7924)
41ec9286
hiyouga [hparam] add enable think argument (#7928)
6a584b40
ericdachen [assets] Warp Support README Update (#7887)
75d7c35f
hiyouga [model] add qwen2 omni 3b (#7945)
52f25651
hiyouga [misc] fix qwen2 omni (#7962)
937447bd
Kuangdd01 [model] add mimo7b (#7946)
5ee9eb64
hiyouga [example] update examples (#7964)
c6bcca4c
hiyouga [misc] update liger kernel patch (#7966)
97e0a4cb
emascarenhas Merge changes from main
82811195
emascarenhas Changes from make style
67f88970
tpoisonooo [docs] add GraphGen (#7974)
f584db50
yunhao-tech [data] Avoid repetitive tool description warp (#8000)
865ac074
Kuangdd01 [scripts] add video params for vllm infer (#7992)
cef3a0b2
hiyouga [data] fix kimi vl template (#8015)
845af89e
emascarenhas Fix failing pytest 2 cases
e5c5b3f5
hiyouga [model] add seed coder and qwen3 quant models (#8039)
8d472c20
hiyouga [assets] update windows installation (#8042)
ab2c0511
hiyouga [assets] update wechat (#8057)
2b23c0a7
Shawn-Tao [infer] Modify vllm_infer.py to batch preprocess to avoid too much fi…
e8a18c17
erictang000 [data] support loading folder from remote (#8078)
130bfaf8
Kuangdd01 [data] add forward compatibility for video_utils in Transformers 4.52…
66f719dd
Kiko-RWan [infer] support lora adapter for SGLang backend (#8067)
820ed764
Wangbiao2 [misc] fix cli (#8095)
89a0d10c
SunnyHaze [trainer] fix KeyError at end of pretrain (#8099)
16e26236
hiyouga [doc] add no build isolation (#8103)
ed2f89ef
piamo [model] update rope kwargs for yarn (#8101)
a6f3adf9
hiyouga [model] switch to gptqmodel (#8108)
f3fd67a9
hiyouga [data] qwen3 fixes (#8109)
b83a38eb
hiyouga [assets] update readme (#8110)
f96c0858
hiyouga [data] llama3 multi tool support (#8124)
b3b2c9f1
hiyouga [deps] update to transformers 4.52 (#8125)
b0c8ba73
hiyouga [misc] update data readme (#8128)
763fbc29
Kuangdd01 [data] fix internvl plugin when using PIL images (#8129)
a9211a73
hiyouga [assets] update wechat (#8156)
52dead87
wangzhanxd [api] support repetition_penalty and align presence_penalty with Open…
c477ae64
akshatsehgal feat: add smollm support (#8050)
e6f45d69
hiyouga [deps] upgrade transformers (#8159)
dc8cca11
hiyouga [model] add smollm2 and medgemma (#8161)
f3a1dc84
hiyouga [webui] fix input args (#8162)
09436c1f
hiyouga [webui] add infer extra args (#8167)
16e1a509
hiyouga [assets] update docker files (#8176)
519ac928
hiyouga [webui] add extra args to export (#8178)
4ecf4dae
hiyouga [data] fix shared file system (#8179)
e542f957
hiyouga [assets] fix docker image (#8180)
07f79214
youngwookim [data] Reading files from cloud is broken (#8182) (#8183)
57c6e232
Muqi1029 [assets] fix incorrect user_tag in dataset_info.json to prevent skipp…
00c4988f
hiyouga [webui] fix skip args (#8195)
54ffd060
hiyouga [assets] update Dockerfile (#8201)
55d37dc4
yzoaim [workflow] auto push docker images (#8181)
73b12baa
hiyouga [assets] fix docker images (#8203)
a4048b7b
emascarenhas Merge branch 'main' into feature/hpu
b9175e42
emascarenhas Update data_args.py - fix typo
94981071
hiyouga [model] add deepseek 0528 models (#8215)
be02003d
emascarenhas Add support for Synapse 1.21.0
a7529961
Kuangdd01 [scripts] specify model class for qwen_omni merge (#8227)
e31afdfd
hiyouga [assets] update readme (#8235)
762c2d77
akshatsehgal [model] add smollm2 support (#8220)
21829b5e
hiyouga [deps] upgrade transformers to 4.52.4 (#8245)
fba9c9d9
Kuangdd01 [model] add MIMO_VL (#8249)
65aa86ed
hiyouga [assets] update wechat (#8270)
16a3f8a7
Zeyi-Lin [tracking] swanlab add llamafactory tag (#8258)
6cc247e8
Kuangdd01 [data] support nested images input for videos (#8264)
3425bc6e
hiyouga [assets] add icon (#8276)
e3d5e0fa
hiyouga [assets] update readme (#8288)
ee676d29
hiyouga [assets] update docker files (#8291)
81c4d9be
Kuangdd01 [script] add Script description for qwen_omni_merge (#8293)
53084247
hubutui [launcher] Add elastic and fault-tolerant training support (#8286)
83688b0b
hiyouga [assets] fix npu docker (#8298)
cecba57b
Kuangdd01 [tests] add visual model save test (#8248)
fcd86623
hiyouga [assets] update readme (#8303)
13fd4361
hiyouga [setup] fix uv (#8311)
f5f35664
hiyouga [data] fix empty template (#8312)
239ced07
Remorax [model] pushing FFT with unsloth (#8325)
d325a1a7
hiyouga [model] fix model generate (#8327)
7ecc2d46
hiyouga [assets] update wechat (#8328)
d8a5050c
MING-ZCH [assets] Add awesome works used LLaMA-Factory (#8333)
8fa55db1
Kuangdd01 [model] support Mistral3.1 small 2503 (#8335)
8ffe7daa
LDLINGLINGLING [model] support MiniCPM4 (#8314)
d39d3106
hiyouga [misc] tiny fixes (#8348)
5ed62a29
amangup [trainer] Add LD-DPO objective (#8362)
1cfe4291
hiyouga [assets] update wechat (#8385)
bb84c3c8
hiyouga [version] release v0.9.3 (#8386)
9f2f12b0
hiyouga [data] fix qwen2vl pos ids (#8387)
af2f75e6
hiyouga [model] fix vlm utils (#8388)
f3d144f0
hiyouga [ci] add docker version (#8390)
cabc9207
hiyouga [misc] set dev version (#8389)
ec04d7b8
hiyouga [assets] update readme (#8396)
0e1fea71
dhiaEddineRhaiem [model] add support for Falcon H1 (#8403)
0d7d0ea9
hiyouga [assets] update wechat (#8414)
3a119ed5
codemayq [assets] update wechat
8a3bddc7
Remorax [model] unsloth resume from checkpoint bug (#8423)
12215335
hiyouga [model] add kimi vl 2506 (#8432)
8ed085e4
Kuangdd01 [model] Add mistral-small 3.2 & kimi-dev (#8433)
fffa43be
hiyouga [misc] fix ci (#8441)
31b0787e
hiyouga [data] fix audio reader (#8448)
be27eae1
hiyouga [assets] update wechat (#8458)
6b46c8b6
hiyouga [model] do not force load processor (#8457)
b10333da
hiyouga [webui] upgrade webui and fix api (#8460)
ed57b7ba
hiyouga [assets] update readme (#8461)
7242caf0
hiyouga Merge commit from fork
bb7bf515
jiajunly [data] fix gemma2 eos token (#8480)
0a004904
Kuangdd01 [model] add GLM-4.1V (#8462)
0b188ca0
injaeryou [parser] update config loading to use OmegaConf #7793 (#8505)
ac6c93df
hiyouga [assets] update wechat (#8517)
544b7dc2
Kuangdd01 [model] add gemma3n (#8509)
c5a08291
Kuangdd01 [assets] update readme (#8519)
4465e434
hiyouga [assets] update readme (#8529)
dcd75e70
hiyouga [assets] update issue template (#8530)
e117e3c2
Zeyi-Lin [tracking] fix swanlab hparams (#8532)
8e7727f4
Remorax [model] add lora dropout to unsloth (#8548)
62c69436
wjunLu [ci] Add workflow for building NPU image (#8546)
d30cbcdf
Remorax Revert "[model] add lora dropout to unsloth" - requested feature alre…
c6290db1
[docs] add nvidia-container-toolkit to Linux Docker setup instruction…
1b549e31
hiyouga [assets] update wechat (#8565)
62bd2c80
hiyouga [deps] bump transformers to 4.49.0 (#8564)
58175836
hiyouga [webui] support other hub (#8567)
043103e1
hiyouga [webui] fix abort finish (#8569)
6a8d8882
Kuangdd01 [data] support glm4.1v video training (#8571)
766884fa
hiyouga [webui] fix elems (#8587)
cf1087d4
emascarenhas Fix conflicts before merge
b72bc09a
emascarenhas Updates to version numbers and fix typos
96ffafba
wenbinc-Bin Add flash attn and token padding to qwen3
66055c92
wenbinc-Bin Fix bug that dpo training fails with 16k seq-len
845fb080
wenbinc-Bin Compute reference first to save memory
0fe0acc0
wenbinc-Bin enable requires_grad=True to input when use gradient_checkpointing
089b8bbd
wenbinc-Bin Add 'LLaMA-Factory/' from commit '089b8bbd70e6b51bdd2d92e5f829ffcbf17…
2fe1b34d
wenbinc-Bin wenbinc-Bin closed this 353 days ago

Login to write a write a comment.

Login via GitHub

Reviewers
No reviews
Assignees
No one assigned
Labels
Milestone