model: support arch `DbrxForCausalLM` #6515
model: dbrx convert to gguf
1d8de315
llama: support dbrx
ed582c1d
gguf-py: remove wrong clip -> clamp
3e3d2d12
model: dbrx, trust remote code
3937100a
llama: add label for model 132B
c0beb3cf
model: dbrx fix python linter in convert-hf-to-gguf.py
09210334
llama: support dbrx fix norm type
e4f8ee4f
dbrx: minor
a7f9a3ea
convert: dbrx: fix mixed up and down expert tensors
e3c1e812
convert: dbrx: fix mixed up and down expert tensors
0a35f588
doc: dbrx: add the model as supported
c8e6f903
convert: dbrx: fix remove wrong ATTN_OUT_NORM tensor, add output laye…
916b9185
llama: dbrx: remove wrong attn output layer in model arch
03da419f
scripts: get-wikitext-2 add unzip
76f266be
llama: dbrx: no attention output layer
9c7dedb0
model: dbrx: fix missing embedding tensor, mix with output layer
fe808987
llama: dbrx: remove not existing condition on empty output layer
4f12a580
slaren
commented
on 2024-04-06
Merge remote-tracking branch 'origin/master' into hp/model/support-dbrx
69856297
llama: dbrx: remove unnecessary optional tensor on FFN_GATE_EXPS
7e7cd53c
llama: increase maximum experts allowed
52c40335
model: dbrx: convert add n_ff
06a59abf
llama: dbrx: quantize fix n_attention_wv tensor name
305ac3b6
model: dbrx: convert fix tokenizer
b6522a9f
llama: dbrx: quantize fix n_attention_wv tensor name
dccb0126
model: convert-hf-to-gguf.py add _set_vocab_tiktoken gpt2 backed on l…
61be4b91
model: convert-hf-to-gguf.py fix classname conflict with qwen2
1fb6d95c
model: dbrx: convert-hf-to-gguf.py fix fix ftype missing, fix tensor …
200ce214
model: dbrx: convert-hf-to-gguf.py add chat template
9e17dad0
llama: quantize: remove wrong look for tensor qkv name as it was badl…
d7546fda
model: dbrx: convert-hf-to-gguf.py fix 'token_embd.weight' has wrong …
3a9dc2ee
model: dbrx: convert-hf-to-gguf.py support python 3.8
8154617f
llama: dbrx: no weight suffix in ffn_gate_exps, ffn_up_exps and ffn_d…
2449ef48
llama: quantize: remove wrong look for tensor qkv name as it was badl…
1bd94270
llama: dbrx: fix tensor qkv number of elements
e9987c66
model: dbrx: convert reshape expert tensors to 3D
d151d8fa
model: dbrx: convert experts to f16
f062b834
model: dbrx: fix tensor names mapping broken
dbfd5911
model: dbrx: fix expert reshape
7dd84b09
model: dbrx: fix expert reshape
c9bddbf2
model: dbrx: fix again sic expert reshape
e2c91996
model: dbrx: weird fix expert reshape
50b43736
llama: dbrx: output norm dim
0ab1bae8
llama: dbrx: fix last normalization
830e46d7
llama: dbrx: revert
2897aa62
llama: dbrx: move norm2 after attention, fix build kv
993f8360
llama: dbrx: fix build kv att out
b01b062a
llama: dbrx: fix build kv att out tensor name
74e6d876
llama: dbrx: hardcode nn.LayerNorm epsilon
f8f97e74
llama: dbrx: Try another rope type
71f9e479
llama: dbrx: fix k scale
52c6276e
llama: dbrx: move norm epsilon to convert. Fix missing normalization.
8e226884
llama: dbrx: rename tensor to actual meaning. Fix normalization in gr…
35dce3e1
llama: dbrx: convert remove previous reverse
506cc2ea
llama: dbrx: load norm eps in hparams
eb0847e6
llama: dbrx: fix experts tensor layout
81f308ad
model: dbrx: convert-hf-to-gguf.py fix experts tensors shapes
21fb24aa
llama: factorize moe graph implementation between grok, mixtral and dbrx
f20c04f0
model: dbrx convert permute experts directly torch, log shape
48909ed2
llama: dbrx: fix experts 3D tensor layout (again)
18a84fed
llama: dbrx: fix experts 3D tensor layout (again)
99689529
llama: dbrx: document changes, permute only FFN_DOWN_EXPS. Add a chec…
e66f1e34
llama: dbrx: rename layer_out_norm to attn_out_norm
f30a73bb
llama: dbrx: first add the residuals and then do the norm
ea8b58c6
model: dbrx: convert fix mixed ffn_gate_exps and ffn_down_exps
55943a28
llama: dbrx: fix ggml context of the attention outputs weight
c7b9a2e8
gguf-py: revert spaces
ac82aa0e
gguf-py: dbrx: reverse again the MOE tensors mapping:
ac75fbd8
Merge remote-tracking branch 'origin/master' into hp/model/support-dbrx
e5631cf2
Merge remote-tracking branch 'origin/master' into hp/model/support-dbrx
6f813dcc
llama: dbrx: use the MOE naming convention for model type
74529e54
Merge remote-tracking branch 'origin/master' into hp/model/support-dbrx
06527c66
model: convert-hf-to-gguf.py remove tiktoken
fc89feed
Is silu activation function applied to MODEL_TENSOR.FFN_GATE_EXP here…
bdc4efe1
Is silu activation function applied to MODEL_TENSOR.FFN_GATE_EXP here…
542585fb
Wrong input was being fed to moe layer. This needs to be corrected
ecbfb1b5
eval-callback: also print last n elements of each dimension
647a11b1
minor spaces
03bdc36e
convert: update comment of MOE tensors mapping
8e6758f2
phymbert
marked this pull request as ready for review 2 years ago
slaren
approved these changes
on 2024-04-12
llama: rename build_moe to build_moe_ffn and fix grok is using gelu i…
f1256dc8
convert-hf-to-gguf.py: fix python linter
e517585f
megha95
approved these changes
on 2024-04-12
ggerganov
approved these changes
on 2024-04-13
minor: fix indent in llama_build_graph
9f77484c
phymbert
merged
4bd0f93e
into master 2 years ago
phymbert
deleted the hp/model/support-dbrx branch 2 years ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub