llama.cpp
213701b5 - Detokenizer fixes (#8039)

Commit

349 days ago

Detokenizer fixes (#8039) * Add llama_detokenize(): - Update header files location - UNKNOWN and CONTROL are 'special pieces' - Remove space after UNKNOWN and CONTROL - Refactor llama_token_to_piece() - Add flag: clean_up_tokenization_spaces - Symmetric params for llama_tokenize() and llama_detokenize() * Update and fix tokenizer tests: - Using llama_detokenize() - Unexpected vocab type as test fail instead of error - Useful when automating tests: - If you don't know in advance the vocab type - Differenciate other loading errors - Skip unicode surrogaes and undefined - Gracefully exit threads - Using exit() is throwing random exceptions - Clean old known problematic codepoints - Minor: confusing hexadecimal codepoint * Update bruteforce random tests - Add detokenizer checks - New generator: ascii_lr_strip - New generator: apostrophe - Add more vocabs files - Detokenize special tokens. - Replace errors with '\uFFFD' when detokenizing to 'utf-8' - More edge cases - Better detokenization results check * Fix add_space_prefix, set false by default * Better leading space removal * Do not remove space when decoding special tokens * Bugfix: custom regexs splits undefined unicode codepoints * 'viking' detokenizer clean spaces

References

#8039 - Detokenizer fixes

Author

jaime-m-p

Parents

be20e7f4

Files11

common
- common.cpp
- common.h
examples
- batched.swift/Sources
  - main.swift
- llama.swiftui/llama.cpp.swift
  - LibLlama.swift
include
- llama.h
src
- llama.cpp
- unicode.cpp
tests
- test-tokenizer-0.cpp
- test-tokenizer-1-bpe.cpp
- test-tokenizer-1-spm.cpp
- test-tokenizer-random.py

llama.cpp 213701b5 - Detokenizer fixes (#8039)

llama.cpp
213701b5 - Detokenizer fixes (#8039)