This a code to train a bpe tokenizer using the amharic alphabets and amharic letters. The trained to
Syllable-aware BPE tokenizer for the Amharic language (አማርኛ) – fast, accurate, trainable. # Amharic
This is a simple script that split an Amharic document into different sentences and tokenes. If you find an issue, please let us know in the GitHub (https://github.com/uhh-lt/amharicprocessor/issues)