Logo Lanfrica

olgapelloni/subword_evenness

Domain:

natural language processing

Record type:

paper
Creator:
olg
Host:
Repository for the paper: Olga Pelloni, Anastassia Shaitarova and Tanja Samardzic (2022). Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource Languages, EMNLP 2022. # Subword Evenness (SuE) Repository for the paper: Olga Pelloni, Anastassia Shaitarova and Tanja Samardzic (2022). Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource Languages, EMNLP 2022. ## Data Data comes from the TeDDi Sample corpus. We create 1M tokens balanced datasets for 19 training languages and 200K tokens datasets for 30 test languages. Links to our datasets: - Train dataset - Valid dataset - Test dataset ## Scripts Most of the scripts are done by Anastassia Shaitarova and me. Scripts for continuous training/fine-tuning are taken and adapted from HuggingFace. Scripts for measuring TTR and unigram entropy come from Ximena Gutierrez-Vasques. The number of BPE merges is calculated following the minimum redundancy approach (Gutierrez-Vasques et al. 2021) and using the scripts from the paper repository. The resulting numbers of merges used for monolingual training on the 19 transfer languages are listed in the file ```measures_scripts/sue/num_bpe_merges.tsv```