Multi-lingual language models (LM), such as mBERT, XLM-R, mT5, mBART, have been remarkably successful in enabling natural language tasks in low-resource languages through cross-lingual transfer from high-resource ones. In this work, we try to better understand how such models, specifically mT5, transfer *any* linguistic and semantic knowledge across languages, even though no explicit cross-lingual signals are provided during pre-training. Rather, only unannotated texts from each language are presented to the model separately and independently of one another, and the model appears to implicitly
Research goal: What is the impact of varying the pre-training data size for each language in mT5 on cross-lingual transfer performance in the XTREME-R benchmark, measured by accuracy and language-specific F1 scores?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.2/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.2/10.