Major advancement in the performance of machine translation models has been made possible in part th
BibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-
The Spoken Portuguese corpus was collected among sociolinguistically diverse speakers having Portugu
The first publicly available sentence-level parallel corpus for Portuguese and Changana (Xichangana/
As part of the Open Language Data Initiative shared tasks, we have expanded the FLORES+ evaluation s