A curated benchmark dataset for Arabic dialect identification, with special focus on Sudanese Arabic
This repository contains a limited sample subset of an organic Sudanese Arabic dataset.
Evaluating multilingual MT systems Notes / challenges: Used in Masakhane benchmarking; supports low
Facebook Low Resource (FLoRes) MT Benchmark # FLORES-200 and NLLB Professionally Translated Dataset
FLORES+ dev and devtest set in Emakhuwa CC-BY-SA-4.0 @inproceedings{ali-etal-2024-expanding, title
A machine translation benchmark for low-resource and multilingual machine translation.