This is a sample of nearly 10M sentence pairs from the NLLB-200
mined dataset allenai/nllb,
scored with the model facebook/blaser-2.0-qe
described in the SeamlessM4T paper.
The sample is not random; instead, we just took the top n sentence pairs from each translation direction.
The number n was computed with the goal of upsamping the directions that contain underrepresented languages.