Logo Lanfrica

No Language Left Behind : Data

Domain:

natural language processing

Record type:

dataset
NLLB project uses data from three sources : public bitext, mined bitext and data generated using backtranslation. Details of different datasets used and open source links are provided in details here.