Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

NLLB: No Language Left Behind

Domaine:

natural language processing

Type de record:

projectmodel
No Language Left Behind (NLLB) is a first-of-its-kind, AI breakthrough project that open-sources models capable of delivering high-quality translations directly between any pair of 200+ languages — including low-resource languages like Asturian, Luganda, Urdu and more. It aims to help people communicate with anyone, anywhere, regardless of their language preferences. To enable the community to leverage and build on top of NLLB, we open source all our evaluation benchmarks(FLORES-200, NLLB-MD, Toxicity-200), LID models and training code, LASER3 encoders, data mining code, MMT training and inference code and our final NLLB-200 models and their smaller distilled versions, for easier use and adoption by the research community. This code repository contains instructions to get the datasets, optimized training and inference code for MMT models, training code for LASER3 encoders as well as instructions for downloading and using the final large NLLB-200 model and the smaller distilled models. In addition to supporting more than 200x200 translation directions, we also provide reliable evaluations of our model on all possible translation directions on the FLORES-200 benchmark. By open-sourcing our code, models and evaluations, we hope to foster even more research in low-resource languages leading to further improvements in the quality of low-resource translation through contributions from the research community.

Visit

github.com

Connected records

paperdatasetpaper

Tasks

machine translation

Languages

AfrikaansAkaAkanAmazighAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenBamanankanBembaBwamu, Cwi+56

Tags

nllb

Licenses

https://github.com/facebookresearch/fairseq/blob/nllb/LICENSE.model.md

Similaires

No Language Left Behind : DataNo Student Left BehindNo Language Left Behind Seed Data (Tamasheq (Tifinagh script))No Language Left Behind Seed Data (Standard Moroccan Tamazight)No Language Left Behind Seed Data (Tamasheq (Latin script))No Language Left Behind: Scaling Human-Centered Machine Translation

No Language Left Behind : Data

NLLB project uses data from three sources : public bitext, mined bitext and data generated using backtranslation. Details of different datasets used and open source links are provided in details here.

No Student Left Behind

The end of 2019 was punctuated by the emergence of an infectious disease spread through human-to-hum

No Language Left Behind Seed Data (Tamasheq (Tifinagh script))

No Language Left Behind Seed Data (Standard Moroccan Tamazight)

No Language Left Behind Seed Data (Tamasheq (Latin script))

No Language Left Behind: Scaling Human-Centered Machine Translation

Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset of languages, leaving behind the va