Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

nllb-200-10M-sample

Domaine:

natural language processing

Type de record:

dataset
Créateur:
slo
Hôte:
This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages.

Visit

huggingface.co

Tasks

machine translation

Languages

AkanAmazighAmharicBamanankanBembaChichewaChokweDholuoDinka, SouthwesternÉwé+38

Licenses

odc-by

Similaires

morlayecis0003/nllb-200-nllb-francais-pulaarSakuzas/nllb-200-wolaytta_to_english_kumorlayecis0003/nllb-200-pulaar-mairieDavey117/nllb-200-stem-yorubaministercmanga/NLLB-200-fine-tuningvpermilp/nllb-200-1.3B-rust

morlayecis0003/nllb-200-nllb-francais-pulaar

Sakuzas/nllb-200-wolaytta_to_english_ku

morlayecis0003/nllb-200-pulaar-mairie

Davey117/nllb-200-stem-yoruba

ministercmanga/NLLB-200-fine-tuning

Fine-tuning a pre-trained NLLB-200 Large Language Model for translating South African languages Pro

vpermilp/nllb-200-1.3B-rust

This is the model card of NLLB-200's 1.3B variant. Here are the metrics for that particular checkpoi