Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

nllb-200-10M-sample

Domain:

natural language processing

Record type:

dataset
Creator:
slo
Host:
This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages.

Visit

huggingface.co

Tasks

machine translation

Languages

AkanAmazighAmharicBamanankanBembaChichewaChokweDholuoDinka, SouthwesternÉwé+38

Licenses

odc-by

Similar

morlayecis0003/nllb-200-nllb-francais-pulaarSakuzas/nllb-200-wolaytta_to_english_kumorlayecis0003/nllb-200-pulaar-mairieDavey117/nllb-200-stem-yorubaministercmanga/NLLB-200-fine-tuningvpermilp/nllb-200-1.3B-rust

morlayecis0003/nllb-200-nllb-francais-pulaar

Sakuzas/nllb-200-wolaytta_to_english_ku

morlayecis0003/nllb-200-pulaar-mairie

Davey117/nllb-200-stem-yoruba

ministercmanga/NLLB-200-fine-tuning

Fine-tuning a pre-trained NLLB-200 Large Language Model for translating South African languages Pro

vpermilp/nllb-200-1.3B-rust

This is the model card of NLLB-200's 1.3B variant. Here are the metrics for that particular checkpoi