Logo Lanfrica

UBC-NLP/afroscope-data

Domain:

natural language processing

Record type:

dataset
Creator:
UBC
Host:
This is the corpus released as part of the AfroScope project for large-scale African language identification, supporting 713 languages. It provides sentence-level text with language labels and linguistic metadata and was used to train the open Afroscope-model. The source project and additional resources are openly available here: github.com sentence (string): The text sample.