Logo Lanfrica

Yoruba Language

Domaine:

natural language processing

Type de record:

dataset
- Training word embeddings (FastText, BERT) - Comparing curated vs. massive embeddings - Language modeling for low-resourced African languages Notes / challenges: Focus on Yoruba language, Contains both clean texts (with proper diacritics) and noisy texts (incorrect or missing diacritics)