Investigating an approach for low resource language dataset creation, curation and classification: Setswana and Sepedi
The recent advances in Natural Language Processing have been a boon for well-represented languages in terms of available curated data and research resources. One of the challenges for low-resourced languages is clear guidelines on the collection, curation and prepa