Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Lucie-Nek/Explore-Academy-AI-South-African-Language-Identification-Hack-2023

Domaine:

natural language processing

Type de record:

datasetproject
Créateur:
Luc
Hôte:
# Explore-Academy-AI-South-African-Language-Identification-Hack-2023 **Overview** South Africa is a multicultural society that is characterised by its rich linguistic diversity. Language is an indispensable tool that can be used to deepen democracy and also contribute to the social, cultural, intellectual, economic, and political life of the South African society. The country is multilingual with 11 official languages, each of which is guaranteed equal status. Most South Africans are multilingual and able to speak at least two or more of the official languages. From South African Government. With such a multilingual population, it is only obvious that our systems and devices also communicate in multi-languages. In this challenge, you will take text which is in any of South Africa's 11 Official languages and identify which language the text is in. This is an example of NLP's Language Identification, the task of determining the natural language that a piece of text is written in. **NB**: The dataset used for this challenge is the NCHLT Text Corpora collected by the South African Department of Arts and Culture & Centre for Text Technology (CTexT, North-West University, South Africa). The training set was improved through additional cleaning done by Praekelt. The data is in the form Language ID, Text. The text is in various states of cleanliness. Some NLP techniques will be necessary to clean up the data. **File descriptions** train_set.csv - the training set test_set.csv - the test set sample_submission.csv -> a sample submission file in the correct format **Language IDs** afr - Afrikaans eng - English nbl - isiNdebele nso - Sepedi sot - Sesotho ssw - siSwati tsn - Setswana tso - Xitsonga ven - Tshivenda xho - isiXhosa zul - isiZulu **Rules for the hackathon** This is a private hackathon whose primary purpose is for the members of the EXPLORE Data Science Academy to apply what they have learned. If you are part of EDSA contact your mentor for the secret …

Visit

github.com

Tasks

language identification

Languages

AfrikaansNdebeleNdebeleSetswanaSotho, NorthernSotho, SouthernSwatiTsongaXhosaZulu

Similaires

Ifeoluwa13/south-african-language-identification-hack-2023setshabapm/south-african-language-identification-hack-2023Buckweed2020/South-African-Language-Identification-Hack-2023Zenani99/South-African-Language-Identification-Hack-2023Maanzak/South-African-Language-Identification-Hack-2023Thato-rabodiba/South-African-Language-Identification-Hack-2023

Ifeoluwa13/south-african-language-identification-hack-2023

In this challenge, we will take text which is in any of South Africa's 11 Official languages and ide

setshabapm/south-african-language-identification-hack-2023

Setshaba Mashigo ExploreAI language classification hackathon submission # South African Language Id

Buckweed2020/South-African-Language-Identification-Hack-2023

# Language Identification in South African Texts This Python notebook aims to develop a language cl

Zenani99/South-African-Language-Identification-Hack-2023

NLP project # South-African-Language-Identification-Hack-2023 NLP project Description ExploreAI Aca

Maanzak/South-African-Language-Identification-Hack-2023

ExploreAI Academy Classification Hackathon Overview South Africa is a multicultural society that is

Thato-rabodiba/South-African-Language-Identification-Hack-2023

# South-African-Language-Identification-Hack-2023 Language Identification Hackathon South African L