Logo Lanfrica

Eliezermga/Lugayetu

Domain:

natural language processing

Record type:

datasetproject
Creator:
eli
Host:
Lugayetu is a research and technology project focused on preserving Congolese low-resource languages and developing artificial intelligence tools for them. The platform collects textual and voice data from native speakers to build ai tools (translator...) # Lugayetu The Ruund-French dataset is available on Hugging Face: k --- ## Project Overview **Lugayetu** is a research and development project dedicated to the **preservation of low-resource Congolese languages**, particularly those spoken in the **Democratic Republic of the Congo**. The project aims to create **digital linguistic resources** and develop **artificial intelligence technologies** capable of processing these languages, especially in the fields of **machine translation** and **Natural Language Processing (NLP)**. Lugayetu combines two main objectives: 1. **Develop a text-to-text automatic translator** 2. **Build a multimodal linguistic corpus (text and speech)** for future AI research --- # Objectives The project pursues several scientific and technological goals: - preserve and digitize local languages - create **structured linguistic datasets** - develop **artificial intelligence models for under-resourced languages** - facilitate research in **Natural Language Processing** --- # Machine Translation (Text-to-Text) The first phase of the project focuses on developing a **text-to-text machine translation system**. The system aims to automatically translate sentences between: - **Ruwund (Rund)** - **French** This part of the project involves: - building a **parallel corpus** - cleaning and aligning texts - training **machine translation models** The goal is to create a **neural translation prototype** capable of understanding and translating Ruwund into French. --- # Linguistic Data Collection One of the main challenges of Congolese languages is the **lack of digital data**. To address this issue, Lugayetu provides a platform to: - collect **sentences in different languages** - associate **corresponding translations** - build a **parallel corpus for model training** These data are essential for developing effective artificial intelligence systems. --- # Voice Data Collection …