Logo Lanfrica

nyacly/open-runyoro-ai

Domain:

natural language processing

Record type:

project
Creator:
nya
Host:
RunyoroAI: An open-source project to create an AI conversational partner for Runyoro using freely available texts like the Bible, fine-tuned with InkubaLM and Hugging Face Transformers, and built with Rasa. Contributions welcome to preserve this low-resource language! # Open Runyoro AI 🇺🇬 **Our Vision:** To build open-source AI tools that can read, understand, and speak Runyoro, helping to preserve and promote the language in the digital age. This project aims to create datasets and models for Natural Language Processing (NLP) and Speech tasks for the Runyoro language. **Current Goals:** 1. **Data Collection:** * **Text Corpus:** Collect a diverse and large corpus of written Runyoro. * **Speech Corpus:** Collect transcribed Runyoro audio from native speakers. 2. **Model Development (Future):** * Text-to-Speech (TTS) for Runyoro *(planned)*. * Automatic Speech Recognition (ASR) for Runyoro. * Machine Translation (e.g., Runyoro English). * Other NLP tools (e.g., part-of-speech taggers, named entity recognizers). * A minimal audio preprocessing example is available in `scripts/preprocess_audio.py`. ## 🚀 How to Contribute We welcome contributions from everyone, especially native Runyoro speakers, linguists, and AI/ML developers! **1. Contributing Data (Most Needed!):** This is the most crucial part of the project right now. High-quality data is the foundation of good AI models. * **Text Data:** * We need plain text files (.txt) containing Runyoro. * Sources can include: books, articles, websites, blogs, proverbs, folk tales, personal writings, etc. * Please ensure the text is in Runyoro and as clean as possible. * **How to submit:** Place your `.txt` files in the `data/text/` directory via a Pull Request. * **Audio Data:** * We need audio recordings (.wav, .mp3, .flac) of spoken Runyoro **along with their accurate transcriptions.** * Ideal audio is clear, with minimal background noise, spoken by a single speaker per file. * **How to submit:** 1. Place your audio files in `data/audio/wavs/` (this directory should now exist). 2. Create/update a `data/audio/metadata.csv` file with the filename and its transcription. Format: `filename|transcription`. Example: `wavs/runyoro_sentence1.wav|Ekic …