Logo Lanfrica

YABKALY/Merq-data

Domaine:

natural language processing

Type de record:

software
Créateur:
YAB
Hôte:
Open source AI dataset collection platform for collecting, reviewing, and exporting high quality Ethiopian datasets. # Merq Data **Merq Data** is an open-source AI dataset collection platform designed to help communities, researchers, developers, universities, and organizations collect, review, manage, and export high-quality Ethiopian datasets for machine learning and artificial intelligence research. The platform focuses on solving one of the biggest challenges in AI development: the lack of clean, local, diverse, and well-structured datasets. Many AI systems are trained using data from other countries, cultures, and languages, which makes them less accurate and less useful for Ethiopian contexts. Merq Data aims to support the creation of datasets that represent Ethiopian languages, communities, environments, and real-world use cases. The first version of Merq Data focuses on collecting Amharic text-based speech data. Admins can upload Amharic text prompts, and volunteers can register through a mobile application, view assigned text, record themselves reading the text, and submit the audio recording. Each submission is stored with useful metadata such as age range, sex, native language, region, and submission status. Admins or reviewers can then listen to the recordings, approve or reject them, and export approved data as an AI-ready dataset. ## Project Goal The main goal of Merq Data is to make Ethiopian AI dataset collection easier, more organized, ethical, and open-source. Merq Data is built to support: * Local language AI research * Speech recognition dataset collection * Text-to-speech dataset preparation * Natural language processing projects * Academic research * Community-based data collection * Open-source AI development * Dataset review and quality control * Structured dataset export for machine learning ## Why Merq Data? AI systems need data to learn. However, high-quality Ethiopian datasets are still limited, especially for local languages such as Amharic, Afan Oromo, Tigrinya, Somali, Sidama, Wolaytta, and others. Without strong local datasets, it becomes …