Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Corpus Voice Dataset Creation in Low-Resource Contexts: A systematic review

Domain:

natural language processing

Record type:

paper
Creator:
Kum
Editor:
Cen
Publisher:
OSF
Host:avatar
Voice corpora are fundamental resources for developing speech technologies, such as automatic speech recognition (ASR), speaker identification, and natural language processing. However, the creation of such datasets remains particularly challenging in low-resource settings, where linguistic diversity, limited technological infrastructure, and financial constraints hinder their systematic development. This systematic literature review aims to synthesize existing research on voice corpus dataset creation in low-resource contexts, focusing on the reported challenges, data collection strategies, and opportunities for advancing speech technologies in under-resourced languages and regions.

Visit

doi.orgosf.io

Tasks

speech processing

Similar

A Systematic Review of AI Usage for Educational Content in Low Resource Contexts: The Case of MadagascarDynAg Open Voice Dataset for Low-Resource Bihari LanguagesCreation of a Nigerian Voice Corpus for Indigenous Speaker RecognitionExploring OCR for the Low-Resource Dogri Language: A Deep Learning Approach Enabled by Comprehensive Dataset Creation and Systematic EvaluationA Systematic Review of AI Usage for Educational Content in Medium and Low Resource Contexts: Global Overview and Case of MadagascarLLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

A Systematic Review of AI Usage for Educational Content in Low Resource Contexts: The Case of Madagascar

Background: Artificial intelligence (AI), particularly natural language processing (NLP

DynAg Open Voice Dataset for Low-Resource Bihari Languages

jjkhkjhkjh

Creation of a Nigerian Voice Corpus for Indigenous Speaker Recognition

Abstract One of the goals of Word Bank’s Identification for Development (ID4D) is t

Exploring OCR for the Low-Resource Dogri Language: A Deep Learning Approach Enabled by Comprehensive Dataset Creation and Systematic Evaluation

A Systematic Review of AI Usage for Educational Content in Medium and Low Resource Contexts: Global Overview and Case of Madagascar

Background: Artificial intelligence (AI), particularly natural language processing (NLP) and deep le

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safet