Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Technological Development Models in the Context of Speech Corpora Imbalance

Domaine:

natural language processing

Type de record:

paper
Créateur:
BaiGavKhaNik
Éditeur:
Pet
Hôte:avatar
The development of speech and language technologies in the era of artificial intelligence critically depends on the availability of large-scale, high-quality linguistic data. While low-resource languages have been widely studied, less attention has been paid to data imbalances among languages that are considered digitally well-supported. This paper examines the uneven distribution of open speech corpora across languages with established infrastructure of speech technologies and available datasets, showing that this disparity creates structural bottlenecks for sovereign AI development. We conduct a comparative analysis of open and non-commercial speech datasets, accounting for demographic factors, licensing conditions, and models of technological development. To quantify resource inequality, we propose the Digital Resource Saturation Index (DRSI), which relates the availability of speech data to the potential for content generation and consumption within language communities. Our findings reveal a strong dominance of English for open speech resources, while many non-Western languages – including Russian – remain systematically underrepresented. While interpreting these results through the lens of Western and non-Western technological modernization models, we suggest that language inequality in AI is not merely a technical or demographic issue, but a self-reinforcing structurally reproduced outcome of data governance, institutional coordination, and political choices regarding openness and digital sovereignty. The study further provides practical recommendations for mitigating these imbalances and fostering a more equitable technological landscape. Technology and Language, 7(1), 80-102

Visit

doi.orgsoctech.spbstu.ru

Tasks

speech processing

Tags

Digital language divideSpeech corpora imbalanceLanguage inequalityTechnological development modelsResource disparity analysisDigital resource saturation indexDRSI

Licenses

Creative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcode

Similaires

Sepedi Speech CorporaDetecting Online Hate Speech Using Context Aware ModelsOpenSLR African Speech CorporaCreation of an Afrikaans Speech Corpora for Speech Emotion RecognitionFirst automatic fongbe continuous speech recognition system: Development of acoustic models and language modelsReflections on the development of corpora for Zimbabwe’s understudied languages

Sepedi Speech Corpora

A corpus of Sesotho sa Leboa telephone speech data collected from mother tongue speakers of the sta

Detecting Online Hate Speech Using Context Aware Models

In the wake of a polarizing election, the cyber world is laden with hate speech. Context accompanyin

OpenSLR African Speech Corpora

ASR system training, speech synthesis Notes / challenges: Includes crowd-sourced recordings for Yor

Creation of an Afrikaans Speech Corpora for Speech Emotion Recognition

First automatic fongbe continuous speech recognition system: Development of acoustic models and language models

This paper reports our efforts toward an ASR system for a new under-resourced language (Fongbe). The aim of this work is to build acoustic models and language models for continuous speech decoding in Fongbe. The problem encountered with Fongbe (an African language

Reflections on the development of corpora for Zimbabwe’s understudied languages

Linguistic corpora are one of the primary research tools in modern-day linguistics. The centrality o