Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
EdeOyoArcBas
Éditeur:
arXiv
Hôte:avatar
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages. 6 pages, 2 figures. Accepted to Interspeech 2026

Visit

doi.org

Tasks

speech processingtext to speech

Languages

Efik

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Efik-TTS-DatasetBuilding a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language ApproachTowards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource LanguageAdapting Foundational ASR Models to Efik: An Empirical Study of an Extremely Low-Resource Tonal LanguageInkubaLM: A small language model for low-resource African languagesTowards a National Framework for Digital Preservation in Nigeria: Technologies and Best Practices

Efik-TTS-Dataset

This dataset comprises audio recordings of Efik speech aligned with textual transcriptions. The data

Building a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language Approach

Banyumasan Javanese, widely recognized through the Ngapak dialect, remains culturally significant bu

Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language

The increase in technological adoption worldwide comes with demands for novel tools to be used by th

Adapting Foundational ASR Models to Efik: An Empirical Study of an Extremely Low-Resource Tonal Language

InkubaLM: A small language model for low-resource African languages

High-resource language models often fall short in the African context, where there is a critical nee

Towards a National Framework for Digital Preservation in Nigeria: Technologies and Best Practices

The need for preserving digital resources (acquired or generated) by institutions in Nigeria becomes