Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

Domain:

natural language processing

Record type:

datasetpaper
Creator:
EdeOyoArcBas
Publisher:
arXiv
Host:avatar
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages. 6 pages, 2 figures. Accepted to Interspeech 2026

Visit

doi.org

Tasks

speech processingtext to speech

Languages

Efik

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Efik-TTS-DatasetBuilding a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language ApproachTowards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource LanguageAdapting Foundational ASR Models to Efik: An Empirical Study of an Extremely Low-Resource Tonal LanguageInkubaLM: A small language model for low-resource African languagesTowards a National Framework for Digital Preservation in Nigeria: Technologies and Best Practices

Efik-TTS-Dataset

This dataset comprises audio recordings of Efik speech aligned with textual transcriptions. The data

Building a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language Approach

Banyumasan Javanese, widely recognized through the Ngapak dialect, remains culturally significant bu

Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language

The increase in technological adoption worldwide comes with demands for novel tools to be used by th

Adapting Foundational ASR Models to Efik: An Empirical Study of an Extremely Low-Resource Tonal Language

InkubaLM: A small language model for low-resource African languages

High-resource language models often fall short in the African context, where there is a critical nee

Towards a National Framework for Digital Preservation in Nigeria: Technologies and Best Practices

The need for preserving digital resources (acquired or generated) by institutions in Nigeria becomes