Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Benchmarking Akan ASR Models Across Domain-Specific Datasets: A Comparative Evaluation of Performance, Scalability, and Adaptability

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
MenWiaEkpApp
Hôte:avatar
Most existing automatic speech recognition (ASR) research evaluate models using in-domain datasets. However, they seldom evaluate how they generalize across diverse speech contexts. This study addresses this gap by benchmarking seven Akan ASR models built on transformer architectures, such as Whisper and Wav2Vec2, using four Akan speech corpora to determine their performance. These datasets encompass various domains, including culturally relevant image descriptions, informal conversations, biblical scripture readings, and spontaneous financial dialogues. A comparison of the word error rate and character error rate highlighted domain dependency, with models performing optimally only within their training domains while showing marked accuracy degradation in mismatched scenarios. This study also identified distinct error behaviors between the Whisper and Wav2Vec2 architectures. Whereas fine-tuned Whisper Akan models led to more fluent but potentially misleading transcription errors, Wav2Vec2 produced more obvious yet less interpretable outputs when encountering unfamiliar inputs. This trade-off between readability and transparency in ASR errors should be considered when selecting architectures for low-resource language (LRL) applications. These findings highlight the need for targeted domain adaptation techniques, adaptive routing strategies, and multilingual training frameworks for Akan and other LRLs. This version has been reviewed and accepted for presentation at the Future Technologies Conference (FTC) 2025, to be held on 6 & 7 November 2025 in Munich, Germany. 17 pages, 4 figures, 1 table

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

Akan

Tags

Computation and LanguageMachine LearningSoundAudio and Speech Processing

Similaires

Adaptability, Scalability and Sustainability of Mhealth Projects Performance in Low and Medium-Income Countries: A Systematic Reviewasr-africa/African-ASR-Domain-Adaptation-EvaluationComparative Internal and External AUC Performance Across Single-Site Models.Comparative Analysis of Cross-Lingual Transfer in Multilingual Versus Monolingual Models on Domain-Specific BenchmarksMEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and TasksAdvancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models

Adaptability, Scalability and Sustainability of Mhealth Projects Performance in Low and Medium-Income Countries: A Systematic Review

International audience Mobile health (mHealth) initiatives have immense potential to

asr-africa/African-ASR-Domain-Adaptation-Evaluation

This dataset forms the the Africa ASR domain adaptation benchmark. The goal of the dataset is to ena

Comparative Internal and External AUC Performance Across Single-Site Models.

The Area Under the Receiver Operating Characteristic Curve (AUC) performance of machine learning

Comparative Analysis of Cross-Lingual Transfer in Multilingual Versus Monolingual Models on Domain-Specific Benchmarks

This paper shows that pretraining multilingual language models at scale leads to significant perform

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. Ho

Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models

Arabic diacritics, similar to short vowels in English, provide phonetic and grammatical information