Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Template-Based Approach to Intelligent Multilingual Corpora Transcription

Domaine:

natural language processing

Type de record:

software
Créateur:
MosEnoAniOlu
Éditeur:
Edi
Hôte:
Emerging linguistic problems are data-driven and multidisciplinary, requiring richly transcribed corpora. Accurate corpus transcription therefore demands intelligent protocols that satisfy the following important criteria: 1) acceptability by end-users, computers/machines; 2) conformity to existing language standards, rules and structures; and 3) representation within the context of the intended language domain. To demonstrate the feasibility of these criteria, a template-based framework for multilingual transcription was proposed and implemented. The first version of the developed transcription tool, also called SCAnnAL (Speech Corpus Annotator for African Languages), applies signal processing to pre-segment waveforms of a recorded speech corpus, into word, syllable and phoneme units, resulting in a pre-segmented TextGrid file with empty labels. Using preformatted templates, the front-end or linguistic aspects/datasets (the text corpus, vowels inventory, consonants inventory, and a set of syllabification rules) are specified in a default language. A Natural Language Understanding (NLU) algorithm then uses these datasets with a data-driven syllabification algorithm to relabel subtrees of the TextGrid file. Tone pattern models were finally constructed from translations of experimental data, using the Ibadan 400 words (a list of basic items of a language), for four Nigerian tone languages. Integration of the tone pattern models into the transcription system is expected in a future paper. This research will benefit emerging digital humanists and computational linguists working on language data, as well as open new opportunities for improved African tone language speech processing systems.

Visit

doi.org

Tasks

speech processing

Licenses

https://www.euppublishing.com/customer-services/librarians/text-and-data-mining-tdm

Similaires

Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence AlignmentSmall-Multilingual-Corporamultilingual Corpora for Ethiopian LanguagesAnalysing the cadastral template using a grounded theory approachGlot500: Scaling Multilingual Corpora and Language Models to 500 LanguagesWebCrawl African : A Multilingual Parallel Corpora for African Languages

Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment

Multilingual sentence representations pose a great advantage for low-resource languages that do not

Small-Multilingual-Corpora

Small Multilingual Pretraining Copora used in the ICML 2025 Paper: Banyan: Improved Representation L

multilingual Corpora for Ethiopian Languages

Analysing the cadastral template using a grounded theory approach

The Cadastral Template contains a wealth of qualitative and quantitative data gathered from 47 na

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., makin

WebCrawl African : A Multilingual Parallel Corpora for African Languages

WebCrawl African is a mixed domain multilingual parallel corpora for a pool of African languages com