Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Pronunciation Modeling for Synthesis of Low Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
Sit
Publisher:
Car
Host:avatar
Natural and intelligible Text to Speech (TTS) systems exist for a number of languages in the world today. However, there are many languages of the world, for which building TTS systems is still prohibitive, due to the lack of linguistic resources and data. Some of these languages are spoken by a large population of the world. Others are primarily spoken languages, or languages with large non-literate populations, which could benefit from speech-based systems. One of the bottlenecks in creating TTS systems in new languages is designing a frontend, which includes creating a phone set, lexicon and letter to sound rules, which contribute to the pronunciation of the system. In this thesis, we use acoustics and cross-lingual models and techniques using higher resource languages to improve the pronunciation of TTS systems in low resource languages. First, we present a grapheme-based framework that can be used to build TTS systems for most languages of the world that have a written form. Such systems either treat graphemes as phonemes or assign a single pronunciation to each grapheme, which may not be completely accurate for languages with ambiguities in their written forms. We improve the pronunciation of grapheme-based voices implicitly by using better modeling techniques. We automatically discover letter-to-sound rules such as schwa deletion using related higher resource languages. We also disambiguate homographs in lexicons in dialects of Arabic to improve the pronunciation of TTS systems. We show that phoneme-like features derived using Articulatory Features may be useful for improving grapheme-based voices. We present a preliminary framework addressing the problem of synthesizing Code Mixed text found often in Social Media. Lastly, we use acoustics and cross-lingual techniques to automatically derive written forms for building TTS systems for languages without a standardized orthography.

Visit

doi.orgkilthub.cmu.edu

Tasks

text to speechspeech processing

Tags

Natural language processingInformation and computing sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Text-to-Speech Synthesis Using Found Data for Low-Resource Languagesnaolbakala/Adapting-Outperformer-Topic-Modeling-for-low-resource-Major-Ethiopian-LanguagesExploiting Cross-Lingual Knowledge in Unsupervised Acoustic Modeling for Low-Resource LanguagesMultilingual Byte2Speech Models for Scalable Low-resource Speech SynthesisModeling Multimodal Discourse Coherence in Low-Resource Languages: Work-in-ProgressMining Large-Scale Low-Resource Pronunciation Data From Wikipedia

Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

Text-to-speech synthesis is a key component of interactive, speech-based systems. Typically, buildi

naolbakala/Adapting-Outperformer-Topic-Modeling-for-low-resource-Major-Ethiopian-Languages

Topic words based dataset extracted from large size of social media data from three major Ethiopian

Exploiting Cross-Lingual Knowledge in Unsupervised Acoustic Modeling for Low-Resource Languages

(Short version of Abstract) This thesis describes an investigation on unsupervised acoustic modeling

Multilingual Byte2Speech Models for Scalable Low-resource Speech Synthesis

To scale neural speech synthesis to various real-world languages, we present a multilingual end-to-e

Modeling Multimodal Discourse Coherence in Low-Resource Languages: Work-in-Progress

Discourse coherence is vital for language understanding but underexplored in low-resource languages.

Mining Large-Scale Low-Resource Pronunciation Data From Wikipedia

Pronunciation modeling is a key task for building speech technology in new languages, and while soli