Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Integrating Human Linguistic Insights into AI: Theory-Driven Representation for Multilingual Text-to-Speech

Domain:

natural language processing

Record type:

paper
Creator:
ZhaZenLiuZhe
Publisher:
arXiv
Host:avatar
This paper explores the integration of human linguistic insights into multilingual text-to-speech (TTS) systems by evaluating the Featurally Underspecified Lexicon (FUL) as a theory-driven input representation. Unlike data-intensive end-to-end models, FUL offers a compact, interpretable feature set grounded in phonological principles, enabling scalable and equitable TTS development for low-resource languages. We provide a mapping from language-specific phones to FUL feature vectors via a SAMPA intermediate and incorporate these features into a modified FastSpeech architecture. Experiments were conducted to evaluate their ability to generate native, non-native, and code-mixed speech in English and Mandarin. We ran an experiment with a small dataset and one with a larger dataset, which showed that TTS with FUL features as input could produce intelligible native speech with as little as 8 hours of training data; with 100 hours of training data, intelligible speech could be generated for a language not present in the training data. The approach further supports code-mixed synthesis while preserving consistent timbre and interpretable phonetic control. These results highlight the potential of theory-driven representations for building efficient, scalable, and linguistically informed TTS systems, demonstrating that phonological features can function as both analytical tools and practical inputs for speech technology. Accepted by Phonetica; earlier version: arXiv:2110.03609

Visit

doi.org

Tasks

speech processingtext to speech

Tags

Computation and Language (cs.CL)Sound (cs.SD)Audio and Speech Processing (eess.AS)FOS: Computer and information sciencesFOS: Electrical engineering, electronic engineering, information engineering

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Multilingual Text Representationsanusishafii989/Multilingual-Speech-to-TextYarnGPT - Nigerian Text-to-Speech AIIntegrating AI-Driven Triage into Digital Pharmacy Systems for Rational Antibiotic Use in Low-Resource Settings.Generating Arabic text in multilingual speech-to-speech machine translation frameworkXLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition

Multilingual Text Representation

Modern NLP breakthrough includes large multilingual models capable of performing tasks across more t

sanusishafii989/Multilingual-Speech-to-Text

A powerful, multilingual speech-to-text application with advanced NLP analytics capabilities. Suppor

YarnGPT - Nigerian Text-to-Speech AI

Generate high-quality, authentic Nigerian text-to-speech audio for your podcasts, videos, and applications.

Integrating AI-Driven Triage into Digital Pharmacy Systems for Rational Antibiotic Use in Low-Resource Settings.

Antimicrobial resistance (AMR) is becoming a bigger threat to global health, especially in low and m

Generating Arabic text in multilingual speech-to-speech machine translation framework

XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition

In this paper, we propose a weakly supervised multilingual representation learning framework, called