Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Leveraging Twitter for Low-Resource Conversational Speech Language Modeling

Domaine:

natural language processing

Type de record:

paper
Créateur:
JaeOst
Hôte:avatar
In applications involving conversational speech, data sparsity is a limiting factor in building a better language model. We propose a simple, language-independent method to quickly harvest large amounts of data from Twitter to supplement a smaller training set that is more closely matched to the domain. The techniques lead to a significant reduction in perplexity on four low-resource languages even though the presence on Twitter of these languages is relatively small. We also find that the Twitter text is more useful for learning word classes than the in-domain text and that use of these word classes leads to further reductions in perplexity. Additionally, we introduce a method of using social and textual information to prioritize the download queue during the Twitter crawling. This maximizes the amount of useful data that can be collected, impacting both perplexity and vocabulary coverage.

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and Language

Similaires

Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentationDziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian DialectEnd-to-End Speech Recognition with Deep Fusion: Leveraging External Language Models for Low-Resource ScenariosForgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English SpeechRideKE: Leveraging Low-Resource, User-Generated Twitter Content for Sentiment and Emotion Detection in Kenyan Code-Switched DatasetLeveraging Language Models for Document Type Classification in Low-Resource Afrikaans Archives

Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation

The availability of data in expressive styles across languages is limited, and recording sessions ar

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

Automatic speech and language technologies are still heavily biased toward high-resource languages,

End-to-End Speech Recognition with Deep Fusion: Leveraging External Language Models for Low-Resource Scenarios

With the rapid development of Automatic Speech Recognition (ASR) technology, end-to-end speech recog

Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech

Dementia detection from spontaneous speech offers a scalable approach to cognitive screening, yet NL

RideKE: Leveraging Low-Resource, User-Generated Twitter Content for Sentiment and Emotion Detection in Kenyan Code-Switched Dataset

Social media has become a crucial open-access platform for individuals to express opinions and share

Leveraging Language Models for Document Type Classification in Low-Resource Afrikaans Archives