Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
BorKudMasGor
Hôte:avatar
We introduce LRLspoof, a large-scale multilingual synthetic-speech corpus for cross-lingual spoof detection, comprising 2,732 hours of audio generated with 24 open-source TTS systems across 66 languages, including 45 low-resource languages under our operational definition. To evaluate robustness without requiring target-domain bonafide speech, we benchmark 11 publicly available countermeasures using threshold transfer: for each model we calibrate an EER operating point on pooled external benchmarks and apply the resulting threshold, reporting spoof rejection rate (SRR). Results show model-dependent cross-lingual disparity, with spoof rejection varying markedly across languages even under controlled conditions, highlighting language as an independent source of domain shift in spoof detection. The dataset is publicly available at \href{huggingface.co and \href{modelscope.cn This paper has been submitted to Interspeech 2026 for review

Visit

arxiv.org

Tasks

speech processing

Tags

SoundAudio and Speech Processing

Similaires

Large Language Models as Detectors or Instigators of Hate Speech in Low-resource Ethiopian LanguagesWhen Distributions Shifts: Causal Generalization for Low-Resource LanguagesLow-Resource Corpus Indonesian Local LanguageDo LLM hallucination detectors suffer from low-resource effect?EthioMT: Parallel Corpus for Low-resource Ethiopian LanguagesParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

Large Language Models as Detectors or Instigators of Hate Speech in Low-resource Ethiopian Languages

When Distributions Shifts: Causal Generalization for Low-Resource Languages

Machine learning models often fail under distribution shifts, a problem exacerbated in low-resource

Low-Resource Corpus Indonesian Local Language

This study departs from the hypothesis that combining Neural Machine Translation (NMT) with the stem

Do LLM hallucination detectors suffer from low-resource effect?

LLMs, while outperforming humans in a wide range of tasks, can still fail in unanticipated ways. We

EthioMT: Parallel Corpus for Low-resource Ethiopian Languages

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT leaves much to be desired for

ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

Description of the Dataset This dataset consists of three parallel corpora involving the Kabyle lan