Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding

Domaine:

natural language processing

Type de record:

papersoftwaredataset
Créateur:
Mar
Hôte:avatar
Syllable-level units offer compact and linguistically meaningful representations for spoken language modeling and unsupervised word discovery, but research on syllabification remains fragmented across disparate implementations, datasets, and evaluation protocols. We introduce findsylls, a modular, language-agnostic toolkit that unifies classical syllable detectors and end-to-end syllabifiers under a common interface for syllable segmentation, embedding extraction, and multi-granular evaluation. The toolkit implements and standardizes widely used methods (e.g., Sylber, VG-HuBERT) and allows their components to be recombined, enabling controlled comparisons of representations, algorithms, and token rates. We demonstrate findsylls on English and Spanish corpora and on new hand-annotated data from Kono, an underdocumented Central Mande language, illustrating how a single framework can support reproducible syllable-level experiments across both high-resource and under-resourced settings. 4 pages + 2 for references, disclosures & acknowledgements; currently under review

Visit

arxiv.org

Tasks

speech processing

Languages

MandinkaManinkakan, EasternNomaandeTobanga

Tags

Computation and LanguageArtificial Intelligence

Similaires

Unsupervised Speech Recognition at the Syllable LevelIntroducing Syllable Tokenization for Low-resource Languages: A Case Study with SwahiliLinguini: A benchmark for language-agnostic linguistic reasoningSyllable-Based Speech Recognition for AmharicYembaTones: A syllable-tone annotated dataset for speech recognition and prosodic analysis of the Yemba languageTokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies

Unsupervised Speech Recognition at the Syllable Level

Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognitio

Introducing Syllable Tokenization for Low-resource Languages: A Case Study with Swahili

Many attempts have been made in multilingual NLP to ensure that pre-trained language models, such as

Linguini: A benchmark for language-agnostic linguistic reasoning

We propose a new benchmark to measure a language model's linguistic reasoning skills without relying

Syllable-Based Speech Recognition for Amharic

Amharic is the Semitic language that has the second large number of speakers after Arabic (Hayward and Richard 1999). Its writing system is syllabic with Consonant-Vowel (CV) syllable structure. Amharic orthography has more or less a one to one correspondence with

YembaTones: A syllable-tone annotated dataset for speech recognition and prosodic analysis of the Yemba language

Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies

Gender-inclusive NLP research has documented the harmful limitations of gender binary-centric large