Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Normalization of Non-Standard Words in Croatian Texts

Domaine:

natural language processing

Type de record:

papersoftware
Créateur:
BelPobMar
Hôte:avatar
This paper presents text normalization which is an integral part of any text-to-speech synthesis system. Text normalization is a set of methods with a task to write non-standard words, like numbers, dates, times, abbreviations, acronyms and the most common symbols, in their full expanded form are presented. The whole taxonomy for classification of non-standard words in Croatian language together with rule-based normalization methods combined with a lookup dictionary are proposed. Achieved token rate for normalization of Croatian texts is 95%, where 80% of expanded words are in correct morphological form. 8 pages, 3 figures in Text, Speech and Dialogue extension to Lecture Notes in Artificial Intelligence LNAI6836. Hebernal, Ivan; Matoušek, Václav (ed). - Plzen: University of West Bohemia, 2011. 1-8 (ISBN: 987-80-261-0069-0)

Visit

arxiv.org

Tasks

text normalization

Tags

Computation and Language

Similaires

Learning to Explain Non-Standard English Words and PhrasesWolof Non-Standard to Standard Parallel PairsComparative Basic-Words of Standard Arabic Palestinian and TunisianAfrikaans in a quantitative microtypology of Germanic standard and non-standard varietiesNine Tashlhiyt texts: Structured representations of 18,000 words10. Non-standard forms of Swahili in west-central Kenya

Learning to Explain Non-Standard English Words and Phrases

We describe a data-driven approach for automatically explaining new, non-standard English expression

Wolof Non-Standard to Standard Parallel Pairs

This dataset contains pairs of non-standard and standard Wolof text, designed for training models to

Comparative Basic-Words of Standard Arabic Palestinian and Tunisian

This paper studies comparative linguistics on the process of word-formation that occurs in Modern St

Afrikaans in a quantitative microtypology of Germanic standard and non-standard varieties

Besides English, Afrikaans is considered “the [Germanic] language which deviates grammatically the f

Nine Tashlhiyt texts: Structured representations of 18,000 words

Digital structured representations of nine texts in the Tashlhiyt language of Morocco

10. Non-standard forms of Swahili in west-central Kenya