Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multilingual Word Segmentation: Training Many Language-Specific Tokenizers Smoothly Thanks to the Universal Dependencies Corpus

Domaine:

natural language processing

Type de record:

papersoftwaremodel
Créateur:
MorVog
Éditeur:
Sch
Éditeur:
CCSD
Hôte:avatar
International audience This paper describes how a tokenizer can be trained from any dataset in the Universal Dependencies 2.1 corpus (UD2) (Nivre et al., 2017). A software tool, which relies on Elephant (Evang et al., 2013) to perform the training, is also made available. Beyond providing the community with a large choice of language-specific tokenizers, we argue in this paper that: (1) tokenization should be considered as a supervised task; (2) language scalability requires a streamlined software engineering process across languages.

Visit

hal.science

Tags

InteroperabilityMultilingualityTokenizationWord SegmentationUniversal Dependencies[INFO.INFO-TT]Computer Science [cs]/Document and Text Processing[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]

Licenses

info:eu-repo/semantics/OpenAccess

Similaires

Universal DependenciesMultilingual Entity and Relation Extraction from Unified to Language-specific TrainingUniversal Dependencies TreebankUniversal Dependencies for AmharicParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource LanguageUniversal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks

Universal Dependencies

Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology pe

Multilingual Entity and Relation Extraction from Unified to Language-specific Training

Entity and relation extraction is a key task in information extraction, where the output can be used

Universal Dependencies Treebank

Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank a

Universal Dependencies for Amharic

ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

Description of the Dataset This dataset consists of three parallel corpora involving the Kabyle lan

Universal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks

In this paper we present a novel lemmatization method based on a sequence-to-sequence neural network