Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilingual Word Segmentation: Training Many Language-Specific Tokenizers Smoothly Thanks to the Universal Dependencies Corpus

Domain:

natural language processing

Record type:

papersoftwaremodel
Creator:
MorVog
Editor:
Sch
Publisher:
CCSD
Host:avatar
International audience This paper describes how a tokenizer can be trained from any dataset in the Universal Dependencies 2.1 corpus (UD2) (Nivre et al., 2017). A software tool, which relies on Elephant (Evang et al., 2013) to perform the training, is also made available. Beyond providing the community with a large choice of language-specific tokenizers, we argue in this paper that: (1) tokenization should be considered as a supervised task; (2) language scalability requires a streamlined software engineering process across languages.

Visit

hal.science

Tags

InteroperabilityMultilingualityTokenizationWord SegmentationUniversal Dependencies[INFO.INFO-TT]Computer Science [cs]/Document and Text Processing[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]

Licenses

info:eu-repo/semantics/OpenAccess