Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Grouping conversational markers across languages by exploiting large comparable corpora and unsupervised segmentation

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
PreStaTse
Éditeur:
LabAcaANR
Éditeur:
CCSD
Hôte:avatar
International audience This work approaches Conversational and Discourse Markers (hereafter DM) from a radical data-driven perspective grounded in large comparable corpora of French, English and Taiwan Mandarin conversations. The key features of our approach are (i) to account for lexicalization as a by-product of unsupervised segmentation applied to our corpora, (ii) to exploit simple metrics for clustering DM (both within a language and within multilingual clusters). We explore the benefits and the drawbacks of such a radical approach to DM. In particular we compare the DM clusters obtained from traditional segmentation into tokens (as given by manual transcription of the corpora) vs. unsupervised segmentation. The metrics on which we ground the clustering experiments are based on contrast between (i) short vs. longer utterances distribution and (ii) position within longer utterances.

Visit

hal.science

Tags

[SCCO.LING]Cognitive science/Linguistics[SCCO.COMP]Cognitive science/Computer science

Licenses

info:eu-repo/semantics/OpenAccess