Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Optimizing Cognacy for Automatic Internal Classifcation of Languages: the Critical Role of Segmentation

Domain:

natural language processing

Record type:

papersoftware
Creator:
Kpo
Editor:
Lan
Publisher:
CCSD
Host:avatar
International audience This paper explores methods for automatic cognate detection. Using a publicly available wordlist, it investigates the relationship between data quality and the performance of cognate detection methods. The study employs the LexStat algorithm, available in Lingpy, to assess the performance of partial and full cognate detection on data segmented according to four parsing criteria: no parsing, phonetic parsing, morphological parsing, and morpho-phonotactic parsing. The results indicate that unparsed data is more suitable for full cognate detection, while parsed data performs better in partial cognate detection. In both cases, higher levels of parsing lead to improved cognate detection. However, increased cognate detection does not necessarily result in improved phylogenetic relationships. The findings suggest that, while efforts to enhance cognate detection are important, it is crucial to recognize thresholds where further parsing could negatively affect overall performance, preventing unintended outcomes.

Visit

hal.science

Tags

word segmentationDogon languagescluster analysiscognate detection[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL][SCCO.LING]Cognitive science/Linguistics