Logo Lanfrica

Analogy-Based Classifier for Text Morphological Disambiguation: A Comparative Study

Domaine:

natural language processing

Type de record:

paper
Créateur:
ElaBouEtt
Éditeur:
UsiLabUni
Éditeur:
CCSDSpr
Hôte:avatar
International audience

Arabic is a linguistically rich and complex language, characterized by extensive morphological and orthographic variation, as well as diverse syntactic and semantic structures. These features often lead to significant morphological ambiguity. In this study, we address the challenge of morphological disambiguation in Arabic texts by formulating it as a classification task. Each morphological feature corresponds to a class, and a classification algorithm is employed to assign the correct class to each word based on its context. Our focus is on evaluating an analogy-based classifier for this task. Specifically, we explore the use of Analogical Proportions (AP), which capture the relationship among four elements-A, B, C, and D-such that "A differs from B as C differs from D" Using AP-based inference, the classifier predicts the unknown fourth element (D) based on the known values of A, B, and C. Since standard machine learning classifiers for morphological disambiguation require perfectly structured data-with training and test instances containing complete and precise information-we propose a method to transform imperfect datasets into ideal forms with accurate attributes and definitive class labels. We then assess the performance of our analogy-based classifier on a corpus of classic Arabic texts and compare its results with those produced by established machine learning and deep learning models.