Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Semi-automatic and Low Cost Approach to Build Scalable Lemma-based Lexical Resources for Arabic Verbs

Domaine:

natural language processing

Type de record:

paperdatasetsoftware
Créateur:
DouLehMauAbd
Éditeur:
UniجامBasQat
Éditeur:
CCSDAIR
Hôte:avatar
International audience This work presents a method that enables Arabic NLP community to build scalable lexical resources. The proposed method is low cost and efficient in time in addition to its scalability and extendibility. The latter is reflected in the ability for the method to be incremental in both aspects, processing resources and generating lexicons. Using a corpus; firstly, tokens are drawn from the corpus and lemmatized. Secondly, finite state transducers (FSTs) are generated semi-automatically. Finally, FSTsare used to produce all possible inflected verb forms with their full morphological features. Among the algorithm’s strength is its ability to generate transducers having 184 transitions, which is very cumbersome, if manually designed. The second strength is a new inflection scheme of Arabic verbs; this increases the efficiency of FST generation algorithm. The experimentation uses a representative corpus of Modern Standard Arabic. The number of semi-automatically generated transducers is 171. The resulting open lexical resources coverage is high. Our resources cover more than 70% Arabic verbs. The built resources contain 16,855 verb lemmas and 11,080,355 fully, partially and not vocalized verbal inflected forms. All these resources are being made public and currently used as an open package in the Unitex framework available under the LGPL license.

Visit

hal.science

Tags

Unitex Finite state transducers Arabic verbs Arabic linguistic resourcesArabic NLP[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Similaires

Using finite-state transducers to build lexical resources for Unitex Arabic packageAn Arabic Transformation Based Approach to Automatic Paraphrasing of Syntactic SentencesSemi-automatic news video annotation framework for Arabic textCollaborative Construction of Arabic Lexical ResourcesA Sentiment analysis approach for Arabic dialects texts analysis based on automatic translation: Application to the Algerian dialect.Developing Language Resources: A Lexical-Diversity-Centric Approach

Using finite-state transducers to build lexical resources for Unitex Arabic package

International audience This paper addresses the issue of generating Arabic verbal inf

An Arabic Transformation Based Approach to Automatic Paraphrasing of Syntactic Sentences

International audience The aim of this paper is to exploit the existing Lexicon-Gramm

Semi-automatic news video annotation framework for Arabic text

Collaborative Construction of Arabic Lexical Resources

International audience no abstract

A Sentiment analysis approach for Arabic dialects texts analysis based on automatic translation: Application to the Algerian dialect.

Developing Language Resources: A Lexical-Diversity-Centric Approach

Languages describe the world in diverse ways, a phenomenon known as linguistic diversity, which is a