Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Organizing Contextual Knowledge for Arabic Text Disambiguation and Terminology Extraction

Domain:

natural language processing

Record type:

paper
Creator:
BouElaEvrSli
Editor:
TunLabInsLog
Publisher:
CCSDErg
Host:avatar
International audience Ontologies have an important role in knowledge organization and information retrieval. Domain ontologies are composed of concepts represented by domain relevant terms. Existing approaches of ontology construction make use of statistical and linguistic information to extract domain relevant terms. The quality and the quantity of this information influence the accuracy of terminologyextraction approaches and other steps in knowledge extraction and information retrieval. This paper proposes an approach forhandling domain relevant terms from Arabic non-diacriticised semi-structured corpora. In input, the structure of documentsis exploited to organize knowledge in a contextual graph, which is exploitedto extract relevant terms. This network contains simple and compound nouns handled by a morphosyntactic shallow parser. The noun phrases are evaluated in terms of termhood and unithood by means of possibilistic measures. We apply a qualitative approach, which weighs terms according to their positions in the structure of the document. In output, the extracted knowledge is organized as network modeling dependencies between terms, which can be exploited to infer semantic relations.We test our approach on three specific domain corpora. The goal of this evaluation is to check if our model for organizing and exploiting contextual knowledge will improve the accuracy of extraction of simple and compound nouns. We also investigate the role of compound nouns in improving information retrieval results.

Visit

hal.science

Tasks

information extraction

Tags

Contextual knowledgeText disambiguationTerminology extractionInformation retrieval[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

Adaptive Auto-encoder for Extraction of Arabic Text: invariant, font, and segmentContextual Text Embeddings for TwiUsing Semi-automated Term Extraction for IsiNdebele Health TerminologyLULCC-KnowText - annotated text segments for knowledge extraction on Land Use and Land Cover changeAnalogy-Based Classifier for Text Morphological Disambiguation: A Comparative StudyFraming Food Online Discourse: Employing Generative AI and Semantic Analysis for Digital Lexicography and Terminology Extraction Journal of Digital Terminology and Lexicography

Adaptive Auto-encoder for Extraction of Arabic Text: invariant, font, and segment

Abstract Adaptive auto-en-codor research strategy for categorizing Arabic text into three

Contextual Text Embeddings for Twi

Transformer-based language models have been changing the modern Natural Language Processing (NLP) la

Using Semi-automated Term Extraction for IsiNdebele Health Terminology

IsiNdebele, also known as Southern isiNdebele, has a limited availability of language resources and

LULCC-KnowText - annotated text segments for knowledge extraction on Land Use and Land Cover change

This dataset contains a corpus of annotated text segments (sentences) extracted from scient

Analogy-Based Classifier for Text Morphological Disambiguation: A Comparative Study

International audience

Arabic is a linguistically rich and complex language,

Framing Food Online Discourse: Employing Generative AI and Semantic Analysis for Digital Lexicography and Terminology Extraction Journal of Digital Terminology and Lexicography

Social media platforms provide vast amounts of authentic, user-generated linguistic data that can