Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Advancing Arabic Reverse Dictionary Systems: A Transformer-Based Approach with Dataset Construction Guidelines

Domain:

natural language processing

Record type:

paperdatasetmodelsoftware
Creator:
SibAhmHarNac
Host:avatar
This study addresses the critical gap in Arabic natural language processing by developing an effective Arabic Reverse Dictionary (RD) system that enables users to find words based on their descriptions or meanings. We present a novel transformer-based approach with a semi-encoder neural network architecture featuring geometrically decreasing layers that achieves state-of-the-art results for Arabic RD tasks. Our methodology incorporates a comprehensive dataset construction process and establishes formal quality standards for Arabic lexicographic definitions. Experiments with various pre-trained models demonstrate that Arabic-specific models significantly outperform general multilingual embeddings, with ARBERTv2 achieving the best ranking score (0.0644). Additionally, we provide a formal abstraction of the reverse dictionary task that enhances theoretical understanding and develop a modular, extensible Python library (RDTL) with configurable training pipelines. Our analysis of dataset quality reveals important insights for improving Arabic definition construction, leading to eight specific standards for building high-quality reverse dictionary resources. This work contributes significantly to Arabic computational linguistics and provides valuable tools for language learning, academic writing, and professional communication in Arabic.

Visit

arxiv.org

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similar

MURAD: A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary DatasetPre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative StudyStreamlining Checks Processing: Advancing Arabic Handwriting Verification with a CNN-Based SystemA Low-Resource Arabic Dataset and Transformer-Based Benchmark for Dark Pattern Detection in E-Commerce Mobile ApplicationsGeoRoBERTa: A Transformer-based Approach for Semantic Address MatchingA transformer-based approach to Nigerian Pidgin text generation

MURAD: A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset

Arabic is a linguistically and culturally rich language with a vast vocabulary that spans scientific

Pre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative Study

Question answering(QA) is one of the most challenging yet widely investigated problems in Natural La

Streamlining Checks Processing: Advancing Arabic Handwriting Verification with a CNN-Based System

A Low-Resource Arabic Dataset and Transformer-Based Benchmark for Dark Pattern Detection in E-Commerce Mobile Applications

Arabic remains underrepresented in many task-specific natural language processing (NLP) resources an

GeoRoBERTa: A Transformer-based Approach for Semantic Address Matching

International audience In this paper, we describe a solution for a specific Entity Ma

A transformer-based approach to Nigerian Pidgin text generation

Abstract This paper describes the development of a transformer-based text generation model for Nige