Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

DiaLEX – Egyptian (DiaLEX-EA)

Domain:

natural language processing

Record type:

dataset
Publisher:
ELR
Host:avatar
The Egyptian Arabic Full-Form Lexicon (DiaLEX-EA) is a comprehensive computational lexicon covering the Egyptian Arabic dialect. Featuring over 93,000,000 forms for 33,000 lemmas, this full-form lexicon provides exhaustive treatment of all inflected forms.DiaLEX-EA has several features that make it ideally suited to support natural language processing applications for Egyptian Arabic, especially morphological analysis and speech technology, including:1.Extremely comprehensive coverage – over 93 million entries2.Comprehensive treatment of all inflected forms, enclitics, proclitics, case endings, declensions, and conjugated forms.3.Full and accurate diacriticization (vocalization), essential for speech technology.4.Extensive coverage of variants which is necessary since dialects don't have a standard orthography.Please note: Phonetic transcriptions, IPA and/or SAMPA, fine-tuned to the licensor’s specifications, are available upon request.Quantity and size: 93,349,857 lines / 13,321 MB (13.0 GB)File format: flat TSV text filesSamples and a specifications document are available upon request.

Visit

catalog.elra.info

Licenses

Rights available for: nonCommercialUse, commercialUse

Similar

DiaLex: A Benchmark for Evaluating Multidialectal Arabic Word EmbeddingsBenchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

DiaLex: A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embedding. DiaLex covers five importan

Benchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embeddings. DiaLex covers five importa