Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Alkhalil Corpus: An Open-Source Thematic and Lemmatized Corpus for Modern Standard Arabic

Domain:

natural language processing

Record type:

dataset
Creator:
AssBELMAZ
Publisher:
Und
Host:avatar
The availability of large annotated corpora remains a major challenge for the development of natural language processing systems for under-resourced languages such as Arabic. In this paper, we present two annotated corpora dedicated to Modern Standard Arabic. These corpora are open-source and freely available on the Hugging Face platform. The first corpus, annotated by theme and designed to provide a balanced representation of contemporary Arabic usage, comprises approximately one hundred million words collected from diverse sources covering multiple domains and geographical regions. The second corpus, containing approximately one million words, is a sub-corpus extracted from the first. It was annotated with lemma tags using a semi-automatic approach that combines automatic annotation with the Alkhalil lemmatizer and MADAMIRA, followed by manual validation.

Visit

doi.orgunderline.io

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similar

4FACTORS Arabic Native-Written Sample Corpus: Modern Standard Arabic, Palestinian Levantine and Egyptian (150 items)Developing an Open-Source Corpus of Yoruba SpeechMAC: An Open and Free Moroccan Arabic Corpus for Sentiment AnalysisAn Open Source System for Crowd Sourcing an African Language Short Story CorpusOpen Corpus for Arabic Maqam Recognition (OCMR)Building and Curating a High-Quality Modern Standard Arabic -Tunisian Arabic Parallel Corpus via Two-Stage LLM Evaluation

4FACTORS Arabic Native-Written Sample Corpus: Modern Standard Arabic, Palestinian Levantine and Egyptian (150 items)

A 150-item demonstration corpus of native-written Arabic across three varieties — Modern Standard Ar

Developing an Open-Source Corpus of Yoruba Speech

This paper introduces an open-source speech dataset for Yoruba - one of the largest low-resource West African languages spoken by at least 22 million people. Yoruba is one of the official languages of Nigeria, Benin and Togo, and is spoken in other neighboring Afri

MAC: An Open and Free Moroccan Arabic Corpus for Sentiment Analysis

An Open Source System for Crowd Sourcing an African Language Short Story Corpus

Many African languages are under resourced in having open access corpora for use in developing technological applications such as grammar checkers, spell checkers, speech to text, text to speech and machine translation tools. This may lead to a decline in all cultu

Open Corpus for Arabic Maqam Recognition (OCMR)

The Open Corpus for Arabic Maqam Recognition (OCMR) is a collection of melodic pitch features and ex

Building and Curating a High-Quality Modern Standard Arabic -Tunisian Arabic Parallel Corpus via Two-Stage LLM Evaluation