Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Building the Oranian-English Parallel Corpus: Methodology and Compilation Process

Domain:

natural language processing

Record type:

dataset
Creator:
AbdKha
Publisher:
Has
Host:
The scarcity of linguistic resources poses a major challenge for automated translation and processing of dialects. These resources are crucial for natural language processing experts conducting research on dialect recognition, processing, and machine translation. This paper describes the compilation of a dataset for an Algerian low-resource language as it emphasizes the importance of developing resources for Algerian dialects. It examines existing relevant corpora and details the creation process and unique features of the pioneering Oranian-English Parallel Corpus (OEPC). OEPC is the first parallel corpus built from scratch that pairs an Algerian dialect with its English counterparts. The paper outlines the criteria and steps involved in compiling a monolingual corpus for the Oranian dialect (ORN), including data sources and formats. ORN comprises 8500 sentences, which were then translated into English to form OEPC. This valuable linguistic resource is a product of the ERAD project, an initiative aimed at providing NLP professionals with diverse Algerian mono-, multi-, and cross-dialectal corpora. The paper also explains the data compilation and augmentation techniques used to expand the project's outputs.

Visit

doi.org

Tasks

machine translation

Languages

Arabic, Algerian Spoken

Licenses

https://creativecommons.org/licenses/by/4.0

Similar

Building an Oranian-English parallel corpus for automated translation trainingBuilding a Parallel Corpus and Training Translation Models Between Luganda and EnglishAkan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel CorpusThe SAWA Corpus: a Parallel Corpus English - SwahiliSomali-English Parallel CorpusNupe-English parallel corpus

Building an Oranian-English parallel corpus for automated translation training

Abstract The main obstacle to automated translation and processing of dialects is t

Building a Parallel Corpus and Training Translation Models Between Luganda and English

Neural machine translation (NMT) has achieved great successes with large datasets, so NMT is more premised on high-resource languages. This continuously underpins the low resource languages such as Luganda due to the lack of high-quality parallel corpora, so even ‘

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel Corpus

This dataset contains a curated bilingual parallel corpus developed to support domain-specific neura

The SAWA Corpus: a Parallel Corpus English - Swahili

Research in data-driven methods for Machine Translation has greatly benefited from the increasing availability of parallel corpora. Processing the same text in two different languages yields useful information on how words and phrases are translated from a source l

Somali-English Parallel Corpus

This dataset contains high-quality parallel sentence pairs, multi-sentence alignments, and paragraph

Nupe-English parallel corpus

This is the first ever Nupe - English Parallel Corpus and Nupe Monolingual Corpora curated from diverse sources including poems,idioms, proverbs, religpoius text etc. The aim of this data collection is to make available a cultural-aware Nupe-english corpus for NLP