Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Building an Oranian-English parallel corpus for automated translation training

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AbdKha
Éditeur:
Oxf
Hôte:
Abstract The main obstacle to automated translation and processing of dialects is their dearth of linguistic resources. The latter provide data to natural language processing professionals to conduct their experiments of dialect recognition, processing, and machine translation. This article highlights the need to resource the Algerian dialects, reviews the use of the available relevant corpora, and describes the process and distinctiveness of the first Oranian-English parallel corpus (OEPC). This is the first parallel corpus that includes one Algerian dialect with its English equivalents made from scratch. Particularly, this article presents the criteria and steps of compiling a monolingual corpus for the Oranian dialect (ORN) with references to data sources and formats. The size of the monolingual corpus ORN reached 8.5K sentences; with their equivalents in English, OEPC has been built. This significant linguistic resource is made under the Empowering and Resourcing Algerian Dialects project. This project is launched to enrich NLP experts with linguistic resources that are different Algerian mono-, bi-, multi-, and cross-dialectal corpora. The mechanism of data compilation and augmentation to extend the products of this project is explained.

Visit

doi.org

Tasks

machine translation

Languages

Arabic, Algerian Spoken

Licenses

https://academic.oup.com/pages/standard-publication-reuse-rights

Similaires

Building the Oranian-English Parallel Corpus: Methodology and Compilation ProcessBuilding a Parallel Corpus and Training Translation Models Between Luganda and EnglishEnglish-Twi Parallel Corpus for Machine TranslationExtended Parallel Corpus for Amharic-English Machine TranslationAmharic-English Parallel Corpus for Neural Machine TranslationAmhEn: Amharic-English Large Parallel Corpus for Machine Translation

Building the Oranian-English Parallel Corpus: Methodology and Compilation Process

The scarcity of linguistic resources poses a major challenge for automated translation and processin

Building a Parallel Corpus and Training Translation Models Between Luganda and English

Neural machine translation (NMT) has achieved great successes with large datasets, so NMT is more premised on high-resource languages. This continuously underpins the low resource languages such as Luganda due to the lack of high-quality parallel corpora, so even ‘

English-Twi Parallel Corpus for Machine Translation

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected where necessary by native

Extended Parallel Corpus for Amharic-English Machine Translation

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for research purposes. Furthermore,

Amharic-English Parallel Corpus for Neural Machine Translation

Amharic is the working language of Ethiopia and, owing to its Semitic characteristics, the language

AmhEn: Amharic-English Large Parallel Corpus for Machine Translation

Recently, using deep neural networks for machine translation (MT) tasks has received great attention. In order for these networks to learn abstract representations of the input and store them as continuous vectors, they need a lot of data. However, very few researc