Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BANTEN: A Parallel Banglish–English Dataset for Machine Translation

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Mia
Éditeur:
Men
Hôte:avatar
BANTEN is a manually curated parallel Banglish–English dataset comprising 14,000 sentence pairs collected from publicly available online sources, including newspapers, Facebook posts and comments, YouTube comments, daily conversations, and blogs. Each instance contains a Banglish sentence written in Roman script and its corresponding human-translated English sentence. The dataset was developed through data collection, filtering, cleaning, manual translation, and expert validation. It is intended to support research in Banglish-to-English machine translation, code-mixed language processing, transliteration, and low-resource natural language processing.

Visit

doi.org

Tasks

machine translation

Tags

Computer ScienceNatural Language ProcessingMachine Translation

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

English-Twi Parallel Corpus for Machine TranslationExtended Parallel Corpus for Amharic-English Machine TranslationParallel Corpora Preparation for English-Amharic Machine TranslationAmharic-English Parallel Corpus for Neural Machine TranslationA PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATIONAmhEn: Amharic-English Large Parallel Corpus for Machine Translation

English-Twi Parallel Corpus for Machine Translation

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected where necessary by native

Extended Parallel Corpus for Amharic-English Machine Translation

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for research purposes. Furthermore,

Parallel Corpora Preparation for English-Amharic Machine Translation

In this paper, we describe the development of an English-Amharic parallel corpus and Machine Translation (MT) experiments conducted on it. Two different tests have been achieved. Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) experiments

Amharic-English Parallel Corpus for Neural Machine Translation

Amharic is the working language of Ethiopia and, owing to its Semitic characteristics, the language

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Machine Translation (MT) poses a significant challenge in developing language corpora for low-resour

AmhEn: Amharic-English Large Parallel Corpus for Machine Translation

Recently, using deep neural networks for machine translation (MT) tasks has received great attention. In order for these networks to learn abstract representations of the input and store them as continuous vectors, they need a lot of data. However, very few researc