Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BANTEN: A Parallel Banglish–English Dataset for Machine Translation

Domain:

natural language processing

Record type:

dataset
Creator:
Mia
Publisher:
Men
Host:avatar
BANTEN is a manually curated parallel Banglish–English dataset comprising 14,000 sentence pairs collected from publicly available online sources, including newspapers, Facebook posts and comments, YouTube comments, daily conversations, and blogs. Each instance contains a Banglish sentence written in Roman script and its corresponding human-translated English sentence. The dataset was developed through data collection, filtering, cleaning, manual translation, and expert validation. It is intended to support research in Banglish-to-English machine translation, code-mixed language processing, transliteration, and low-resource natural language processing.

Visit

doi.org

Tasks

machine translation

Tags

Computer ScienceNatural Language ProcessingMachine Translation

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode