Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

OasisSimp: An Open-source Asian-English Sentence Simplification Dataset

Domain:

natural language processing

Record type:

paperdataset
Creator:
LiuTiaAliGao
Host:avatar
Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains limited for mid-resource and low-resource languages due to the scarcity of high-quality data. To address this gap, we introduce the OasisSimp dataset, a multilingual dataset for sentence-level simplification covering five languages: English, Sinhala, Tamil, Pashto, and Thai. Among these, no prior sentence simplification datasets exist for Thai, Pashto, and Tamil, while limited data is available for Sinhala. Each language simplification dataset was created by trained annotators who followed detailed guidelines to simplify sentences while maintaining meaning, fluency, and grammatical correctness. We evaluate eight open-weight multilingual Large Language Models (LLMs) on the OasisSimp dataset and observe substantial performance disparities between high-resource and low-resource languages, highlighting the simplification challenges in multilingual settings. The OasisSimp dataset thus provides both a valuable multilingual resource and a challenging benchmark, revealing the limitations of current LLM-based simplification methods and paving the way for future research in low-resource sentence simplification. The dataset is available at oasissimpdataset.github.io. Accepted at LREC 2026

Visit

arxiv.org

Tags

Computation and Language

Similar

English-Giriama Parallel Sentence DatasetSagalee: An Open Source ASR Dataset for Oromo LanguageSagalee: An Open Source ASR Dataset for Oromo LanguageAn Open-Source Annotated Dataset of Ghana Currency Images Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo LanguageAmgyna/Africa-text-Open-Source-Dataset-

English-Giriama Parallel Sentence Dataset

This dataset consists of sentence pairs in English and their corresponding translations in Giriama (

Sagalee: An Open Source ASR Dataset for Oromo Language

Sagalee is Speech Recognition Dataset for Oromo language Presented in the paper: Sagalee: an Open So

Sagalee: An Open Source ASR Dataset for Oromo Language

Sagalee is Speech Recognition Dataset for Oromo language Presented in the paper: Sagalee: an Open So

An Open-Source Annotated Dataset of Ghana Currency Images

The field of deep learning has led to remarkable advancements in many areas, including banking. Iden

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoke

Amgyna/Africa-text-Open-Source-Dataset-

# Africa-text-Open-Source-Dataset- ## 🇹🇿Maelezo ya Kiswahili Karibu kwenye **African Text Datasets*