Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SofPavJayWee
Hôte:avatar
Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative expressions of high-resource languages, it often faces challenges when dealing with low-resource languages like Sinhala due to limited available data. To address this limitation, we introduce a corpus of 2,344 Sinhala figures of speech with cultural and cross-lingual annotations. We examine this dataset to classify the cultural origins of the figures of speech and to identify their cross-lingual equivalents. Additionally, we have developed a binary classifier to differentiate between two types of FOS in the dataset, achieving an accuracy rate of approximately 92%. We also evaluate the performance of existing LLMs on this dataset. Our findings reveal significant shortcomings in the current capabilities of LLMs, as these models often struggle to accurately convey idiomatic meanings. By making this dataset publicly available, we offer a crucial benchmark for future research in low-resource NLP and culturally aware machine translation. 19 pages, 6 figures, 8 tables, Accepted paper at the 22nd Workshop on Multiword Expressions (MWE 2026) @ EACL 2026

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language

Similaires

Sinhala-English Parallel Word Dictionary DatasetKasem Speech-Text Parallel DatasetVai Speech-Text Parallel DatasetGa Speech-Text Parallel DatasetDeg Speech-Text Parallel DatasetVagla Speech-Text Parallel Dataset

Sinhala-English Parallel Word Dictionary Dataset

Parallel datasets are vital for performing and evaluating any kind of multilingual task. However, in

Kasem Speech-Text Parallel Dataset

This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Gha

Vai Speech-Text Parallel Dataset

This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana

Ga Speech-Text Parallel Dataset

This dataset is made available because of Ghana NLP's volunteer driven research work. Please conside

Deg Speech-Text Parallel Dataset

This dataset contains 125958 parallel speech-text pairs for Deg, a language spoken primarily in Ghan

Vagla Speech-Text Parallel Dataset

This dataset contains 48605 parallel speech-text pairs for Vagla, a language spoken primarily in Gha