Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BenCoref: A Multi-Domain Dataset of Nominal Phrases and Pronominal Reference Annotations

Domain:

natural language processing

Record type:

paperdataset
Creator:
RohHosRasMoh
Host:avatar
Coreference Resolution is a well studied problem in NLP. While widely studied for English and other resource-rich languages, research on coreference resolution in Bengali largely remains unexplored due to the absence of relevant datasets. Bengali, being a low-resource language, exhibits greater morphological richness compared to English. In this article, we introduce a new dataset, BenCoref, comprising coreference annotations for Bengali texts gathered from four distinct domains. This relatively small dataset contains 5200 mention annotations forming 502 mention clusters within 48,569 tokens. We describe the process of creating this dataset and report performance of multiple models trained using BenCoref. We expect that our work provides some valuable insights on the variations in coreference phenomena across several domains in Bengali and encourages the development of additional resources for Bengali. Furthermore, we found poor crosslingual performance at zero-shot setting from English, highlighting the need for more language-specific resources for this task.

Visit

arxiv.org

Tasks

coreference resolution

Tags

Computation and Language

Similar

Nominal and Pronominal Coordination in GreboPronominal System and Reference in PulaarNominal Phrases in LongudaYoNER: A New Yorùbá Multi-domain Named Entity Recognition DatasetReference assignment in pronominal argument languagesMURAD: A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset

Nominal and Pronominal Coordination in Grebo

Pronominal System and Reference in Pulaar

Nominal Phrases in Longuda

This scholarly endeavor delves into a thorough investigation of nominal phrases in Longuda, a minori

YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset

Named Entity Recognition (NER) is a foundational NLP task, yet research in Yorùbá has been constrain

Reference assignment in pronominal argument languages

Using data from Toposa and Kiswahili, this article demonstrates that the reference assignment in pro

MURAD: A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset

Arabic is a linguistically and culturally rich language with a vast vocabulary that spans scientific