Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World

Domaine:

natural language processingpeace and security

Type de record:

paperdataset
Créateur:
SemZhaZhaLi,
Hôte:avatar
This paper presents LEMONADE, a large-scale conflict event dataset comprising 39,786 events across 20 languages and 171 countries, with extensive coverage of region-specific entities. LEMONADE is based on a partially reannotated subset of the Armed Conflict Location & Event Data (ACLED), which has documented global conflict events for over a decade. To address the challenge of aggregating multilingual sources for global event analysis, we introduce abstractive event extraction (AEE) and its subtask, abstractive entity linking (AEL). Unlike conventional span-based event extraction, our approach detects event arguments and entities through holistic document understanding and normalizes them across the multilingual dataset. We evaluate various large language models (LLMs) on these tasks, adapt existing zero-shot event extraction systems, and benchmark supervised models. Additionally, we introduce ZEST, a novel zero-shot retrieval-based system for AEL. Our best zero-shot system achieves an end-to-end F1 score of 58.3%, with LLMs outperforming specialized event extraction models such as GoLLIE. For entity linking, ZEST achieves an F1 score of 45.7%, significantly surpassing OneNet, a state-of-the-art zero-shot baseline that achieves only 23.7%. However, these zero-shot results lag behind the best supervised systems by 20.1% and 37.0% in the end-to-end and AEL tasks, respectively, highlighting the need for further research. Findings of ACL 2025

Visit

arxiv.org

Tasks

information extraction

Tags

Computation and Language

Similaires

XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 LanguagesTIGQA:An Expert Annotated Question Answering Dataset in TigrinyaCSympData: An Expert-Annotated Dataset of Symptoms and Demographics for Clinical RecommendationsSYNAPSE: An Expert-Annotated Dataset of Symptoms and Demographics for Triage RecommendationsMPBD-18: A Large-Scale Real-World Medicinal Plant Image Dataset from Bangladesh for Automated Plant Species IdentificationTowards Globally Inclusive Multilingual Dialogue Systems for Real-World Applications

XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages

Contemporary works on abstractive text summarization have focused primarily on highresource languages like English, mostly due to the limited availability of datasets for low/midresource ones. In this work, we present XLSum, a comprehensive and diverse dataset comp

TIGQA:An Expert Annotated Question Answering Dataset in Tigrinya

The absence of explicitly tailored, accessible annotated datasets for educational purposes presents

CSympData: An Expert-Annotated Dataset of Symptoms and Demographics for Clinical Recommendations

Access to timely and accurate healthcare guidance remains a challenge, particularly in Low- and Midd

SYNAPSE: An Expert-Annotated Dataset of Symptoms and Demographics for Triage Recommendations

Access to timely and accurate healthcare guidance remains a challenge, particularly in Low- and Midd

MPBD-18: A Large-Scale Real-World Medicinal Plant Image Dataset from Bangladesh for Automated Plant Species Identification

MPBD-18 is a large-scale real-world medicinal plant image dataset collected from diverse environment

Towards Globally Inclusive Multilingual Dialogue Systems for Real-World Applications

With the advent of large language models (LLMs), dialogue systems have become the primary interface