Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection

Domain:

natural language processing

Record type:

paperdataset
Creator:
GhoSalTiwEkb
Host:avatar
Detecting novelty of an entire document is an Artificial Intelligence (AI) frontier problem that has widespread NLP applications, such as extractive document summarization, tracking development of news events, predicting impact of scholarly articles, etc. Important though the problem is, we are unaware of any benchmark document level data that correctly addresses the evaluation of automatic novelty detection techniques in a classification framework. To bridge this gap, we present here a resource for benchmarking the techniques for document level novelty detection. We create the resource via event-specific crawling of news documents across several domains in a periodic manner. We release the annotated corpus with necessary statistics and show its use with a developed system for the problem in concern. Accepted for publication in Language Resources and Evaluation Conference (LREC) 2018

Visit

arxiv.org

Tasks

text classification

Tags

Computation and Language

Similar

AFRIDOC-MT: Document-level MT Corpus for African LanguagesOpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource LanguagesNovelty detection in event surveillance documentsMada-French Parallel Corpus 1.0TypeCraft Akan Corpus, Release 1.0DocHPLT: A Massively Multilingual Document-Level Translation Dataset

AFRIDOC-MT: Document-level MT Corpus for African Languages

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-t

OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages

In machine translation (MT), health is a high-stakes domain characterised by widespread deployment a

Novelty detection in event surveillance documents

Source Agritrop Cirad (https://agritrop.cirad.fr/617038/) International audience Even

Mada-French Parallel Corpus 1.0

This dataset comprises a parallel corpus of Mada–French literary text translations totalling 2,154 l

TypeCraft Akan Corpus, Release 1.0

The Release 1.0 of the TC Akan Corpus consists of 41 short texts, mostly linguistic sentence collections, corresponding to 669 sentences. Two of the released texts are transcribed recordings of students narrating a video. The students doing the original work wer

DocHPLT: A Massively Multilingual Document-Level Translation Dataset

Existing document-level machine translation resources are only available for a handful of languages,