Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

VOC GM NER corpus

Domain:

natural language processing

Record type:

dataset
Editor:
PetVosRoode
Publisher:
VU
Host:avatar
Corpus and training data for Named-Entity Recognition from the VOC General Letters. The corpus consist of a selection of letters from the Generale Missiven, a subset of the Overgebleven Brieven en Papieren corpus of the United East India Company (VOC). These letters were reports sent by governor generals and administrators of the VOC to the board, from locations where the VOC was active (Indonesia and other parts of Asia as well as South Africa). The letters for the current corpus were edited and digitalized by the Huygens Institute of Netherlands History between 1960 and 2007 as part of the Rijks Geschiedkundige Publicatiën (RGP) series. In this edition, letters were transcribed in part, while other parts were summarized. The data in the current package consist of a selection of these letters, spread in time, where the original text and modern additions (notes and passage summaries) are extracted into separate documents to allow for training on either the historical text or modern additions. The entities identified in the data are: persons, locations, organisations (mainly the VOC itself) and ships. These are completed with forms derived from location or religion names. The 'corpus' folder contains files for the historical text and modern notes of each letter, in CoNLL 2002 format, taking paragraphs or separate notes as units for segmentation. The 'datasplit_all_standard' folder contains training, validation and test data for the 'standard' NER experiment on all the data referred in the companion publication, splitting sequences longer than 256 subtokens. For more information, see Arnoult et. al, 2021. Batavia asked for advice. Pretrained language models for Named Entity Recognition in historical texts. In Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature. Code, intermediary data and more information on the collection process can be found on the cltl/voc-missives package on Zenodo and GitHub.

Visit

doi.orgpublication.yoda.vu.nl

Tasks

information extractionnamed entity recognition

Tags

Natural Sciences - Computer and information sciences (1.2)Humanities - History and archaeology (6.1)Humanities - Languages and literature (6.2)NERDigital HumanitiesEarly modern Dutch

Licenses

Open - freely retrievableinfo:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 International Public Licensehttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Isizulu Ner CorpusSetswana Ner CorpusSesotho Ner CorpusAfrikaans Ner CorpusSiswati Ner CorpusIsixhosa Ner Corpus

Isizulu Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Setswana Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Sesotho Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Afrikaans Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Siswati Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Isixhosa Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.