Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MigraAnno: Creating a topic-specific corpus from a large newspaper collection

Domain:

natural language processing

Record type:

dataset
Creator:
Bro
Publisher:
Zenodo
Host:avatar

This poster presents MigraAnno, a topic-specific corpus of migration- and minority-related discourse extracted from Austrian newspapers (1703–1938) using seeded BERTopic modelling, human evaluation, and transformer-based relevancy classification. The resulting 96,871-sentence corpus is openly available and supports historical, linguistic, and computational research on migration and minority discourses.

Visit

doi.org

Tasks

topic classificationtext classification

Languages

Ndasa

Tags

Unsupervised Machine LearningSupervised Machine LearningCorpus CreationNewspapers as TopicHuman Migration

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Automatically Creating a Large Number of New Bilingual DictionariesSAE Newspaper Text CorpusTranslation features in a comparable corpus of Afrikaans newspaper articlesVerb-Noun Collocations In Newspaper Editorials In Ghana: A Corpus-Based Analysisjkayobotsi/kinyarwanda-topic-afroxlmr-largeCreating a Parallel Corpus for Machine Translation: A Case Study of Kru and Krio

Automatically Creating a Large Number of New Bilingual Dictionaries

This paper proposes approaches to automatically create a large number of new bilingual dictionaries

SAE Newspaper Text Corpus

Newspaper text in electronic format obtained from Avusa Media through a licensing agreement (renewed

Translation features in a comparable corpus of Afrikaans newspaper articles

Verb-Noun Collocations In Newspaper Editorials In Ghana: A Corpus-Based Analysis

This paper is a corpus-based study which aims at profiling the most frequent verb-noun collocations

jkayobotsi/kinyarwanda-topic-afroxlmr-large

Creating a Parallel Corpus for Machine Translation: A Case Study of Kru and Krio