Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Afrikaans Domain corpus POS annotated (5 domains)

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Gaustad, Tanja
Éditeur:
McKellar, CindyGent, Sunny
Éditeur:
North-West University - Centre for Text Technology (CTexT)
Hôte:avatar
This deliverable contains part-of-speech tagged data from five different text types for Afrikaans. The text types included are: - CAPS gr12 (Academic) - MA/PhD Theses (Academic) - Magazines (Non-Academic) - News (Non-Academic) - Novels (Fiction) The data is given as txt files where each line contains a token and the corresponding POS tag, tab separated. Each text type data file contains 11,000+ tokens, amounting to a total of 60,809 tokens for the language. Please see the included protocol for more details on the POS tags used.

Visit

hdl.handle.net

Tasks

part of speech tagging

Languages

Afrikaans

Tags

Afrikaans, POS annotated, domain-specific, annotated corpus

Licenses

Creative Commons Attribution 4.0 International