Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Afrikaans word2vec embeddings trained on OpenSubtitles

Domaine:

natural language processing

Type de record:

datasetmodel
Créateur:
GriBuc
Éditeur:
Zenodo
Hôte:avatar
This dataset contains the subs2vec embeddings for Afrikaans, as presented in zenodo.org. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: opus.nlpl.eu.  For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of: Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …) Window size: varying context windows (e.g., 2, 5, 10, …) Each file corresponds to a unique configuration (dimension × window size).  Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1).  If you use this dataset, please cite: Manuscript: doi.org  Data: This Zenodo dataset (using the DOI provided here)

Visit

doi.orgzenodo.org

Tasks

embeddings

Languages

Afrikaans

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode