Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Afrikaans word2vec embeddings trained on OpenSubtitles

Domain:

natural language processing

Record type:

datasetmodel
Creator:
GriBuc
Publisher:
Zenodo
Host:avatar
This dataset contains the subs2vec embeddings for Afrikaans, as presented in zenodo.org. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: opus.nlpl.eu.  For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of: Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …) Window size: varying context windows (e.g., 2, 5, 10, …) Each file corresponds to a unique configuration (dimension × window size).  Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1).  If you use this dataset, please cite: Manuscript: doi.org  Data: This Zenodo dataset (using the DOI provided here)

Visit

doi.orgzenodo.org

Tasks

embeddings

Languages

Afrikaans

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode