Logo Lanfrica

NCHLT Tshivenḓa RoBERTa language model

Domaine:

natural language processing

Type de record:

model
Créateur:
Roald Eiselen
Éditeur:
Rico KoenAlbertus KrugerJacques van Heerden
Éditeur:
North-West University - Centre for Text Technology (CTexT)
Hôte:avatar
Contextual masked language model based on the RoBERTa architecture (Liu et al., 2019). The model is trained as a masked language model and not fine-tuned for any downstream process. The model can be used both as a masked LM or as an embedding model to provide real-valued vectorised respresentations of words or string sequences for Tshivenḓa text.