Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Viability of Neural Networks for Core Technologies for Resource-Scarce Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
MelMar
Éditeur:
MDP
Hôte:
In this paper, the viability of neural network implementations of core technologies (the focus of this paper is on text technologies) for 10 resource-scarce South African languages is evaluated. Neural networks are increasingly being used in place of other machine learning methods for many natural language processing tasks with good results. However, in the South African context, where most languages are resource-scarce, very little research has been done on neural network implementations of core language technologies. In this paper, we address this gap by evaluating neural network implementations of four core technologies for ten South African languages. The technologies we address are part of speech tagging, named entity recognition, compound analysis and lemmatization. Neural architectures that performed well on similar tasks in other settings were implemented for each task and the performance was assessed in comparison with currently used machine learning implementations of each technology. The neural network models evaluated perform better than the baselines for compound analysis, are viable and comparable to the baseline on most languages for POS tagging and NER, and are viable, but not on par with the baseline, for Afrikaans lemmatization.

Visit

doi.org

Tasks

information extractionnamed entity recognitionpart of speech tagging

Languages

Afrikaans

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Developing Core Technologies for Resource-Scarce Nguni LanguagesNLP Web Services for Resource-Scarce LanguagesBuilding CLIA for Resource-Scarce African LanguagesCore technologies for conjunctively written South African languagesA word‐level language identification strategy for resource‐scarce languagesText Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: Tigrinya

Developing Core Technologies for Resource-Scarce Nguni Languages

The creation of linguistic resources is crucial to the continued growth of research and development

NLP Web Services for Resource-Scarce Languages

In this paper, we present a project where existing text-based core technologies were ported to Java-based web services from various architectures. These technologies were developed over a period of eight years through various government funded projects for 10 resou

Building CLIA for Resource-Scarce African Languages

Since most of the existing major search engines and commercial Information Retrieval (IR) systems ar

Core technologies for conjunctively written South African languages

During this SADiLaR funded project, enriched corpora for the four official South African languages w

A word‐level language identification strategy for resource‐scarce languages

ABSTRACT This study is based on the premise that it is possible to train compute

Text Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: Tigrinya

This article studies convolutional neural networks for Tigrinya (also referred to as Tigrigna), whic