Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Contextualising Levels of Language Resourcedness that affect NLP tasks

Domain:

natural language processing

Record type:

paper
Creator:
KeeKhumalo, Langa
Host:avatar
Several widely used software applications involve some form of processing of natural language, with tasks ranging from digitising hardcopies and text processing to speech generation. Varied language resources are used to develop software systems to accomplish a wide range of natural language processing (NLP) tasks, such as the ubiquitous spellcheckers and chatbots. Languages are typically characterised as either low (LRL) or high resourced languages (HRL) with African languages having been characterised as resource-scarce languages and English by far the most well-resourced language. But what lies in-between? We argue that the dichotomous typology of LRL and HRL for all languages is problematic. Through a clear understanding of language resources situated in a society, a matrix is developed that characterises languages as Very LRL, LRL, RL, HRL and Very HRL. The characterisation is based on the typology of contextual features for each category, rather than counting tools. The motivation is provided for each feature and each characterisation. The contextualisation of resourcedness, with a focus on African languages in this paper, and an increased understanding of where on the scale the language used in a project is, may assist in, among others, better planning of research and implementation projects. We thus argue in this paper that the characterisation of language resources within a given scale in a project is an indispensable component, particularly for those in the lower half of the scale. 27 pages, 2 tables

Visit

arxiv.org

Tags

Computation and LanguageI.2.7

Similar

Urhobo language training text for NLP, ASR and TTS tasksMultilingual Prompt Engineering in Large Language Models: A Survey Across NLP TasksNiger Volta LTI: Yorùbá language training text for NLP, ASR and TTS tasksNiger Volta LTI: Ị̀gbò language training text for NLP, ASR and TTS tasksNiger Volta LTI: Fon language training text for NLP, ASR and TTS tasksHow does scaling the number of intermediate language-understanding tasks affect the performance of zero-shot cross-lingual

Urhobo language training text for NLP, ASR and TTS tasks

Urhobo language training text for NLP, ASR and TTS tasks

Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks

Large language models (LLMs) have demonstrated impressive performance across a wide range of Natural

Niger Volta LTI: Yorùbá language training text for NLP, ASR and TTS tasks

Yorùbá language training text for NLP, ASR and TTS tasks

Niger Volta LTI: Ị̀gbò language training text for NLP, ASR and TTS tasks

Igbo language training text for NLP, ASR and TTS tasks

Niger Volta LTI: Fon language training text for NLP, ASR and TTS tasks

Fon language training text for NLP, ASR and TTS tasks

How does scaling the number of intermediate language-understanding tasks affect the performance of zero-shot cross-lingual

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni