Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

GiellaLT: an infrastructure for rule-based language technology tool development

Domain:

natural language processing

Record type:

software
Creator:
FlaTroSjuLin
Publisher:
Dig
Host:
Currently, machine learning is presented as the ultimate solution for language technology regardless of use case and application, however, it requires as a starting point a massive amount of curated linguistic data in electronic form that is expected to be high quality and representative of the kind of language usage that the tools will follow. For minority and indigenous languages, this can be an insurmountable task, as digital materials of the necessary sizes do not exist and can not easily be produced. In this article we present an approach we have successfully used for supporting indigenous languages to survive and grow in digital contexts for years, and describe the potential of our approach for African contexts. Our technological solution is a free and open-source infrastructure that enables language experts and users to cooperate on creating linguistic resources like dictionaries and grammatical descriptions. In addition we provide language-independent frameworks to build these into applications that are needed by the language community.

Visit

doi.org

Similar

UzbekTagger: The rule-based POS tagger for Uzbek languageDevelopment of an Evidence-Based Clinical Tool for Childhood Cancer Survivorship Care in TanzaniaLexicon and Rule-based Word Lemmatization Approach for the Somali LanguageDevelopment of a rule-based Yorùbá numerals to digits translatorDevelopment of a rule-based Arabic numerals to Yorùbá translatorTowards Development of an Indigenous African Language-based Programming Language

UzbekTagger: The rule-based POS tagger for Uzbek language

This research paper presents a part-of-speech (POS) annotated dataset and tagger tool for the low-re

Development of an Evidence-Based Clinical Tool for Childhood Cancer Survivorship Care in Tanzania

CONCLUSIONS Three in every four childhood cancer survivors experience late effects

Lexicon and Rule-based Word Lemmatization Approach for the Somali Language

Lemmatization is a Natural Language Processing (NLP) technique used to normalize text by changing mo

Development of a rule-based Yorùbá numerals to digits translator

The importance of numbers in commerce, education, and everyday conversation cannot be overemphasised

Development of a rule-based Arabic numerals to Yorùbá translator

Number conversion and normalisation are among major pre-processing steps required by a text-to-speec

Towards Development of an Indigenous African Language-based Programming Language

Programming languages based on the lexicons of indigenous African languages are rare to come by unli