Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Dictionaries for Under-Resourced Languages: from Published Files to Standardized Resources Available on the Web

Domain:

natural language processing

Record type:

paperdatasetsoftware
Creator:
ManEng
Editor:
LabUniGroLab
Publisher:
CCSD
Host:avatar
Most work in the feld of natural language processing focuses on well-resourced languages. However, much remains to be done on under-resourced ones: there are few dictionaries, parsers, etc. Nevertheless, when published dictionaries are available, it is sometimes possible to fnd the data fles used to print the dictionar. (usuall. in Word format). A conversion process can then be applied to these fles in order to obtain standardized XML lexical data. Attention must be paid to specifc problems such as a lack of standardization in the alphabets or the use of hacked fonts for displa.ing specifc characters. Next, the standardized XML data can be imported into an online lexical resources management platform. It is then available online for lookup and editing. A fnal step can also be performed to automaticall. export the data into interchange formats such as Lexical Markup Framework or lemon in order to produce linked data.

Visit

hal.science

Tags

Jibiki platformXMLLMFBambaraKhmerWolofNigerNational Language Processinglexical databaseDictionary+1

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

A primer on getting neologisms from foreign languages to under-resourced languagesDatasheets for Under-resourced Languages: An ExampleToward More Meaningful Resources for Lower-resourced LanguagesThe Multilingual Semantic Web as Virtual Knowledge Commons: The Case of the Under-Resourced South African LanguagesStrategies for building wordnets for under-resourced languages: The case of African languagesEnd-to-End Text-To-Speech synthesis for under resourced South African languages

A primer on getting neologisms from foreign languages to under-resourced languages

Mainly due to lack of support, most under-resourced languages have a reduced lexicon in most realms

Datasheets for Under-resourced Languages: An Example

The datasheet provides an example of how to use the Datasheet standard for describing and sharing un

Toward More Meaningful Resources for Lower-resourced Languages

In this position paper, we describe our perspective on how meaningful resources for lower-resourced

The Multilingual Semantic Web as Virtual Knowledge Commons: The Case of the Under-Resourced South African Languages

Strategies for building wordnets for under-resourced languages: The case of African languages

The African Wordnet Project (AWN) aims at building wordnets for five African languages: Setswana, is

End-to-End Text-To-Speech synthesis for under resourced South African languages