Logo Lanfrica

Qaamuuska-NLP: A Structured 46K-Entry Somali Lexicon Extracted from a Print Dictionary

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AbdAbdAbdAbd
Éditeur:
Spr
Hôte:
Abstract Somali NLP still has relatively few structured lexical resources that preserve the linguistic information found in traditional dictionaries. We present Qaamuuska-NLP, a machine-readable reconstruction of Qaamuuska Af-Soomaaliga, the 2012 monolingual Somali dictionary by Puglielli and Mansuur. The resource contains 46,314 dictionary records extracted directly from the PDF's native text layer using a deterministic, rule-based pipeline, without OCR or machine-learning-based extraction. Where available in the source, records include part of speech, noun gender, verb class and transitivity, plural information, subject domains, synonym references, cross-references, homonym indices, and numbered definitions. The original entry text is retained alongside the parsed representation to support inspection against the source.We describe the extraction pipeline, the structure and coverage of the resulting resource, and the main classes of residual parsing errors. We also examine the distribution of grammatical and lexical information in the dictionary and use a simple surface-form prediction probe to test how strongly selected morphological labels are reflected in Somali word forms. Qaamuuska-NLP complements existing Somali morphological, lemmatization, corpus, and multi-dictionary lexical resources by preserving the structure of a single major monolingual dictionary in a reproducible computational form. The extraction code and aggregate statistics are being prepared for release; public redistribution of the complete structured lexical dataset remains subject to rightsholder permission.