Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Meidad1/Khoekhoe-Morphological-Analyzer

Domaine:

natural language processing

Type de record:

software
Créateur:
Mei
Hôte:
This script performs morphological analysis (including segmentation into morphemes) of written and spoken texts Khoekhoe. # Khoekhoe Grammatical Analyzer ## Description This project is responsible for the processing pipeline of spoken and written texts (given as ELAN files) in Khoekhoe, as part of a corpus construction in a linguistic project of the Hebrew University of Jerusalem. The analyzer consists of two main modules: preprocessing (cleaning and morpheme segmentation) and glossing (i.e., annotation of morphological functions and part-of-spheech tagging). ### Preprocessing This package handles the preprocessing of texts in Khoekhoe, which includes the following steps: * Validation of character encodings * Copy \tx tiers to their corresponding \orig tiers * Cleaning the tx annotations: - Punctuation cleaning - Decapitalization of words which are not proper names - Spelling check and correction * Morpheme segmentation of the text in all tx tiers ### Glossing This code section consists of _single_eaf_glosser.py_, _glossing_dict_generator.py_ and _ambiguity_removing.py_, which handle the process of glossing the texts in Khoekhoe, i.e., performing linguistic analysis. It fills the morphological meaning/function and part-of-speech of morphemes in \ge and \ps tiers of the input ELAN files. ## Usage **INPUT** The folder _input_ should contain the following files: 1. All ELAN files (.eaf) you want to gloss. _pfsx_ files are allowed to be in this folder, they are ignored. Each eaf file should be time aligned and contain annotations in \tx (text) and \fte (translation) tiers. In addition, backchannel notations should be written in \fte tiers. It is also assumed that all the input files are created with the project's templates (that define the tiers structures). 2. _ISF_Khoekhoe_dictionary_for_glosses.xlsx_ file which contains a glossing dictionary of Khoekhoe. This file must contain a sheet called: "KK_dict_for_glosses". 3. _misspellings_correction.json_: includes all the replacements needed for misspelled forms in the texts. Be careful when editing it and do not change its structure. 4. …

Visit

github.com

Tasks

part of speech tagging

Languages

Khoekhoe

Similaires

kipprice/morphological-analyzerTewodrosAbebe/Wolaytta-Morphological-AnalyzerGeek0254/Iteso-Morphological-AnalyzerWeb-Based Arabic Morphological AnalyzerAdrian-Silver/Gikuyu-Morphological-AnalyzerNadiaBMKarmani/Tunisian-Arabic-Morphological-Analyzer-evaluation-corpus

kipprice/morphological-analyzer

Morphological analyzers for Xhosa and Japanese Old Morphological files Written in XFST Author: Kip

TewodrosAbebe/Wolaytta-Morphological-Analyzer

Learning morphology for Goffa using Morphological Analyzer of Wolaytta.

Geek0254/Iteso-Morphological-Analyzer

# Iteso-Morphological-Analyzer This repository contains a morphological analyzer model for the Teso

Web-Based Arabic Morphological Analyzer

Adrian-Silver/Gikuyu-Morphological-Analyzer

A simple Gikuyu Morphological Analyzer using XFST tools (foma) # Gikuyu-Morphological-Analyzer A si

NadiaBMKarmani/Tunisian-Arabic-Morphological-Analyzer-evaluation-corpus

1000 Tunisian Arabic words