Logo Lanfrica

liulalemx/felig-toolkit

Domaine:

natural language processing

Type de record:

software
Créateur:
liu
Hôte:
A toolset for Amharic Language pre-processing. Includes an Amharic Stemmer, Transliterator, Stopword remover , Lexical analyzer, Corpus indexer and Term weighter. Felig Toolkit A toolset for Amharic Language pre-processing 🔧 Felig Toolkit Web ​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ Now with Typescript support! --- ## What is felig-toolkit? It is a toolset for Amharic Language pre-processing. It includes an Amharic Stemmer, Amharic Transliterator, Amharic Stopword remover, Amharic Lexical analyzer, Amharic Corpus indexer and Term weighter. ### Amharic Lexical Analyzer Breaks down Amharic language corpus and returns tokens by removing any whitespace, expanding abbreviations(`አ.አ -> አዲስ አበበ`), removing numbers, breaking up hyphenated words, and removing punctuation (`፡ ። ! ? `...). ### Amharic Stopword remover Removes commonly occuring words that have no contribution to the semantics of the corpus. Eg: `እና ፡ ስለዚህ ፡ በመሆኑም`... ### Amharic Transliterator Changes Unicode Amharic characters to ASCII. Exmaple: `ልጆች -> ልጅኦች -> ljoc`. This tool implements two types of Amharic transliteration lookup tables. - SERA (System for Ethiopic Representation in ASCII) - This system maps alphabets with similar sounds separately. Eg: `(ሀ፣ሐ፡ኀ)፣(ሰ፡ሠ)፡(ጸ፡ፀ)፡(ዐ፡አ)`. However, in practice, these alphabets are used interchangeably and use of SERA would greatly decrease recall. **NOT RECOMMENDED!** - Felig - Normalizes the redundant symbols into a common symbol. **RECOMMENDED!** ### Amharic Stemmer LIVE DEMO Reduces the different morphological (e.g. inflectional or derivational) variations of Amharic word forms by taking an Amharic word and returning the stem through affix-removal with longest match. Exmaple: `ልጆች -> ልጅኦች -> ljoc -> lj -> ልጅ` ### Amharic Corpus Indexer Produces an index file for the stemmed words in a corpus and relates them with the files they are found in. It also stores their frequencies per file. ### Term Weighter Calculates the weight of words from the index file using product of their length normalized Term frequency and Inverse document frequency (`tf*idf`). ## Installation Felig Toolkit is available as a pack …