Logo Lanfrica

liulalemx/felig-toolkit

Domain:

natural language processing

Record type:

software
Creator:
liu
Host:
A toolset for Amharic Language pre-processing. Includes an Amharic Stemmer, Transliterator, Stopword remover , Lexical analyzer, Corpus indexer and Term weighter. Felig Toolkit A toolset for Amharic Language pre-processing πŸ”§ Felig Toolkit Web ​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ Now with Typescript support! --- ## What is felig-toolkit? It is a toolset for Amharic Language pre-processing. It includes an Amharic Stemmer, Amharic Transliterator, Amharic Stopword remover, Amharic Lexical analyzer, Amharic Corpus indexer and Term weighter. ### Amharic Lexical Analyzer Breaks down Amharic language corpus and returns tokens by removing any whitespace, expanding abbreviations(`አ.አ -> αŠ α‹²αˆ΅ αŠ α‰ α‰ `), removing numbers, breaking up hyphenated words, and removing punctuation (`ፑ ፒ ! ? `...). ### Amharic Stopword remover Removes commonly occuring words that have no contribution to the semantics of the corpus. Eg: `αŠ₯αŠ“ ፑ αˆ΅αˆˆα‹šαˆ… ፑ α‰ αˆ˜αˆ†αŠ‘αˆ`... ### Amharic Transliterator Changes Unicode Amharic characters to ASCII. Exmaple: `αˆαŒ†α‰½ -> αˆαŒ…αŠ¦α‰½ -> ljoc`. This tool implements two types of Amharic transliteration lookup tables. - SERA (System for Ethiopic Representation in ASCII) - This system maps alphabets with similar sounds separately. Eg: `(αˆ€α£αˆα‘αŠ€)፣(ሰፑሠ)ፑ(αŒΈα‘α€)ፑ(α‹α‘αŠ )`. However, in practice, these alphabets are used interchangeably and use of SERA would greatly decrease recall. **NOT RECOMMENDED!** - Felig - Normalizes the redundant symbols into a common symbol. **RECOMMENDED!** ### Amharic Stemmer LIVE DEMO Reduces the different morphological (e.g. inflectional or derivational) variations of Amharic word forms by taking an Amharic word and returning the stem through affix-removal with longest match. Exmaple: `αˆαŒ†α‰½ -> αˆαŒ…αŠ¦α‰½ -> ljoc -> lj -> αˆαŒ…` ### Amharic Corpus Indexer Produces an index file for the stemmed words in a corpus and relates them with the files they are found in. It also stores their frequencies per file. ### Term Weighter Calculates the weight of words from the index file using product of their length normalized Term frequency and Inverse document frequency (`tf*idf`). ## Installation Felig Toolkit is available as a pack …