A toolset for Amharic Language pre-processing. Includes an Amharic Stemmer, Transliterator, Stopword remover , Lexical analyzer, Corpus indexer and Term weighter.
Felig Toolkit
A toolset for Amharic Language pre-processing π§
Felig Toolkit Web ββββββββββββββββββββββββββββββββββββ
Now with Typescript support!
---
## What is felig-toolkit?
It is a toolset for Amharic Language pre-processing. It includes an Amharic Stemmer, Amharic Transliterator, Amharic Stopword remover, Amharic Lexical analyzer, Amharic Corpus indexer and Term weighter.
### Amharic Lexical Analyzer
Breaks down Amharic language corpus and returns tokens by removing any whitespace, expanding abbreviations(`α .α -> α α²α΅ α α α `), removing numbers, breaking up hyphenated words, and removing punctuation (`α‘ α’ ! ? `...).
### Amharic Stopword remover
Removes commonly occuring words that have no contribution to the semantics of the corpus. Eg: `α₯α α‘ α΅ααα
α‘ α αααα`...
### Amharic Transliterator
Changes Unicode Amharic characters to ASCII. Exmaple: `ααα½ -> αα
α¦α½ -> ljoc`. This tool implements two types of Amharic transliteration lookup tables.
- SERA (System for Ethiopic Representation in ASCII) - This system maps alphabets with similar sounds separately. Eg: `(αα£αα‘α)α£(α°α‘α )α‘(αΈα‘α)α‘(αα‘α )`. However, in practice, these alphabets are used interchangeably and use of SERA would greatly decrease recall. **NOT RECOMMENDED!**
- Felig - Normalizes the redundant symbols into a common symbol. **RECOMMENDED!**
### Amharic Stemmer LIVE DEMO
Reduces the different morphological (e.g. inflectional or
derivational) variations of Amharic word forms by taking an Amharic word and returning the stem through affix-removal with longest match.
Exmaple:
`ααα½ -> αα
α¦α½ -> ljoc -> lj -> αα
`
### Amharic Corpus Indexer
Produces an index file for the stemmed words in a corpus and relates them with the files they are found in. It also stores their frequencies per file.
### Term Weighter
Calculates the weight of words from the index file using product of their length normalized Term frequency and Inverse document frequency (`tf*idf`).
## Installation
Felig Toolkit is available as a pack β¦