Logo Lanfrica

yordanoswuletaw/amharic-ngram-autocomplete

Domain:

natural language processing

Record type:

softwaremodel
Creator:
yor
Host:
Amharic N-Gram Language Model for Auto-Completion, Implemented in Python and NumPy from Scratch # Amharic N-Gram Language Model for Auto-Complete This project implements an N-Gram language model for the Amharic language using only Python and NumPy, designed to provide auto-completion functionality. The model employs a simple N-Gram probabilistic approach to predict and suggest the most probable next words based on input sequences. ## Features - **N-Gram Based Predictions**: Supports unigram, bigram, trigram and n-gram models to generate context-aware suggestions. - **Amharic Language Support**: Handles the structure and highly morphological nature of Amharic text. - **Tokenization**: Includes an Amharic-specific tokenizer to handle words and punctuation correctly. - **Smoothing Techniques**: Implements smoothing methods (e.g., Laplace smoothing) to address the issue of zero probabilities. - **Scalable Design**: Can be trained on large datasets for improved accuracy. ## Installation 1. Clone this repository: ```bash git clone github.com ``` 2. Navigate to the project directory: ```bash cd amharic-ngram-autocomplete ``` 3. Install the required dependencies: ```bash pip install -r requirements.txt ``` ## Notebooks 1. **Amharic Auto Complete** ## Repository Structure ```plaintext ├── .vscode/ │ └── settings.json # VS Code settings for environment setup ├── .github/ │ └── workflows/ │ ├── unittests.yml # CI/CD pipeline for unit tests ├── .gitignore # Ignored files and folders ├── requirements.txt # Dependencies for the project ├── README.md # Documentation of the repository ├── data/ # Dataset for training, dev and testing ├── src/ # Source code for analysis and processing ├── notebooks/ │ ├── __init__.py # Package initialization │ └── README.md # Documentation for the notebooks ├── tests/ │ ├── __init__.py # Test initialization └── s …