# Amazigh–English Parallel Corpus for NLP
Introductory data analysis project using Python and pandas. The goal is to practice data preparation, data analysis, EDA, data cleaning and working with multilingual Unicode text.
## Current Status
- Currently cleaning and preprocessing
- Removed duplicate translation pairs, unnecessary whitespace, and punctuation
- Next: Analyze the translation variations and the ambiguous Unicode characters and capitalized words
## Dataset
**Amazigh-English Dictionary: 160k+ Sentences for NLP**
Largest Amazigh-English parallel corpus for NLP (from Tatoeba & ManyThings.org)
- **Creator:** Jasmine Mohamed Fahmy
- **Source:** Amazigh-English 160k Parallel Sentences
## Features
- Python
- pandas
- Jupyter Notebook
- KaggleHub
## License
Creative Commons Attribution 4.0 International (CC BY 4.0)