Logo Lanfrica

rumihoney/amazigh-english-corpus-analysis

Domain:

natural language processing

Record type:

dataset
Creator:
rum
Host:
# Amazigh–English Parallel Corpus for NLP Introductory data analysis project using Python and pandas. The goal is to practice data preparation, data analysis, EDA, data cleaning and working with multilingual Unicode text. ## Current Status - Currently cleaning and preprocessing - Removed duplicate translation pairs, unnecessary whitespace, and punctuation - Next: Analyze the translation variations and the ambiguous Unicode characters and capitalized words ## Dataset **Amazigh-English Dictionary: 160k+ Sentences for NLP** Largest Amazigh-English parallel corpus for NLP (from Tatoeba & ManyThings.org) - **Creator:** Jasmine Mohamed Fahmy - **Source:** Amazigh-English 160k Parallel Sentences ## Features - Python - pandas - Jupyter Notebook - KaggleHub ## License Creative Commons Attribution 4.0 International (CC BY 4.0)