Project Nightingale is an open source project aimed at providing a platform for gathering, documenting and annotating textual data on local African languages.
# Project-Nightingale
Over the past few decades, research in the NLP has been greatly facilitated by the availability and sharing of large annotated and well documented open source textual datasets such as those distributed by the Linguistic Data Consortium (LDC).
The availability of these datasets have led to the development of better translation systems for most global languages. However, this is not the case for African languages. Currently, only 13 languages of the estimated 1500 local African languages in Africa are supported by Google translate. This is mainly due to the lack of reliable, well documented and insufficient annotated datasets on local African languages.
**Project Nightingale is an open source project aimed at providing a platform for gathering, documenting and annotating textual data on local African languages.**
The project will move through three main phases:
+ Phase I: This the data curation phase. In this phase, the project will focus on building data curation tools such as web scraping bots, data entry platforms, ios/andriod apps that will be used to gather and collect textual data in local languages. The data gathered will undergo rigourous data quality and data normlization processes to ensure availability of acurate and rich datasets for the ML/ NLP research community. All the data collected will be open source and released for R & D purposes.
+ Phase II: Phase Two will focus on developing ML/NLP language algorithms and techniques that will leverage NLP and ML Tools
+ The data quality phase (Phase II) involves performing parity and sanity checks on the data and corresponding documentation to ensure conformity to the DST.
+ Finally, all datasets that passed through the quality checks will be ingested to a cloud repository for public usage.
Project webpage is currently under development