Logo Lanfrica

dagisky/Amharic-Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
dag
Host:
# Amharic Corpus Python based code to collect and preprocess Amharic Language Corpus ## Table of Contents 1. Project Motivation 3. Installation 4. Usage 5. Licensing and Acknowledgements ### Project Motivation superstratum to enable communication between people who spoke a mix of different languages. The language serves as the working language of Ethiopia and is also the working language of several of the states within the Ethiopian federal system. However Amharic has a very small publically available corpus for downstream natural language tasks. As such different NLP tasks are considerably difficult for such under-resourced languages. ### Installation The code requires Java and python 3.x installation. In addition use the requirements and install all the required python dependencies. ### Usage ``` python main.py --root_dir Root_Dir_To_Be_Explored --extract_pdf ``` To copy pdf files ``` python main.py --copy_pdf --root_dir Root_Dir_To_Be_Explored --output_dir Directory_to_store ``` ### Licensing and Acknowledgements Credit is due to UESTC, Capital Printing Press # visual geeze formating correction

Languages