Logo Lanfrica

AmEn: Amharic-English Large Parallel Corpus for Machine Translation

Domain:

natural language processing

Record type:

dataset
Creator:
AssAyeBelTon
Publisher:
Und
Host:avatar
Recently, using deep neural networks for machine translation (MT) tasks has received great attention. To learn more abstract representations of the input and store them as continuous vectors, these networks need a large amount of data. However, very few research studies have been conducted, and only a few data points are available for Amharic. The progress of the Amharic to English MT task is affected by a lack of a relatively large and available benchmark dataset. This paper presents the first relatively large-scale Amharic-English parallel corpora (1.1M) for the MT task. We ran experiments using a pre-trained language model (M2M100_48M) and Transformer from scratch. The pre-trained models outperformed the baseline Transformer model.

Languages

Similar