Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Amharic-English bilingual corpus

Domain:

natural language processing

Record type:

dataset
Publisher:
ELR
Host:avatar
The Amharic-English bilingual corpus contains parallel text from legal and news domains in Amharic script, in transliterated form and in English. The size of the corpus is of 232,653 words in Amharic and 291,701 in English.This parallel corpus contains documents from two domains, namely legal and news, in English and Amharic language. The two domains are separately processed. In addition, for Amharic language, documents were prepared using its own script which is different from Latin alphabet. For easy of use and processing, as well as normalization purposes, the Amharic documents are transliterated and the English documents are converted into lower case format. Furthermore, clean documents were prepared without considering the two domains separately.Amharic is a Semitic language spoken in Ethiopia.

Visit

catalog.elra.info

Tasks

machine translation

Languages

Amharic

Licenses

Rights available for: nonCommercialUse, commercialUse