Logo Lanfrica

fgaim/tigrinya-analogy-test

Domain:

natural language processing

Record type:

dataset
Creator:
fga
Host:
Tigrinya Analogy Test for evaluating Word Embeddings and Language Models # Tigrinya Analogy Test Tigrinya Analogy Test for evaluating word embedding models. Download the dataset HERE. ## Introduction This is a Tigrinya version of the Google Analogy Test set, which is used to evaluate English word-embedding models. The analogy test is a well-established strategy to empirically evaluate the quality of word-embedding models. More information about the English task can be found at the ACL Wiki). The data was first machine translated then manually verified by a native speaker to reduce errors. Some aspects of the original analogy test is focused on English and may not transfer well to other languages, such as those related to grammar or morphology. Therefore, when adapting the task we have discarded examples that were discovered as irrelevant in Tigrinya. Finally, there are a total of `18465` entries in the Tigrinya Analogy Test set, while the source English data has `19544` entries. An entry is dropped if the translations led to one of the following conditions: 1. If the source word pair map to one Tigrinya word, for example, both *lucky* and *luckiest* correspond to *ዕድለኛ*. 2. If the source word results in a multi-word expression. For example, *grandson* (*ወዲ ጓል* / *ወዲ ወዲ)*, *granddaughter* (*ጓል ጓል* / *ጓል ወዲ*). This because the typical word-embedding approaches such as word2vec are not designed to predict multi-word phrases. ## Test Sections The test includes a series of semantic and syntactic analogies divided up into subsections including world capitals, currencies, family, tense, and plurality. The test contains the following sections: 1. capital-world 1. currency 1. city-in-state 1. family 1. gram1-adjective-to-adverb 1. gram2-opposite 1. gram3-comparative 1. gram4-superlative 1. gram5-present-participle 1. gram6-nationality-adjective 1. gram7-past-tense 1. gram8-plural 1. gram9-plural-verbs **Examples:** - Semantic section of World Capitals: “ኣስመራ: ኤርትራ as ፓሪስ: ?” and if the model responds correctly it will return: “ፈረንሳ”. …

Languages