Logo Lanfrica

TyDiP

Domain:

natural language processing

Record type:

dataset
Creator:
Gen
Host:
The TyDiP dataset is a dataset of requests in conversations between wikipedia editors that have been annotated for politeness. The splits available below consists of only requests from the top 25 percentile (polite) and bottom 25 percentile (impolite) of politeness scores. The English train set and English test set that are adapted from the Stanford Politeness Corpus, and test data in 9 more languages (Hindi, Korean, Spanish, Tamil, French, Vietnamese, Russian, Afrikaans, Hungarian) was annotated by us.