The TyDiP dataset is a dataset of requests in conversations between wikipedia editors
that have been annotated for politeness. The splits available below consists of only
requests from the top 25 percentile (polite) and bottom 25 percentile (impolite) of
politeness scores. The English train set and English test set that are
adapted from the Stanford Politeness Corpus, and test data in 9 more languages
(Hindi, Korean, Spanish, Tamil, French, Vietnamese, Russian, Afrikaans, Hungarian)
was annotated by us.