This is a multilingual reasoning dataset covering more than 30 languages.
This dataset was made by:
Sampling prompts from English datasets and translating them to various languages
Generating responses to these prompts 8 times using deepseek-ai/DeepSeek-R1-Distill-Llama-70B
Filtering out sections with incorrect language, non-fluent language, and incorrect answers