This dataset is a filtered subset of [juletxara/mgsm] with an added integer id per language.
English questions overlapping with a PolyMath English set were removed, and the same ids were excluded from all languages to avoid cross-language overlaps.
Source: juletxara/mgsm (test split)
Filtering date: 2025-09-15
Splits: each language is exposed as a split (en, bn, de, es, ja, sw, te, th).
Fields