This benchmark is made up of "targeted syntactic evaluation tests for three low-resource languages (Basque, Hindi, and Swahili)" from the paper 'Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models' (Kryvosheieva & Levy, 2025).
Repository:
github.com