This is a filtered subset of the openai/MMMLU dataset. This dataset only included mathematics (abstr
This is a filtered subset of the CohereLabs/Global-MMLU dataset. This dataset only included mathemat
Global MMLU-Lite is a multilingual evaluation benchmark for LLMs covering 18 languages. This dataset
Authors: Tuka Alhanai tuka@ghamut.com, Adam Kasumovic adam.kasumovic@ghamut.com, Mohammad Ghassemi g
As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are