This dataset captures blind spots of the DeepBrainz-R1-0.6B model: Model Link.
I generated 10 diverse prompts to test the model’s responses against expected answers. The prompts cover factual knowledge, math, language, reasoning, and commonsense questions. This helps highlight where the model may give incorrect or misleading outputs.
"What is the capital of Burkina Faso?"
"If 10y + 5 = 30, find y."