This dataset evaluates the EleutherAI GPT-Neo 1.3B base model by testing 10 diverse prompts in reasoning, translation, arithmetic, factual knowledge, and scientific explanation. Each prompt is evaluated against the expected correct output and blind-spot category.
GPT-Neo 1.3B (Base Pre-trained Model)
Source:
huggingface.co
Methodology