This dataset provides a targeted, interpretable checklist of reasoning failures and blindspots discovered in the CohereLabs/tiny-aya-base model when evaluated on Amharic language tasks across arithmetic, logic, science, history, and geography domains.