This dataset documents failure cases of the base modelQwen3.5-0.8B.
The model has approximately 0.8 billion parameters and is designed as a lightweight multilingual language model.
The goal of this dataset is to highlight cases where the model produces incorrect, misleading, or suboptimal outputs across diverse tasks including translation, reasoning, and regional knowledge.