This dataset is an annotated collection of Arabic user interface (UI) text strings related to dark patterns in e-commerce mobile applications operating in Saudi Arabia.
Dataset contents:The dataset comprises 223 manually annotated Arabic UI text strings collected from nine e-commerce-related mobile applications operating in Saudi Arabia: Noon, Keeta, SHEIN KSA, Temu, Almosafer, Namshi, HungerStation, Amazon SA, and Booking SA. Text strings were collected across five user flows per application: account registration, product search, add to cart, checkout, and subscription/cancellation management.
Label categories:The dataset includes five dark pattern categories and one non-dark-pattern class:
Urgency / Scarcity: 55 instances
Misdirection: 30 instances
Hidden Costs: 31 instances
Forced Continuity / Roach Motel: 22 instances
Social Proof Manipulation: 31 instances
None — not a dark pattern: 54 instances
Total: 223 instances
Annotation details:Two independent annotators labeled the Arabic UI text strings using the five dark pattern categories and the non-dark-pattern class. Inter-annotator agreement was measured using Cohen’s kappa, yielding κ = 0.89, indicating almost perfect agreement. A total of 20 disagreement cases were resolved through joint adjudication. Two instances originally marked as “Uncertain” during annotation were excluded before the final dataset release, model training, and evaluation.
Intended use:This dataset is intended for research on Arabic dark pattern detection, Arabic natural language processing, consumer protection, ethical interface design, and e-commerce auditing. It may be used to train, evaluate, or benchmark machine learning models for Arabic UI text classification.