This dataset is derived from Tulu-v2-SFT-Mixture-English-Darija. It splits the original data by role (user and assistant), creating a separate row for each role with its corresponding English and Darija content.
Dataset: Original source (e.g., flan_v2, sharegpt).
ID: Unique identifier from the original dataset.
Role: user or assistant.
English: English content for the role.
Darija: Translated Darija content for the role.