This dataset supports Direct Preference Optimization (DPO) fine-tuning using off-policy alignment signals to enhance stylistic control, code-switching behavior, and instruction adherence and on-policy alignment signals to improve instruction-following and safety behavior.