Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ismailSadouki/dziri-dpo

Domain:

natural language processing

Record type:

softwareproject
Creator:
ism
Host:
End-to-end DPO alignment pipeline from scratch. Implements DPO mechanics in pure PyTorch with numerical verification against TRL, followed by SFT and DPO training, preference data construction, annotation agreement, and evaluation for Algerian Darija alignment. # DziriDPO **End-to-end DPO alignment pipeline from scratch.** Implements DPO mechanics in pure PyTorch with numerical verification against TRL, followed by SFT and DPO training, preference data construction, annotation agreement, and evaluation for Algerian Darija alignment. > **Public framing:** verified DPO mechanics + Darija alignment methodology — not simply “I fine-tuned a model.” --- ## Overview DziriDPO is a two-stage project for understanding and building Direct Preference Optimization (DPO) systems. The project separates **DPO mechanics verification** from **production Darija alignment**: ```text DziriDPO │ ┌────────────┴────────────┐ │ │ ▼ ▼ Scratch DPO Darija Alignment Pure PyTorch TRL │ │ ▼ ▼ DPO mechanics Data + SFT + DPO │ │ ▼ ▼ TRL numerical oracle Triple-axis evaluation │ │ └────────────┬────────────┘ ▼ Reproducible results ``` The central correctness oracle is: $$ |L_{\text{scratch}} - L_{\text{TRL}}| **What makes the chosen response better than the rejected response in Algerian Darija?** The guideline considers properties such as: * instruction following * factual correctness * relevance * clarity * natural Darija usage * appropriate code-switching * cultural/contextual appropriateness * harmful or misleading content The detailed guideline is maintained in: ```text darija_alignment/data/guideline.md ``` --- # Inter-Annotator Agreement A subset of at least **100 preference pairs** is independently double-annotated. Cohen's kappa is computed as: $$ \kappa ====== \frac{p_o-p_e}{1-p_e} $$ where: * $p_o$ is observed agreement * $p_e$ is expected agreement by chance The report includes: * number of double-annotated pairs * observed agreement * expected agreement * Cohen's $\kappa$ * disagreement categories * re …

Visit

github.com

Tasks

language modeling

Languages

Arabic, Algerian Spoken

Similar

ismailSadouki/mini-infirence-engineikraam05/dziri-sentiment-analysisDziriOFN Corpus (Dziri Offensive corpus) v1.0afrisynt/amharic_llama3-dpo-rgeminireuben256/alpaca-DPO-lugandaNeurotech-HQ/python-dpo

ismailSadouki/mini-infirence-engine

End-to-end LLM inference engine from scratch. Implements KV caching, prefill/decode execution, conti

ikraam05/dziri-sentiment-analysis

Sentiment Analysis on Algerian Dialect using DziriBERT # Algerian Dialect Sentiment Analysis using

DziriOFN Corpus (Dziri Offensive corpus) v1.0

Dziri refers to the name of the Algerian dialectal Arabic. DziriOFN is a new corpus dedicated to offensive language detection on this under-resourced language. Dziri dialect if known as a complex socio-linguistic situation, where the latter is known by the code-

afrisynt/amharic_llama3-dpo-rgemini

reuben256/alpaca-DPO-luganda

This dataset is a Luganda version of the Alpaca dataset, formatted for Direct Preference Optimizatio

Neurotech-HQ/python-dpo

A python package to easy the integration with Direct Online Pay (Mpesa, TigoPesa, AirtelMoney, Card