Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

jjlalli/Tunisian-Derja-NLP-Resources

Domain:

natural language processing

Record type:

software
Creator:
jjl
Host:
Open, maintained inventory of NLP resources for Tunisian Arabic (aeb). Contributions welcome. # Tunisian Arabic NLP Resources A curated list of datasets, models, tools, and papers for the natural language processing of **Tunisian Arabic** (Tunisian Derja / Tounsi / تونسي, ISO 639-3 code **`aeb`**). Also on Hugging Face: the inventory as a loadable table at datasets/fatmajlali/tunisian-nlp-resources (regenerated from these files by `build-dataset-csv.py`), and the open Tunisian datasets and models gathered in one place in the Tunisian Arabic (Derja) collection. The goal is to be the single most complete inventory of what exists for Tunisian NLP: text and speech, open and gated, so that researchers, students, and engineers can find what is out there and see clearly where the gaps are. **This is meant to stay current, and that only works if it isn't maintained by one person.** If you know a resource that is missing, or you built one, or something here is wrong or out of date: - **Add a resource** — paste a link, that's enough; formatting and verification are my job - **Report something wrong** — dead links and overstated numbers make this worse than useless - Or open a pull request directly — see CONTRIBUTING.md Multi-dialect and pan-Arabic resources are welcome; the map records the Tunisian portion honestly rather than counting the whole thing. ## At a glance | | Entries | Where | |---|---|---| | Text datasets, benchmarks, lexicons & papers | 67 | this file | | Speech corpora (ASR, SLU, translation, TTS) | 25 | SPEECH.md | | Pretrained models (LLMs, encoders, ASR, TTS) | 16 | MODELS.md | | Researchers, labs & companies | 30 | PEOPLE.md | *Counted as one `###` heading each, so the figures above sum to the entries badge and anyone can reproduce them with `grep -c '^### '`. Two caveats in opposite directions: a few headings group several related items (for example "Classic ASR systems (papers)"), which undercounts individual resources; and a handful of resources are cross-listed under a second category with a pointer to the full entry (PADIC, TArC), wh …

Visit

github.com

Languages

Arabic, Tunisian Spoken

Tags

arabicawesome-listdatasetsderjalow-resource-languagesnlptunisian-arabic

Similar

csisc/Tunisian-Derja-NormalizationTunisian Derja Unified Raw CorpusTunisian Arabic NLP Resources: an access-verified inventorychiraz/Definitive-Guide-of-Tunisian-Dialect-NLP-ResourcesBuilding Bi-script Language Resources for the Tunisian Dialect’s NLPjjlalli/Anzar

csisc/Tunisian-Derja-Normalization

Evidence-Based Normalization of Tunisian Derja # Tunisian-Derja-Normalization Evidence-Based Normal

Tunisian Derja Unified Raw Corpus

Tunisian Derja Unified Raw Corpus Dataset Description Repository: hamzabouajila/tunisian-derja-unifi

Tunisian Arabic NLP Resources: an access-verified inventory

An access-verified inventory of natural language processing resources for Tunisian Arabic (Derja, IS

chiraz/Definitive-Guide-of-Tunisian-Dialect-NLP-Resources

Comprehensive list of resources for automated processing of Tunisian dialect text. ## Introduction

Building Bi-script Language Resources for the Tunisian Dialect’s NLP

jjlalli/Anzar

On-device acoustic leak detection for water pipes — ESP32 → STM32, no cloud. Named for the Amazigh r