Logo Lanfrica

jjlalli/Tunisian-Derja-NLP-Resources

Domaine:

natural language processing

Type de record:

software
Créateur:
jjl
Hôte:
Open, maintained inventory of NLP resources for Tunisian Arabic (aeb). Contributions welcome. # Tunisian Arabic NLP Resources A curated list of datasets, models, tools, and papers for the natural language processing of **Tunisian Arabic** (Tunisian Derja / Tounsi / تونسي, ISO 639-3 code **`aeb`**). Also on Hugging Face: the inventory as a loadable table at datasets/fatmajlali/tunisian-nlp-resources (regenerated from these files by `build-dataset-csv.py`), and the open Tunisian datasets and models gathered in one place in the Tunisian Arabic (Derja) collection. The goal is to be the single most complete inventory of what exists for Tunisian NLP: text and speech, open and gated, so that researchers, students, and engineers can find what is out there and see clearly where the gaps are. **This is meant to stay current, and that only works if it isn't maintained by one person.** If you know a resource that is missing, or you built one, or something here is wrong or out of date: - **Add a resource** — paste a link, that's enough; formatting and verification are my job - **Report something wrong** — dead links and overstated numbers make this worse than useless - Or open a pull request directly — see CONTRIBUTING.md Multi-dialect and pan-Arabic resources are welcome; the map records the Tunisian portion honestly rather than counting the whole thing. ## At a glance | | Entries | Where | |---|---|---| | Text datasets, benchmarks, lexicons & papers | 67 | this file | | Speech corpora (ASR, SLU, translation, TTS) | 25 | SPEECH.md | | Pretrained models (LLMs, encoders, ASR, TTS) | 16 | MODELS.md | | Researchers, labs & companies | 30 | PEOPLE.md | *Counted as one `###` heading each, so the figures above sum to the entries badge and anyone can reproduce them with `grep -c '^### '`. Two caveats in opposite directions: a few headings group several related items (for example "Classic ASR systems (papers)"), which undercounts individual resources; and a handful of resources are cross-listed under a second category with a pointer to the full entry (PADIC, TArC), wh …