Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

linuxscout/algerian-name-transcription-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
lin
Host:
Aldjia: Algerian-name-transcription-corpus # Algerian Name Transcription Corpus The **Algerian Name Transcription Corpus (Aldjia علجية)** is an open bilingual corpus of Algerian personal names. It provides aligned Arabic and French spellings of first names and family names collected from real administrative data. The corpus is intended for research and applications involving Arabic–Latin transliteration, named entity recognition, information retrieval, record linkage, and multilingual natural language processing. ## Contents The corpus is distributed as two CSV files. ### First Names `first_names.csv` | Column | Description | | -------------- | ------------------------------------- | | `firstname_fr` | French (Latin) transcription | | `firstname_ar` | Arabic spelling | | `sexe` | `feminin` or `masculin` | | `generation` | `gen1` (parents) or `gen2` (children) | Example: ```text firstname_fr,firstname_ar,sexe,generation Mohamed,محمد,homme,gen2 Fatima,فاطمة,fem,gen1 ``` ### Family Names `family_names.csv` | Column | Description | | --------------- | ---------------------------- | | `familyname_fr` | French (Latin) transcription | | `familyname_ar` | Arabic spelling | Example: ```text familyname_fr,familyname_ar Benali,بن علي Amrani,عمراني ``` ## Corpus Generation The corpus was generated from Students registry data. The generation process includes: 1. Extraction of relevant columns from Excel files. 2. Construction of bilingual name pairs. 3. Removal of duplicate entries. 4. Data normalization. 5. Automatic quality verification. 6. Generation of corpus statistics. ## Validation Several automatic checks are performed to detect potential errors, including: - inconsistent spacing between Arabic and French names; - Latin letters appearing in Arabic fields; - Arabic letters appearing in French fields; - duplicate entries; - additional validation rules can easily be ad …

Visit

github.com

Tasks

named entity recognitioninformation extraction

Languages

Arabic, Algerian Spoken

Licenses

CC0-1.0

Similar

Algerian Corpus (Algerian Dataset)Lwazi II Cross-lingual Proper Name CorpusSouth African Directory Enquiries (SADE) Name CorpusLwazi II Proper Name Call Routing Telephone CorpusA Southern African corpus for multilingual name pronunciationSpoken Tunisian Arabic Corpus “STAC”: Transcription and Annotation

Algerian Corpus (Algerian Dataset)

Lwazi II Cross-lingual Proper Name Corpus

Prompted audio recordings of personal names in different languages, produced by 20 speakers with dif

South African Directory Enquiries (SADE) Name Corpus

"Audio and tagged orthographic transcriptions of South African names produced by first-language spea

Lwazi II Proper Name Call Routing Telephone Corpus

Short prompts of proper names and language names collected via the telephone network.

A Southern African corpus for multilingual name pronunciation

Spoken Tunisian Arabic Corpus “STAC”: Transcription and Annotation