L’ajami, c’est-à-dire l’usage de l’alphabet arabe pour transcrire des langues africaines, a produit sur plusieurs siècles un patrimoine documentaire considérable, longtemps tenu à l’écart des dispositifs classiques de traitement de l’écrit. Ce corpus pose, à l’ingénierie documentaire, des problèmes que les chaînes standardisées peinent à résoudre : variabilité graphique forte, plurilinguisme intra-document, conventions orthographiques locales, ancrage rituel ou pédagogique des textes. Cet article propose une chaîne de traitement intégrée pour ces écritures hybrides, fondée sur l’ingénierie documentaire et un usage circonstancié des technologies récentes, en particulier l’intelligence artificielle, les architectures de type Retrieval-Augmented Generation (RAG) et les mécanismes de traçabilité issus de la blockchain documentaire.
Le matériau de réflexion réunit des manuscrits numérisés, des fragments iconographiques diffusés sur le web et des archives privées. Sur cette base, l’article examine d’abord les limites des outils standards d’OCR et de HTR appliqués à ces écritures, puis articule reconnaissance graphique, translittération assistée, traduction contextualisée et enrichissement sémantique multilingue. Une attention particulière est portée à la modélisation des métadonnées : il s’agit d’associer les normes bibliothéconomiques et archivistiques (Dublin Core, TEI, ISBD, RiC-O) aux exigences culturelles propres aux corpus afro-arabes. La recherche augmentée par le contexte et les registres distribués sont ici envisagés non comme des solutions, mais comme des leviers : leviers de gouvernance des versions, des droits et des interprétations. Au total, l’ambition est de fonder une ingénierie documentaire cognitive de ces traditions scripturales, attentive à la fois à la rigueur technique, au respect des valeurs culturelles et à la participation des communautés détentrices.
Ajami, the practice of writing African languages in Arabic script, has produced over several centuries a substantial body of written heritage that has long remained outside the reach of conventional text-processing systems. This corpus confronts documentary engineering with problems that standardised workflows struggle to resolve: pronounced graphic variability, multilingualism within single documents, local orthographic conventions, and the ritual or pedagogical settings in which the texts are embedded. This article proposes an integrated processing chain for these hybrid scripts, grounded in documentary engineering and in a measured use of recent technologies, notably artificial intelligence, Retrieval-Augmented Generation (RAG) architectures, and the traceability mechanisms derived from documentary blockchain. The material under study brings together digitised manuscripts, iconographic fragments circulating on the web, and private archives. On this basis, the article first examines the limits of standard OCR and HTR tools when applied to these scripts, then articulates graphic recognition, assisted transliteration, contextualised translation, and multilingual semantic enrichment. Particular attention is given to metadata modelling, where the aim is to combine library and archival standards (Dublin Core, TEI, ISBD, RiC-O) with the cultural requirements specific to Afro-Arabic corpora. Context-augmented retrieval and distributed ledgers are treated here not as solutions but as instruments of governance, for versions, for rights, and for interpretations. The overall ambition is to lay the foundations of a cognitive documentary engineering for these scribal traditions, one that holds together technical rigour, respect for cultural values, and the participation of the communities that hold these materials."