Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilingual AI pipeline for museum knowledge and translation: Terminology curation, semantic structuring, and LLM-enhanced analysis

Domain:

natural language processing

Record type:

paper
Creator:
AldAlbAhmRos
Editor:
SorMohLouObs
Publisher:
CCSD
Host:avatar
This paper presents a multilingual AI pipeline designed to strengthen museum knowledge engineering by combining human-curated terminology, semantic consolidation, and LLM-enhanced document analysis. Developed as part of a 2025 collaboration between the Louvre Abu Dhabi (LAD), the Sorbonne Center for Artificial Intelligence (SCAI), Sorbonne Université, and Sorbonne University Abu Dhabi, the project addresses long-standing gaps in multilingual museum documentation across English, French, and Arabic. Challenges include terminology heterogeneity, limited availability of specialized Arabic resources, and inconsistencies in cross-lingual metadata-all of which constrain translation, cataloguing, and digitization workflows.To address these issues, we introduce a three-module pipeline:1. Multilingual Terminological Resource: a validated corpus of 400 entries and 68 bibliographic records, enriched with definitions, authoritative sources, and images.2 . Semantic Structuring and Termbase Integration: the consolidation of nearly 1,000 terms into the Louvre Abu Dhabi Termbase, following LAD's concept-based, metadata-rich model and hybrid prescriptive/descriptive methodology.3 . LLM-Enhanced OCR and Metadata Extraction: a document analysis workflow combining OCR engines with multimodal LLMs for transcription, post-correction, and structured artifact metadata extraction across all three languages.Experiments demonstrate that curated terminology and semantic relations significantly improve the accuracy of OCR post-correction and LLM extraction-especially for Arabic, where language-specific thresholds and manual gold-standard pages (294 corrected pages) were essential. The resulting workflow provides a scalable methodological framework for multilingual museum documentation and a transferable blueprint for heritage institutions.The resulting workflow provides a scalable methodological framework for multilingual museum documentation and demonstrates how concept-based terminological resources can function as effective semantic constraints for LLM-driven document analysis in low-resource multilingual settings.

Visit

hal.science

Tasks

computer visionmachine translationoptical character recognition

Tags

Museum terminologyMultilingual information extractionOCRNamed entitiesLarge language models[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

Framing Food Online Discourse: Employing Generative AI and Semantic Analysis for Digital Lexicography and Terminology Extraction Journal of Digital Terminology and LexicographyAdapters for Enhanced Modeling of Multilingual Knowledge and TextReproducibility Package for "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis"Semantic analysis of witchcraft terminology in Northern SothoCollectivIA: Two-Pipeline Multilingual Legal RAG for Moroccan Territorial Governance with LLM-Assisted and Regex-Based Chunkingbandym05/AI-Enhanced-Health-Management-Combining-Predictive-Analytics-LLM-and-RAG-Driven-Support-for-Diabe

Framing Food Online Discourse: Employing Generative AI and Semantic Analysis for Digital Lexicography and Terminology Extraction Journal of Digital Terminology and Lexicography

Social media platforms provide vast amounts of authentic, user-generated linguistic data that can

Adapters for Enhanced Modeling of Multilingual Knowledge and Text

Large language models appear to learn facts from the large text corpora they are trained on. Such fa

Reproducibility Package for "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis"

This deposit contains the reproducibility materials for the manuscript "A Governance Archit

Semantic analysis of witchcraft terminology in Northern Sotho

CollectivIA: Two-Pipeline Multilingual Legal RAG for Moroccan Territorial Governance with LLM-Assisted and Regex-Based Chunking

Retrieval grounding is crucial for high-stakes administrative applications, since large language mod

bandym05/AI-Enhanced-Health-Management-Combining-Predictive-Analytics-LLM-and-RAG-Driven-Support-for-Diabe

An intelligent system that combines predictive analytics, LLMs, and RAG to assess diabetes risk, off