Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Parallel Bible Corpus in the TextGrid Repository

Domaine:

natural language processing

Type de record:

dataset
Créateur:
CalChrBarBre
Éditeur:
Zenodo
Hôte:avatar
Benefits of Research Data Repositories: Showcasing the Parallel Bible Corpus and the TextGrid Repository This poster shows the benefits of integrating resources previously available in GitHub repositories into research data repositories. More specifically, we will show the different ways in which the quality of the data of the Multilingual Parallel Bible Corpus could be improved through its integration in the TextGrid Repository (TGR). The corpus contains biblical texts spanning over 100 languages, encoded originally in Corpus Encoding Standard (CES) as XML files (Christodouloupoulos & Steedman 2015). The project aimed to create a corpus aligned at the verse level for comparing methods in highly multilingual contexts, including those involving low-resource languages. In 2025, the resource has been integrated into the TGR, being part of Text+’s portfolio, the consortium for text- and language-based research data within the German National Research Data Infrastructure (NFDI). The TGR aims to integrate existing resources improving the quality of the data in line with the FAIR principles, and providing long-term archiving. As a TEI-specific repository, the TGR has been equipped with new features, such as project-specific options, the use of library classification and authority file systems (Calvo Tello et al. 2023), a Python library for accessing data (Hynek et al. 2024), and a new publication workflow (Veentjer et al. 2025). Other project with XML-TEI data can benefit similarly from these features when publishing in the TGR. During its integration, several aspects of the data were updated following the FAIR-principles (Wilkinson et al. 2016). This proposal highlights only a few. Considering the text, the original CES files have been encoded into TEI. Considering the metadata, some information has been enriched: Integration of metadata about the specific translation (e.g. year). Identification of entities in the metadata: Each biblical work and author (such as Paul or John) has been identified using identifiers from Wikidata and the authority-file system of the German-speaking libraries (GND). Categorization of the works by both testament and genre (e.g. epistles, prophecy, history, gospels) through GND entities. Identification of the languages are through the ISO 639-3 standard and grouped into families through the Basic Classification. Thus, this corpus is one of those that adhere most closely to the FAIR principle.1 The Bible offers a unique chance for analysis across languages. To demonstrate the previously mentioned characteristics of the resource and the TGR, we are currently creating Jupyter Notebooks applying Named Entity Recognition algorithms in various languages.2 The poster will highlight different aspects: First, it will provide a quantitative overview of the corpus, presenting data on various measures such as documents, words, languages, works, and authors. Secondly, a visual representation of the connections across different levels will demonstrate the various types of resources to which the corpus is currently linked. Thirdly, the benefits of the corpus within the TGR will be displayed, with several QR-codes linking to different functionalities. Finally, the poster will give access and show the main results of the above-mentioned Jupyter Notebooks. Furthermore, this contribution serves as an invitation for other projects with TEI files to import their data into the TGR and enjoy similar advantages. References Calvo Tello, José, Stefan E. Funk, Mathias Göbel, Daniel Kurzawe, Nanette Rißler-Pipka, and Ubbo Veentjer. 2023. Between Corpora, Tools, and Authority Files: TextGrid Repository for Hispanic Studies. Revista de Humanidades Digitales, vol. 8: 90–108. doi.org. Christodouloupoulos, Christos, and Mark Steedman. 2015. ‘A Massively Parallel Corpus: The Bible in 100 Languages’. Language Resources and Evaluation 49 (2): 375–95. doi.org. Hynek, Stefan, Ubbo Veentjer, José Calvo Tello, et al. 2024. ‘TextGrid Python Clients: Making the Repository Programmable’. Paper presented at Digital Humanities im deutschsprachigen Raum, Passau. Quo Vadis DH?, February 21. doi.org. Kraft, Sophie, Angela Schmalen, Hendrik Seitz-Moskaliuk, et al. 2021. ‘Nationale Forschungsdateninfrastruktur (NFDI) e. V.: Aufbau und Ziele’. Bausteine Forschungsdatenmanagement, no. 2 (July): 2. doi.org. Veentjer, Ubbo, Stefan Buddenbohm, José Calvo Tello, et al. 2025. ‘Fluffy Publication Workflow: Preserving Humanities Research Data with the TextGrid Repository’. Transformations: A DARIAH Journal Workflows (Textual data-based workflows). doi.org. Wilkinson, Mark D., Michel Dumontier, IJsbrand Jan Aalbersberg, et al. 2016. ‘The FAIR Guiding Principles for Scientific Data Management and Stewardship’. Scientific Data 3 (March). doi.org. 1 textgridrep.org 2 gitlab.gwdg.de

Visit

doi.orgzenodo.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode