Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Using large language models to automate herbarium specimen transcription: A case study at the Missouri Botanical Garden

Domaine:

environment and energy

Type de record:

software
Créateur:
MatAleMikHei
Éditeur:
WILEY
Hôte:
Societal Impact Statement Biological specimens housed in natural history collections are indispensable resources for documenting where species occur and how they have changed through time, and are thus vital for combating biodiversity loss. Digitization of these collections promises to make these critical resources globally available. However, manually transcribing specimen labels is a time‐intensive process, which considerably delays their being available online. Here, we report a case study of using an automated pipeline that harnesses the large language model, ChatGPT, to transcribe specimens in the Missouri Botanical Garden Herbarium, and encourage other institutions to consider adopting similar approaches for digitizing their collections. Summary The online mobilization of natural history collections is critical for expanding access to specimen data and combating the ongoing loss of biodiversity. However, specimen digitization is often time and labor intensive, necessitating the development of high‐throughput digitization workflows. Here, we report a case study detailing the development and use of a novel pipeline for automating transcription of specimen labels at the Missouri Botanical Garden Herbarium (MO) and offer lessons learned for institutions embarking on similar digitization efforts. The pipeline harnesses optical character recognition (OCR), large language model (LLM) guided parsing, and post‐process data cleaning on a batch of specimen images, and returns a spreadsheet formatted for upload to an institutional database. This pipeline can optionally be set to recognize the language of OCR‐derived text and translate it into English. We implemented this workflow for two digitization projects, one for Asia and one for tropical Africa. The pipeline successfully transcribed all tested fields in a majority of cases, with many common specimen‐related variables achieving accuracies >76%. Implementing this pipeline reduced the time digitization staff invested in transcriptions by 12.0% and 28.8%, and decreased transcription cost by 10.2% and 28.5%, for the Asia and Africa projects, respectively. This case study provides valuable lessons for implementing automated transcription pipelines in large‐scale digitization projects and demonstrates the value of harnessing LLM‐based pipelines for digitization at scale.

Visit

doi.org

Tasks

optical character recognitioncomputer vision

Licenses

http://creativecommons.org/licenses/by/4.0/http://creativecommons.org/licenses/by/4.0/http://doi.wiley.com/10.1002/tdm_license_1.1

Similaires

DELSUH Herbarium Specimen Catalogue and Analysis CodeAutomate to Innovate or Automate to Stagnate? A Comparative Case Study of AI Adoption Paradoxes in Nigerian Fintech SMEsLeveraging Large Language Models to Preserve Indigenous Games: A Case Study of a Chatbot for the Kenyan Game BanoLow-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-StudyLarge Language Model-assisted EIA screening: a case study using GPTUsing large language models and speech-to-text models to facilitate the assessment of basic literacy in Ghana

DELSUH Herbarium Specimen Catalogue and Analysis Code

Cleaned and standardised specimen catalogue and analysis code supporting the study “T

Automate to Innovate or Automate to Stagnate? A Comparative Case Study of AI Adoption Paradoxes in Nigerian Fintech SMEs

This paper was presented at the ACM International Conference on AI in Finance 2025 in Singa

Leveraging Large Language Models to Preserve Indigenous Games: A Case Study of a Chatbot for the Kenyan Game Bano

This paper looks at how large language models can be used to help preserve indigenous cultural knowl

Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study

Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain

Large Language Model-assisted EIA screening: a case study using GPT

Large Language Models (LLMs) have developed rapidly in recent years and are increasingly used for

Using large language models and speech-to-text models to facilitate the assessment of basic literacy in Ghana

This dissertation examines how recent advances in artificial intelligence, particularly in Natural L