Societal Impact Statement
Biological specimens housed in natural history collections are indispensable resources for documenting where species occur and how they have changed through time, and are thus vital for combating biodiversity loss. Digitization of these collections promises to make these critical resources globally available. However, manually transcribing specimen labels is a time‐intensive process, which considerably delays their being available online. Here, we report a case study of using an automated pipeline that harnesses the large language model, ChatGPT, to transcribe specimens in the Missouri Botanical Garden Herbarium, and encourage other institutions to consider adopting similar approaches for digitizing their collections.
Summary
The online mobilization of natural history collections is critical for expanding access to specimen data and combating the ongoing loss of biodiversity. However, specimen digitization is often time and labor intensive, necessitating the development of high‐throughput digitization workflows. Here, we report a case study detailing the development and use of a novel pipeline for automating transcription of specimen labels at the Missouri Botanical Garden Herbarium (MO) and offer lessons learned for institutions embarking on similar digitization efforts. The pipeline harnesses optical character recognition (OCR), large language model (LLM) guided parsing, and post‐process data cleaning on a batch of specimen images, and returns a spreadsheet formatted for upload to an institutional database. This pipeline can optionally be set to recognize the language of OCR‐derived text and translate it into English. We implemented this workflow for two digitization projects, one for Asia and one for tropical Africa. The pipeline successfully transcribed all tested fields in a majority of cases, with many common specimen‐related variables achieving accuracies >76%. Implementing this pipeline reduced the time digitization staff invested in transcriptions by 12.0% and 28.8%, and decreased transcription cost by 10.2% and 28.5%, for the Asia and Africa projects, respectively. This case study provides valuable lessons for implementing automated transcription pipelines in large‐scale digitization projects and demonstrates the value of harnessing LLM‐based pipelines for digitization at scale.