Whole-genome sequencing (WGS) is increasingly central to antimicrobial resistance (AMR) surveillance and bacterial epidemiology. However, transforming raw sequencing data into reproducible and scalable analyses remains challenging in many resource-constrained research environments due to limited computational infrastructure, fragmented workflows, and inconsistent access to training and high-performance computing resources.
This presentation describes the development of a reproducible bioinformatics pipeline for Enterococcus genomic surveillance in Ghana using Linux- and Python-based research software practices. The pipeline integrates short-read WGS processing, automated metadata organization, comparative genomics, and pangenome analysis using both locally generated and publicly available datasets. It is designed to support scalable characterization of AMR determinants, population structure, and plasmid diversity within a One-Health framework.
A key component of this work is the integration of a machine learning-based approach for plasmid typing from short-read sequencing data, aimed at improving genomic interpretation in settings where long-read sequencing is not routinely available. The workflow is implemented using accessible consumer-grade computing infrastructure, demonstrating that robust and reproducible genomic analyses can be achieved without reliance on high-performance computing systems.
The presentation highlights practical lessons in workflow design, including reproducibility, portability, automation, metadata harmonization, and sustainable software practices. By sharing experiences from Ghana, this work demonstrates how research software can be designed to be both locally feasible and globally reusable, contributing to equitable access to computational genomics tools and strengthening bioinformatics capacity across Africa and other resource-limited settings.