Logo Lanfrica

Long-Read Sequencing of Mycobacterial Tuberculosis Is Comparable to Short-Read Sequencing for Antimicrobial Resistance Prediction and Epidemiological Studies

Domaine:

healthcare

Type de record:

paper
Créateur:
MatCatEloHan
Éditeur:
Elsevier BV
Hôte:
Background: Short-read genetic sequencing technologies (mainly Illumina) have been extensively used for around a decade for Mycobacterium tuberculosis complex (MTBC) outbreak analysis and genomic drug susceptibility testing (gDST) with the result that Illumina has become the de facto gold standard. Long-read sequencing, as exemplified by Oxford Nanopore Technologies (ONT), offer the prospect of faster, simpler, and portable sequencing. In this work, we carry out the largest to date comparison of how well Illumina and ONT technologies sequence MTBC samples, making use of R10.4 ONT flowcells, updated basecalling models and deep-learning variant calling.

Methods: A total of 508 samples were sequenced using both short and long-read platforms. All samples originated from South Africa or Vietnam and were over-selected for drug resistance and also included several local outbreaks and a range of lineages. The South African and Vietnamese samples had already been Illumina sequenced. Samples with ≥50 read depth by Illumina were selected for sequencing by ONT using one of the GridION or PromethION platforms. Bioinformatics processing was done using a modified online cloud platform which included reference-based variant calling, catalogue-based gDST and identified related samples via SNP counting to inform outbreak detection. The lineages and gDST predictions obtained by short- and long-sequencing were compared for all samples as were all putative clusters identified via SNP counting. For convenience Illumina was used as the reference method.

Findings: Of the 508 samples, 425 (83.7%) had sufficient read depths to permit comparison between the two sequencing technologies. The assigned lineages were identical for 407/425 (95.8%) samples and all discordances were due to mixed lineages being identified by one technology. Evidence of non-tuberculous mycobacterium (NTM) subpopulations were found in nine samples. Using Illumina as the reference method, the very major error (VME) rate of ONT for predicting resistance to all 15 drugs is 1.0% (0.6-1.5%) whilst the major error (ME) rate is 1.7% (1.3-2.2%) with an unclassified rate of 6.9% (6.3-7.5%). This is below the thresholds specified by the CLSI. Considering each of the 15 drugs individually they had VME and ME point estimates below ≤3% in 29/30 cases; and most 25/30 below ≤1.5%.Filtering out all samples containing mixtures left 382 isolates. By appropriate masking of the reference genome we were able to obtain a mean SNP distance between the two platforms of 0.13 (median of zero) for the same sample and for 376/382 samples (98.4%, CI:96.6-99.4%) the difference was ≤1 SNPs. The high concordance in SNP identification ensured that few differences in the 43 putative clusters among 172 isolates were observed. 

Interpretation: The differences between the two sequencing platforms for the key clinical outputs is so small that it is now within the tolerances set by regulatory agencies. Provided the sequencing is of sufficient quality, we have therefore reached a threshold whereby sequencing data from long- and short-read platforms can be aggregated. This will enable large scale analyses by national and international public health agencies whilst allowing the MTBC community to take advantage of the portability and speed of long-read sequencing.

Similaires