Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Artificial real metagenomic reads

Domain:

healthcare

Record type:

dataset
Creator:
Hal
Publisher:
Zenodo
Host:avatar
We created Nanopore and Illumina metagenomic datasets by combining real (non-simulated) sequencing reads into an artificial metagenomic dataset. By doing this, we can be highly confident of the true taxa for each read in the dataset. We used samples which have matched Illumina and Nanopore sequencing to ensure that differences are purely driven by technological differences and not composition differences. For the human component we combined reads from three individuals in the 1000 Genomes Project with Nanopore data for the human component downloaded from the 1000G ONT Sequencing Consortium (millerlaboratory.com) and we provide URLs for each sample: HG00277 Finnish Male with Illumina NovaSeq 6000 (accession: ERR3241786) and Nanopore R10.4; NA19318 Luhya, Kenya Male with Illumina NovaSeq 6000 (accession: ERR3239713) and Nanopore R10.4 (basecalled with Dorado v0.3.4); HG03611 Bengali, Bangladesh Female with Illumina NovaSeq 6000 (accession: ERR3243073) and Nanopore R10.4 (basecalled with Dorado v0.3.4). Each human readset was randomly downsampled to 1Gbp using rasusa (v0.7.1). For the M. tuberculosis component we used Illumina HiSeq 4000 (accession: ERR245682) and Nanopore R10.3 (accession: ERR8170871) (note: we used R10.3 as there are no R10.4 M. tuberculosis WGS datasets publicly available). For the bacterial component, we used Illumina MiSeq (accession: ERR7255689) and Nanopore R10.4 (accession: ERR7287988) reads from the ZymoBIOMICS HMW DNA Standard D6322 (Zymo Research), which contains seven bacterial and one fungal strain(s) - none of which are Mycobacterium. We removed Nanopore reads from all datasets with a length less than 500bp and the M. tuberculosis and Zymo datasets were downsampled to 3Gbp with rasusa. All human, M. tuberculosis, and Zymo reads were combined into a single artificial metagenomic file.

Visit

doi.orgzenodo.org

Languages

Luhya

Tags

bioinformaticsfastqmetagenomicilluminananopore

Licenses

Creative Commons Zero v1.0 Universalhttps://creativecommons.org/publicdomain/zero/1.0/legalcodeMIT Licensehttps://opensource.org/licenses/MIT

Similar

High Quality Project Acheron 16S ReadsDetecting copy number variation with mated short readsAdoption of Artificial Intelligence (AI) in Real Estate Valuation Practice in Lagos, NigeriaArtificial Intelligence Chatbot Adoption Framework for Real-Time Customer Care Support in KenyaArtificial intelligence-enabled screening for diabetic retinopathy: a real-world, multicenter and prospective studyFrangiPANe, a tool for creating a panreference using left behind reads

High Quality Project Acheron 16S Reads

High quality 16S V4 forward read DNA sequences from each Project Acheron sample.

Detecting copy number variation with mated short reads

The development of high-throughput sequencing (HTS) technologies has opened the door to novel method

Adoption of Artificial Intelligence (AI) in Real Estate Valuation Practice in Lagos, Nigeria

The market value of real estate assetsis one of the major determinants influencing real

Artificial Intelligence Chatbot Adoption Framework for Real-Time Customer Care Support in Kenya

In today’s society, most if not all sectors digitize and automate in order to become more e

Artificial intelligence-enabled screening for diabetic retinopathy: a real-world, multicenter and prospective study

Introduction Early screening for diabetic retinopathy (DR) with an efficient and

FrangiPANe, a tool for creating a panreference using left behind reads

International audience We present here FrangiPANe, a pipeline developed to build panr