Introduction
This dataset contains 263 metagenome-assembled genomes (MAGs) recovered from the Nested Agroecology Meta-Omic Studies of Crop–Shrub–Microbiome Interactions in a Sahelian Intercropping System project. These MAGs represent some of the first metagenome-assembled genomes recovered from agricultural soils in the West African Sahel and were instrumental in identifying microbial taxa and functions associated with crop–shrub–microbiome interactions in this system.
Methods
Metagenomic libraries were prepared and sequenced at the Department of Energy Joint Genome Institute (JGI). Metagenomic libraries from the simulated drought experiment were prepared at The Ohio State University using the Illumina Nextera XT DNA Library Prep Kit (Illumina, San Diego, CA, USA) according to the manufacturer's instructions with minor modifications. Simulated drought experiment metagenomes were sequenced at the Columbia Genomics Core on the Illumina NovaSeq S4 platform.
All metagenomic samples were assembled using MEGAHIT (v1.2.9) with default settings (1). Unmapped reads from the field study assemblies were indexed using Bowtie2 (v2.5.2) (2) and reassembled with MEGAHIT (v1.2.9). These assemblies were then combined with the original assemblies and deduplicated using DeDupe (3). Because this approach resulted in only minor improvements, it was not repeated for the simulated drought experiment assemblies. Assembly quality was assessed using MetaQUAST (4), and eukaryotic contamination was quantified using EukRep (v0.6.6) (5). Virtually no eukaryotic signal was detected in simulated drought metagenomes (714 of 12,366,520 scaffolds) or metatranscriptomes (518 of 42,813 scaffolds). In the field study, 42,813 scaffolds out of 297,435,843 total scaffolds were identified as eukaryotic.
Reads were mapped to assemblies using CoverM (v0.6.1) (6) with the parameter --min-covered-fraction 10, and the trimmed mean method was used to further assess assembly quality. Functional annotation of all open reading frames (ORFs) was performed using DRAM (7), and all proteins from both studies were clustered using the Markov Cluster Algorithm (MCL) (8), resulting in 1,583,284 protein clusters (PCs). Read coverage of PCs was quantified using CoverM (6) and used to assess differential protein-cluster enrichment using LEfSe (9), with treatment designated as the class variable and replicate as the subclass variable.
Recovery of Metagenome-Assembled Genomes
Binning and refinement of field-study metagenomic assemblies were conducted using two complementary approaches: (i) the JGI standard metagenome analysis pipeline, which utilized metaSPAdes (v3.13.0) (10) and MetaBAT (11) with a 3,000-bp minimum contig cutoff and the parameter -superspecific to maximize specificity, and (ii) an in-house workflow implemented through MetaWRAP (12), which used MaxBin2 (v2.12.1) (13) and MetaBAT2 (11) with a minimum contig length of 500 bp. Bin quality was evaluated using CheckM (v1.1.6) (14), and bins meeting MIMAG medium-quality standards (>70% completeness and <10% contamination) were retained as MAGs (15).
A total of 1,180 MAGs were recovered, including 819 medium-quality MAGs (>70% complete, <10% contaminated) and 361 high-quality MAGs (>90% complete, <5% contaminated). Of these, 989 were recovered using in-house workflows, 166 were recovered through the JGI pipeline, and 25 were recovered from the simulated drought experiment. These 1,180 MAGs were dereplicated at 95% average nucleotide identity (ANI) using dRep (16), resulting in a final set of 263 nonredundant MAGs. Taxonomic assignments were generated using GTDB-Tk v2.3.0 (17). All MAGs recovered from the simulated drought experiment were distinct from those recovered from the field study. Functional annotation of MAGs was performed using DRAM v1.4 (7).
MAG abundance was quantified as reads per million trimmed metagenomic reads using CoverM (v0.6.1) with the parameters --min-read-aligned-percent 75 --min-read-percent-identity 95 (6). The proportion of the microbial community represented by the recovered MAG set was estimated using SingleM appraise (18). Genus-level recovery estimates were generated using the parameters --imperfect --sequence_identity 0.89, while approximate species-level recovery estimates were generated using --imperfect --sequence_identity 0.95.
Data Availability
Raw metagenomic sequencing reads are available through NCBI BioProject accessions:
PRJNA928765
PRJNA930014
Genome annotations are here: 263 MAG annotations for three nested metagenomic studies describe crop-shrub-microbe interactions in an agroecology system in the Sahel
References
Li D, Liu CM, Luo R, Sadakane K, Lam TW. 2015. MEGAHIT: An ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics 31:1674–1676.
Langmead B, Salzberg SL. 2012. Fast gapped-read alignment with Bowtie 2. Nature Methods 9:357–359.
Bushnell B. BBMap/DeDupe. Lawrence Berkeley National Laboratory. Available at:
sourceforge.net
Mikheenko A, Saveliev V, Gurevich A. 2016. MetaQUAST: Evaluation of metagenome assemblies. Bioinformatics 32:1088–1090.
West PT, Probst AJ, Grigoriev IV, Thomas BC, Banfield JF. 2018. Genome reconstruction for eukaryotes from complex natural microbial communities. Genome Research 28:569–580.
Woodcroft BJ. CoverM: Read coverage calculator for metagenomics. Available at:
github.com
Shaffer M, Borton MA, McGivern BB, et al. 2020. DRAM for distilling microbial metabolism to automate the curation of microbiome function. Nucleic Acids Research 48:8883–8900.
van Dongen S. 2008. Graph clustering via a discrete uncoupling process. SIAM Journal on Matrix Analysis and Applications 30:121–141.
Segata N, Izard J, Waldron L, et al. 2011. Metagenomic biomarker discovery and explanation. Genome Biology 12:R60.
Nurk S, Meleshko D, Korobeynikov A, Pevzner PA. 2017. metaSPAdes: A new versatile metagenomic assembler. Genome Research 27:824–834.
Kang DD, Li F, Kirton E, et al. 2019. MetaBAT 2: An adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies. PeerJ 7:e7359.
Uritskiy GV, DiRuggiero J, Taylor J. 2018. MetaWRAP—A flexible pipeline for genome-resolved metagenomic data analysis. Microbiome 6:158.
Wu YW, Simmons BA, Singer SW. 2016. MaxBin 2.0: An automated binning algorithm to recover genomes from multiple metagenomic datasets. Bioinformatics 32:605–607.
Parks DH, Imelfort M, Skennerton CT, Hugenholtz P, Tyson GW. 2015. Assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Research 25:1043–1055.
Bowers RM, Kyrpides NC, Stepanauskas R, et al. 2017. Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea. Nature Biotechnology 35:725–731.
Olm MR, Brown CT, Brooks B, Banfield JF. 2017. dRep: A tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication. ISME Journal 11:2864–2868.
Chaumeil PA, Mussig AJ, Hugenholtz P, Parks DH. 2022. GTDB-Tk v2: Memory-friendly classification with the Genome Taxonomy Database. Bioinformatics 38:5315–5316.
Woodcroft BJ. SingleM: A tool for estimating microbial community composition from metagenomic data. Available at:
wwood.github.io