Microbiota have emerged as fundamental regulators of host physiology,
shaping both ecological interactions and evolutionary trajectories. Yet,
the determinants of microbiota diversity and structure in wild
populations—particularly the respective roles of host genetics and
environmental context—are still poorly understood. In this study, we
investigated these influences in the freshwater snail Bulinus truncatus, a
key intermediate host for human and animal Schistosoma parasites, using a
multifactorial approach. We developed 31 new microsatellite markers to
resolve population genetic structure across nine sites in Senegal.
Metabarcoding methods were then employed to profile the bacterial
microbiota of individual snails and to characterize environmental
bacterial assemblages from each location via environmental DNA. Shell
measurements and molecular diagnostics for trematode infection status were
included to assess additional potential contributors. Employing multiple
regression on distance matrices (MRM), we quantified how snail population
genetics, site-specific environmental bacterial communities, spatial
patterns, and infection status shape microbiota composition. Our analyses
reveal that snail geographic distribution and population genetic structure
drive the composition of Bulinus truncatus microbiota, with environmental
bacterial communities exerting a weaker but still significant effect. In
contrast, neither shell size nor trematode infection status impacted
microbiota structure significantly. Notably, a considerable fraction of
variation remains unexplained, indicating the likely involvement of other
ecological or intrinsic factors. These results advance understanding of
microbiota determinants in natural populations and underscore the
intricate interplay between host genetics, environment, and microbial
communities. This dataset comprises fastq files from 16S v3-v4 rDNA
metabarcoding data of Bulinus truncatus and environmental DNA samples from
Senegal. There is one TAR archive containing the fastq files of eDNA
samples (raw_edna.tar) and another containing the fastq files of B.
truncatus individual snails (raw_bulinus_truncatus.tar). The samples were collected during February 2022 in Senegal. At
each site, we first filtered water along the water column from the surface
to the water-sediment interface. We used filtration capsules of 0.45 µM
mesh size (Waterra USA Inc.) connected to an electric water pump as
previously described in (Douchet
et al. 2022). Filtration was performed
until the filtration capsule was clogged and the final filtration volume
was retrieved. Once the filtration completed, the filtration capsule was
depleted from its water content, filled with 50 mL of Longmire solution,
vigorously shaken, and preserved at ambient temperature until used for
environmental DNA (eDNA) extractions. Following eDNA sampling,
Bulinus truncatus were sampled at each site. Total
eDNA from water filtrations were extracted following (Douchet
et al. 2022) using DNeasy PowerSoil Pro kit
(Qiagen) according to manufacturer’s protocol. Total genomic DNA from each
of the snails removed from their shell was extracted using DNeasy 96 Blood
and Tissue Kit (Qiagen), according to the manufacturer’s
protocol. We targeted the variable V3-V4 loops region
of the
16S sDNA gene using the 341F
(5’-CCTACGGGNGGCWGCAG-3’) and 805R (5’-GACTACHVGGGTATCTAATCC-3’) primers
(Klindworth
et al. 2013) combined with universal
Illumina adapters. The first PCRs were performed using the Q5®
High-Fidelity 2X Master Mix (New England BioLabs), in a 25µL final volume
containing 2 µL of template DNA and using a PCR program consisting in an
initial denaturation step of 30s at 98°C followed by 32 cycles containing
a denaturation step of 6 sec at 98°C, an annealing step of 30 sec at 55°C,
and an elongation step of 8 sec at 72°C and ending with a final elongation
step of 60 sec at 65°C. We used 5 µL of the PCR products to check the
PCR products quality and integrity through electrophoresis on agarose
gels. The 20µL left were sent to the Bio-Environment platform (University
of Perpignan Via Domitia, France). Libraries were performed using Illumina
Nextera index kit and Q5 high fidelity DNA polymerase (New England
Biolabs). Indexed PCR products were then normalized with SequalPrep plates
(ThermoFisher) and paired-end sequenced on a MiSeq instrument using v2
chemistry (2 x 250 bp). # Senegalese *Bulinus truncatus* and environmental samples bacterial
microbiota
[
doi.org](
doi.org) ## Description of the data and file structure This OTUs table and fastq sequences underlie the main results of the study "Evidences that host genetic background more than the environment shapes the microbiota of the snail *Bulinus truncatus*, an intermediate host of *Schistosoma* species.". In this study, our aim was to characterise and compare the bacterial communities found in *B. truncatus* snails and in the aquatic environment at 9 ecologically contrasted. The field work was conducted in February 2022 and we focused on 9 natural sites from Northern Senegal that differ in terms of habitats. At each of these sites, we sampled the water-sediment interface and all *B. truncatus* snails found. We also filtered comercial spring water at each site to have a field negative control. We caracterised the microbiota using v3-v4 16s rDNA metabarcoding a MiSeq amplicons sequencing approach. Those approaches led to the processing of 198 bacterial 16S metabarcoding libraries including the DNA extracts of the 124 snails, 70 eDNA replicates consisting in 54 eDNA samples (each extraction triplicate was amplified in duplicate) and 16 technical field negative controls (each control was amplified in duplicate), and 4 negative PCR controls (milliρ water). ### Files and variables #### File: raw_edna.tar **Description:** fastq sequences coming from the 16s rDNA sequencing of environmental samples (54 samples + 16 field negative controls) #### File: raw_bulinus_truncatus.tar **Description:** fastq sequences coming from the 16s rDNA sequencing of *Bulinus truncatus* snails #### File: controls.tar **Description:** fastq sequences coming from the 16s rDNA sequencing of the PCR negative controls (water) #### Naming convention for FASTQ files The FASTQ files follow an extended Illumina naming structure that encodes sample type, site, replicate information, and sequencing metadata. General structure: `[SampleCode]_[InternalID]_L001_R[ReadNumber]_001.fastq.gz` Where: * **SampleCode** = descriptive identifier indicating sample type and origin * `S1-1` = snail sample 1 from site 1 * `C-S1-1` = field eDNA control 1 from site 1 * `S1-1-1` = eDNA sample 1, subsample 1 from site 1 * `neg-9D-pl3` = PCR negative control (plate 3) * **InternalID** = sequencing facility internal sample ID (e.g., `S1`, `S101`, `S103`, `S194`) * **L001** = sequencing lane * **R1 / R2** = forward and reverse reads * **001** = file index (always 001) ### **Examples** * `S1-1_S1_L001_R1_001.fastq.gz` → snail sample 1 from site 1, forward read * `C-S1-1_S101_L001_R1_001.fastq.gz` → field eDNA control 1 from site 1, forward read * `S1-1-1_S103_L001_R1_001.fastq.gz` → eDNA sample 1, subsample 1 from site 1, forward read * `neg-9D-pl3_S194_L001_R1_001.fastq.gz` → PCR negative control (plate 3), forward read All reverse reads follow the same structure with `R2`. All FASTQ files contained in the `.tar` archives follow this naming convention.
| site | longitude | latitude | habitat | site\_name | | :--- | :--------
| :------- | :------- | :---------- | | S1 | -14.84952 | 16.58457 |
river\_2 | Ndiawara | | S2 | -14.93601 | 16.59875 | river\_2 | Ouali Diala
| | S3 | -14.92521 | 16.59762 | canal | Guia | | S4 | -14.88948 | 16.59714
| river\_2 | Dioundou | | S5 | -14.96063 | 16.60565 | river\_2 | Fonde Ass
| | S6 | -14.94503 | 16.59575 | river\_2 | Khodit | | S7 | -15.80198 |
16.27096 | lake | Mbane | | S8 | -15.80171 | 16.24226 | lake | Saneinte |
| S9 | -16.34956 | 16.10936 | river\_1 | Lampsar | #### File:
Suppl_table_2.csv **Description:** file containing the individual snail
shell size (length and width). **Variables**: * `sample_id` : unique
snail identifier * `shell_length` : shell length (mm) * `shell_width` :
shell width (mm) **Delimiter**: `;` **Decimal notation**: `,` #### File:
suppl_table_4.csv **Description:** asv table containing all samples, prior
to any cleaning step. **Variables**: * `asv` : ASV identifier *
`sample_id` : sample name * `sample_type` : sample type (eDNA control,
negative control, eDNA, snail) * `count` : read count **Delimiter**: `;`
#### File: suppl_table_5.csv **Description:** table showing all ASVs
present in the negative controls, with taxonomy details. **Variables**: *
`asv_id` * `sample_id` * `count` * `taxonomy` **Delimiter**: `;` ####
File: suppl_table_6.csv **Description:** table showing the read counts per
sample during cleaning steps. **Variables**: * `sample_id` * `Read count
after preprocessing` * `Read count after removing mitochondria reads` *
`Read count after removing chloroplast reads` * `Read count after removing
kingdom = NA reads` * `Read count after removing phylum = NA reads` *
`Read count after global abundance filter (0.005%)` * `Read count after
individual abundance filter (0.1%)` * `Read counts after rarefaction to
5000 reads` **Delimiter**: `;` ##### **File:**
Suppl._table_2-_Microsatellite_primers_sequences_W1_W2.xlsx
**Description:** Microsatellite loci used in the study, including locus
name, motif, number of repeats, amplicon size, primer sequences,
sequencing coverage, and genotyping metrics. **Column description** *
**Multiplex** — multiplex panel used for amplification (e.g., W1, W2). *
**Locus name** — unique identifier of the microsatellite locus. *
**ReadID+** — reference sequence ID used for locus mapping. *
**Microsatellite motif** — repeated motif (e.g., ATAG). * **Number of
repeats** — number of motif repetitions. * **Amplicon size** — expected
PCR product length (bp). * **Forward primer sequence*** — forward primer
(5’→3’). * **Reverse primer sequence** — reverse primer (5’→3’). * **Mean
coverage/loci/sample** — average sequencing depth per locus per sample. *
**SD coverage/loci/sample** — standard deviation of sequencing depth. *
**Analysis strategy** — genotyping strategy (e.g., FullLength). *
**Parameter set*** — parameter set used for genotyping. *
**NallelesSequence** — number of alleles detected by sequence. *
**NalleleSize** — number of alleles detected by size. * **Raw missing
genotype $** — proportion of missing genotypes. * **Allelic error** —
allelic error rate.