Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

ยฉ 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Meti2000-ke/ethiopia-tb-genomic-data-curation-ahri-project

Domain:

healthcare

Record type:

software
Creator:
Met
Host:
A reproducible bioinformatics pipeline for quality control, trimming, and curation of Mycobacterium tuberculosis sequencing data from Ethiopian studies (AHRI project). # ๐Ÿงฌ Ethiopia TB Genomic Data Curation Pipeline (AHRI Project) This project presents a reproducible bioinformatics workflow for the curation and preprocessing of *Mycobacterium tuberculosis* sequencing data associated with the Armauer Hansen Research Institute (AHRI) project in Ethiopia. ## ๐Ÿ“Œ Project Overview The goal of this project is to transform raw sequencing data (FASTQ files) into a clean, structured, and analysis-ready dataset through: - Quality control (QC) - Read trimming and filtering - Metadata organization - Data validation - Summary reporting This pipeline ensures reproducibility and provides a strong foundation for downstream genomic analysis such as alignment, variant calling, and lineage identification. ## ๐Ÿงช Data Source - Sequencing data obtained from publicly available studies (SRA) - Samples correspond to *Mycobacterium tuberculosis* isolates from Ethiopia - Curated in the context of an AHRI-related research project - Sample identifiers (SRR IDs) are listed in the metadata file ## โš™๏ธ Pipeline Steps ### 1. Quality Control (Raw Data) - Tool: **FastQC** - Output: `qc_raw/` ### 2. Quality Summary - Tool: **MultiQC** - Aggregates QC reports ### 3. Read Trimming - Tool: **Trimmomatic** - Removes low-quality bases and short reads ### 4. Post-trimming QC - Tool: **FastQC + MultiQC** - Ensures data quality improvement ### 5. Metadata Curation - Structured sample information stored in `metadata.csv` ### 6. Data Validation - Python script verifies: - Missing values - File consistency - Sample integrity ### 7. Summary Statistics - Tool: **SeqKit** - Reports sequence-level statistics ## ๐Ÿš€ How to Run the Pipeline ### 1. Clone the repository git clone github.com

Visit

github.com

Languages

Amharic

Similar

AHRI Data RepositoryDATA CURATION: IMPEDIMENTS TO DIGITAL CURATION AND THEIR PRACTICAL SOLUTIONS IN MALAWIAfro-TB dataset: a large scale genomic data of Mycobacterium tuberculosis in Africainternews-ke/data-workshopGuidelines for shipping and curation of dataAfro-TB dataset as a large scale genomic data of Mycobacterium tuberuclosis in Africa

AHRI Data Repository

The Africa Health Research Institute (AHRI) has published its updated analytical datasets for 2016.

DATA CURATION: IMPEDIMENTS TO DIGITAL CURATION AND THEIR PRACTICAL SOLUTIONS IN MALAWI

Afro-TB dataset: a large scale genomic data of Mycobacterium tuberculosis in Africa

This dataset contains genetic diversity, classification, and resistance data for over 13,753 tubercu

internews-ke/data-workshop

[curriculum] Materials from Internews-Kenya's Data Journalism Workshop (May 26-30, 2014) ##data-wor

Guidelines for shipping and curation of data

Course overview This course provides a basic understanding of data privacy and security in the cont

Afro-TB dataset as a large scale genomic data of Mycobacterium tuberuclosis in Africa

Abstract Mycobacterium tuberculosis (MTB) is a pathogenic bacterium accountable for 10.6 million