---
title: README.md
tags: ["SARS-CoV-2", "Genomics", "Bioinformatics", "Metadata", "Linux", "Analysis", "Activity"]
---
# **Building capacity in SARS-CoV-2 genomics in Africa**
---
###### ***Trainers***: John Juma, Ouso Daniel & Gilbert Kibet
---
- Introduction
- Scope
- Background
- Prerequisite
- Set-Up
- Preparations
- Log into the HPC
- Project Organisation
- Data retrieval and integrity checks
- Analysis
- Loading modules
- Prepare the reference genome
- Quality assessment
- Quality and Adapter filtering
- Decontamination
- Alignment
- Sort and Index alignment map
- Primer trimming
- Compute coverage
- Variant calling
- Variant annotation
- Consensus genome assembly
- Pangolin: Lineage assignment
- Nextclade: Clade assignment
- Summarize results
- Download reports
- Data Retrieval and Review
- Transfer of data: MiSeq
- Transfer of data: MinION
- Reviewing data: Illumina
- Reviewing data: ONT
- Working with metadata
- Galaxy workflows for SARS-CoV-2 data analysis
- Activity
- Galaxy Exercise
## Introduction
In early January 2020, the novel coronavirus (SARS-CoV-2) responsible for a
pneumonia outbreak in Wuhan, China, was identified using next-generation
sequencing (NGS) and readily available bioinformatics pipelines. In addition to
virus discovery, these NGS technologies and bioinformatics resources are
currently being employed for ongoing genomic surveillance of SARS-CoV-2
worldwide, tracking its spread, evolution and patterns of variation on a global
scale.
## Scope
In this short workshop we will tackle, hands-on, the basic principles employed by the numerous bioinformatic pipelines:
to generate consensus genome sequences of SARS-CoV-2 and identify variants using
an actual dataset generated in our facility.
> **Note**
> This is part of the initiative fronted by the Africa
CDC with generous support from the Rockeffeler
foundation to build capacity in pathogen genomics
in Africa.
## Background
We will use a dataset comprising of raw sequence reads of S …