Skip to content

Pipeline overview

What You Get

Category What It Finds Steps
Variant Calling SNPs, indels, structural variants, copy number variants 3, 4, 4b, 18, 19
Clinical Screening Pathogenic variants, carrier status, cancer predisposition (CPSR panels), SMN1/SMN2 copy number (opt-in) 6, 17, 35
Pharmacogenomics Drug-gene interactions (23+ genes, CYP2C19, CYP2D6 SV, DPYD, etc.), with HLA-A/B from T1K and CYP2D6 only when two callers agree 7, 21, 27, 32, 36
Structural Variants Deletions, duplications, inversions, translocations (4 callers + consensus) 4, 4b, 5, 15, 18, 19, 22
Functional Annotation Impact prediction for every variant (VEP + CADD, SpliceAI, REVEL, AlphaMissense) 13, 30
Variant Prioritization Rare deleterious variants, compound hets, gene constraint filtering 31
Repeat Expansions Huntington's, Fragile X, ALS and the other disorders at the 31 loci of ExpansionHunter's bundled catalog 9, 9b
Ancestry & Haplogroups Mitochondrial haplogroup (with an mtDNA contamination check), Y-chromosome haplogroup (opt-in), consanguinity check, projection onto a 1000 Genomes reference panel 11, 12, 26, 37
Telomere Length Relative telomere content estimation from WGS reads 10
Mitochondrial Heteroplasmy detection, mitochondrial disease variants 12, 20
Polygenic Risk Scores for 9 common conditions (CAD, T2D, cancers, etc.) with pgsc_calc; a percentile among the most similar ancestry group with the panel installed 25
Quality Control Adapter trimming, coverage statistics, aggregated QC report, sex check, sample identity and contamination, SV filtering 1b, 15, 16, 16b, 28, 33
Storage Alignments kept as a checked CRAM, about half the size of the BAM 34

Pipeline Overview

graph LR
    FASTQ["FASTQ"] --> fastp["fastp<br/><small>QC + trim</small>"]
    fastp --> align["minimap2<br/><small>Alignment</small>"]
    align --> BAM["Sorted BAM"]
    BAM --> DV["DeepVariant<br/><small>SNPs + indels</small>"]
    DV --> VCF["VCF"]

    %% VCF-based analyses
    VCF --> clinvar["ClinVar Screen"]
    VCF --> pharmcat["PharmCAT<br/><small>PGx</small>"]
    pharmcat --> cpic["CPIC<br/><small>Drug recs</small>"]
    VCF --> vep["VEP<br/><small>Annotation</small>"]
    vep --> vcfanno["vcfanno<br/><small>CADD / SpliceAI<br/>REVEL / AlphaMissense</small>"]
    vcfanno --> slivar["slivar<br/><small>Prioritization</small>"]
    vcfanno --> clinical["Clinical Filter"]
    VCF --> cpsr["CPSR<br/><small>Cancer predisposition</small>"]
    VCF --> roh["ROH Analysis"]
    VCF --> prs["PRS<br/><small>Polygenic risk</small>"]
    VCF --> ancestry["Ancestry<br/><small>panel projection</small>"]

    %% BAM-based analyses
    BAM --> manta["Manta<br/><small>SVs</small>"]
    BAM --> delly["Delly<br/><small>SVs</small>"]
    BAM --> cnvpytor["CNVpytor<br/><small>CNVs</small>"]
    manta --> duphold["duphold"]
    duphold --> annotsv["AnnotSV"]
    manta --> consensus["SV Consensus"]
    delly --> consensus
    cnvpytor --> consensus

    BAM --> eh["ExpansionHunter<br/><small>STRs</small>"]
    BAM --> pypgx["pypgx<br/><small>23-gene PGx<br/>+ CYP2D6 SV</small>"]
    BAM --> cyrius["Cyrius<br/><small>CYP2D6, opt-in</small>"]
    BAM --> hla["T1K<br/><small>HLA (+ KIR, opt-in)</small>"]
    BAM --> paralogs["Parascopy<br/><small>SMN1/SMN2, opt-in</small>"]
    hla --> consensus36["PGx consensus<br/><small>outside calls</small>"]
    pypgx --> consensus36
    cyrius --> consensus36
    consensus36 --> pharmcat
    BAM --> telomere["TelomereHunter"]
    BAM --> coverage["mosdepth<br/>+ indexcov"]
    BAM --> mito["Mutect2<br/><small>Mitochondrial</small>"]
    BAM --> haplo["Haplogrep3<br/><small>mtDNA haplogroup</small>"]
    BAM --> sampleqc["somalier + VerifyBamID2<br/><small>Identity + contamination</small>"]
    BAM --> cram["CRAM archive"]

    %% Reporting
    clinical --> report["HTML Report<br/>+ MultiQC"]
    slivar --> report
    clinvar --> report
    pharmcat --> report
    cpsr --> report
    sampleqc --> report

    %% Styling
    classDef input fill:#0ea5e9,stroke:#0284c7,color:#fff
    classDef core fill:#8b5cf6,stroke:#7c3aed,color:#fff
    classDef analysis fill:#10b981,stroke:#059669,color:#fff
    classDef sv fill:#f59e0b,stroke:#d97706,color:#fff
    classDef annotation fill:#ec4899,stroke:#db2777,color:#fff
    classDef report fill:#ef4444,stroke:#dc2626,color:#fff

    class FASTQ,BAM,VCF input
    class fastp,align,DV core
    class clinvar,pharmcat,cpic,cpsr,eh,roh,prs,ancestry,pypgx,cyrius,hla,paralogs,consensus36,telomere,coverage,mito,haplo,sampleqc,cram analysis
    class manta,delly,cnvpytor,consensus,duphold,annotsv sv
    class vep,vcfanno,slivar,clinical annotation
    class report report

All Steps

# Step Tool Image variable Required?
1 ORA to FASTQ orad orad binary Only for Illumina ORA files
1b QC & Trimming fastp FASTP_IMAGE Recommended
2 Alignment minimap2 + samtools MINIMAP2_IMAGE + SAMTOOLS_IMAGE Yes (if starting from FASTQ)
3 Variant Calling DeepVariant DEEPVARIANT_IMAGE Yes
4 Structural Variants Manta MANTA_IMAGE Recommended
5 SV Annotation AnnotSV ANNOTSV_IMAGE If step 4 run
6 ClinVar Screen bcftools isec BCFTOOLS_IMAGE Yes
7 Pharmacogenomics PharmCAT PHARMCAT_IMAGE Yes
8 HLA Typing T1K T1K_IMAGE Optional
9 STR Expansions ExpansionHunter EXPANSIONHUNTER_IMAGE Recommended
9b STR Annotation Stranger STRANGER_IMAGE If step 9 run
10 Telomere Length TelomereHunter TELOMEREHUNTER_IMAGE Optional
11 ROH Analysis bcftools roh BCFTOOLS_IMAGE Recommended
12 Mito Haplogroup haplogrep3 + haplocheck HAPLOGREP3_IMAGE + HAPLOCHECK_IMAGE Optional (reads step 20's calls when they exist)
13 VEP Annotation VEP VEP_IMAGE Recommended
14 Imputation Prep bcftools BCFTOOLS_IMAGE Optional, opt-in (IMPUTATION=true)
15 SV Quality duphold DUPHOLD_IMAGE If step 4 run
16 Coverage QC indexcov GOLEFT_IMAGE Recommended
16b Coverage Stats mosdepth MOSDEPTH_IMAGE Recommended
17 Cancer Predisposition CPSR PCGR_IMAGE Recommended
18 CNV Calling CNVpytor CNVPYTOR_IMAGE Optional
19 SV Calling (Delly) Delly DELLY_IMAGE Optional
20 Mitochondrial GATK Mutect2 GATK_IMAGE Optional

Post-Processing Steps

These run after the core pipeline completes and combine outputs from earlier steps.

# Step Tool Image variable Required?
21 CYP2D6 Star Alleles Cyrius PYTHON_IMAGE Opt-in (TOOLS=...,cyrius; non-commercial licence)
22 SV Consensus Merge SURVIVOR SURVIVOR_IMAGE If two SV callers ran
23 Clinical Filter bcftools +split-vep BCFTOOLS_IMAGE If step 13 run
24 HTML Report bash + bcftools BCFTOOLS_IMAGE Recommended
25 Polygenic Risk Scores pgsc_calc PGSC_UTILS_IMAGE, PLINK2_IMAGE Exploratory; percentiles need the ancestry panel
26 Ancestry pgsc_calc PGSC_FRAPOSA_IMAGE, PGSC_UTILS_IMAGE After setup.sh --ancestry-panel (runs inside step 25)
27 CPIC Recommendations Python + CPIC PYTHON_IMAGE If step 7 run
28 MultiQC Report MultiQC MULTIQC_IMAGE Recommended
29 Somatic Variants GATK Mutect2 GATK_IMAGE Experimental, opt-in (SOMATIC=true)
30 Annotation Enrichment vcfanno VCFANNO_IMAGE If step 13 run
31 Variant Prioritization slivar SLIVAR_IMAGE If step 13 run
32 pypgx Pharmacogenomics pypgx PYPGX_IMAGE Recommended
33 Sample Identity and Contamination somalier + VerifyBamID2 SOMALIER_IMAGE + VERIFYBAMID2_IMAGE Recommended
34 CRAM Archive samtools SAMTOOLS_IMAGE Optional, when the analysis is done
35 Paralog Genes: SMN1/SMN2 Parascopy PARASCOPY_IMAGE Opt-in (TOOLS=...,parascopy)
36 PGx Consensus Python PYTHON_IMAGE Runs with step 7 when step 8, 21 or 32 ran
37 Y-Chromosome Haplogroup Yleaf YLEAF_IMAGE Opt-in (TOOLS=...,y_haplogroup, after setup.sh --yleaf-data; male samples)

What a default run covers

A default ./scripts/run-all.sh <sample> <sex> runs 32 numbered steps: 1b and 2 (only when there is no BAM yet), 3 (only when there is no VCF yet), 4, 5, 6, 7, 8, 9, 9b, 10, 11, 12, 13, 15, 16, 16b, 17, 18, 19, 20, 22, 23, 24, 25, 26, 27, 28, 30, 31, 32 and 36. Steps 5, 8, 13, 17, 18 and 32 are reported as skipped when their data is not installed, 25 when no score file is, 26 (inside the PRS run of step 25) when the ancestry panel is not, 23, 30 and 31 when step 13 did not run, and 36 (inside the PharmCAT stage) runs when step 8 or 32 did. It ends with the summary report (generate-report.sh).

Off unless you ask for them: 21 (TOOLS=...,cyrius, after setup.sh --cyrius), 35 (TOOLS=...,parascopy, after setup.sh --parascopy-data), KIR typing in step 8 (KIR=true, after setup.sh --kir-data), 4b (GRIDSS=true), 14 (IMPUTATION=true), 29 (SOMATIC=true), 37 (TOOLS=...,y_haplogroup, after setup.sh --yleaf-data, male samples), the alternative callers 3a to 3d (EXTRA_CALLERS=gatk,freebayes,strelka2,octopus) and the caller comparison (BENCHMARK=true). Step 1 (ORA input) and the other alternative scripts (2a, 2b, 3e, 4a, 4c) run only by hand. Steps 33 and 34 run by hand, or in the Nextflow pipeline with sample_qc and cram_archive in --tools.

The Nextflow pipeline runs the same chain from a samplesheet, from FASTQ (steps 1b, 2, 16 and 3) to the report, with the steps above that have a module.

The minimum useful run is steps 2, 3, 6 and 7 (alignment, variant calling, ClinVar, PharmCAT). Runtimes per step and for a whole run are on Hardware and storage requirements.

Alternative Tools (Benchmarking)

No single variant caller is universally best. The pipeline includes alternative tools that output to separate directories so you can compare results without overwriting the defaults.

Script Tool Alternative To Output Directory
02a BWA-MEM2 minimap2 (step 2) aligned_bwamem2/
03a GATK HaplotypeCaller DeepVariant (step 3) vcf_gatk/
03b FreeBayes DeepVariant (step 3) vcf_freebayes/
04a TIDDIT Manta (step 4) sv_tiddit/
03c Strelka2 DeepVariant (step 3) vcf_strelka2/
03d Octopus DeepVariant (step 3) vcf_octopus/
04b GRIDSS Manta (step 4) sv_gridss/
benchmark bcftools isec / hap.py — benchmark/

See docs/benchmarking.md for how to run and interpret results, and docs/tool-rationale.md for why each default was chosen.