Skip to content

Personal Genome Pipeline

Personal Genome Pipeline: no cloud, no subscription, just your data

Release CI GitHub Stars License: GPL-3.0-or-later


Personal Genome Pipeline turns the files a consumer sequencing vendor gives you (FASTQ, BAM, VCF or Illumina ORA) into a full genomic profile on your own computer: small and structural variants, ClinVar and cancer-predisposition screening, pharmacogenomics, repeat expansions, HLA type, telomere content, mitochondrial haplogroup and heteroplasmy, runs of homozygosity, ancestry and polygenic risk scores. A vendor's own report covers a fraction of this and keeps your genome on their servers; a clinical lab charges per panel. Here every step is one Docker container with CPU and memory limits, run by a Nextflow pipeline (scripts/run-all.sh starts it for one sample) or from one bash script per step, and no script sends your data anywhere. Start with Getting started, then run the quick test on public data before your own.

  • Getting started


    Prerequisites, the four entry paths (FASTQ, BAM, VCF, ORA), the first run and what "it worked" looks like.

  • Quick test


    Run the pipeline on a small public sample first, so a setup problem shows up in minutes, not after twelve hours.

  • Pipeline overview


    Every step with its tool, its image and whether it is required, which steps a default run includes, plus the pipeline graph.

  • Interpreting your results


    What each report means, what is normal, what needs a genetic counsellor, and what to re-run when databases update.

What it finds

  • Small variants: SNPs and indels with DeepVariant; structural and copy-number variants with Manta, Delly and CNVpytor merged into one consensus set.
  • Clinical screening: known pathogenic variants from ClinVar, cancer predisposition panels with CPSR, VEP annotation with CADD, SpliceAI, REVEL and AlphaMissense, and slivar prioritisation of rare deleterious variants.
  • Pharmacogenomics: PharmCAT and pypgx star alleles across 23 genes, CYP2D6 structural alleles with Cyrius, CPIC dosing recommendations.
  • Ancestry and lineage: mitochondrial haplogroup and heteroplasmy, runs of homozygosity, ancestry SNP intersection, polygenic risk scores for ten conditions.
  • Everything else the reads hold: HLA type, repeat expansions at the 31 loci of ExpansionHunter's catalog (Huntington's, Fragile X, ALS among them), telomere content, sex-chromosome check, coverage statistics, one HTML report and a MultiQC summary.
The step 24 HTML report: cards for quality control, variant calling, ClinVar, pharmacogenomics, CPIC, CYP2D6 across callers, HLA typing, structural variants, cancer predisposition, repeat expansions, runs of homozygosity, the mitochondrial and Y haplogroups, telomere length, mitochondrial variants, the clinical filter and slivar, then the polygenic risk scores table
The HTML report from step 24, run on DEMO-001, an invented sample. Every number in it is made up.

Pipeline steps

One page per step, grouped by stage. The step number is the script name and the output directory.

Alternative callers for benchmarking (BWA-MEM2, GATK HaplotypeCaller, FreeBayes, Strelka2, TIDDIT) are scripts without a page of their own; Variant caller benchmarking covers how to run and compare them.

How it runs

graph LR
    FASTQ["FASTQ / ORA"] --> fastp["fastp"]
    fastp --> align["minimap2"]
    align --> BAM["Sorted BAM"]
    BAM --> DV["DeepVariant"]
    DV --> VCF["VCF"]
    VCF --> vcfsteps["ClinVar, PharmCAT, VEP, CPSR,<br/>ROH, PRS, ancestry, haplogroup"]
    BAM --> bamsteps["Manta, Delly, CNVpytor, ExpansionHunter,<br/>HLA, telomeres, pypgx, Cyrius, mosdepth"]
    vcfsteps --> report["HTML report + MultiQC"]
    bamsteps --> report
  • Each step is one docker run with an image tag or digest from versions.env (listed on Image versions), a CPU limit and a memory limit, so a step cannot take the machine down.
  • Two ways to run it: ./scripts/run-all.sh <sample> <male|female> starts the Nextflow pipeline, which runs independent steps in parallel and resumes after a failure (Full run); or one bash script per step under scripts/, with Docker alone.
  • The minimum useful run is alignment, DeepVariant, ClinVar and PharmCAT, about 4 to 7 hours on a 16-core desktop; a default run is 6 to 12 hours. Hardware and storage requirements gives the per-step time, memory and disk figures; a 30X sample needs about 500 GB.
  • After reference data setup a run downloads only a few public files, listed in Why run locally?. A BAM or VCF from your vendor skips alignment or variant calling; Getting started has the entry paths and Vendor compatibility the per-vendor notes.

What it does not do

  • It is not a medical device and has not been clinically validated. Findings go to a genetic counsellor or physician before any decision; the pipeline uses the same tools a clinical lab uses, but without a lab's validation and confirmation.
  • It does not give ancestry percentages or population percentiles for polygenic risk scores: both need a reference cohort the pipeline does not ship. Step 25 and step 26 say what the raw numbers can and cannot support.
  • It does not call CYP2D6 reliably from a VCF alone; that gene needs the BAM-based callers (Cyrius, pypgx), and short-read WGS still misses some alleles.
  • It does not impute. Step 14 prepares files for an external imputation server; sending them there is your decision.

Privacy

  • Your reads, alignments and variants stay in GENOME_DIR on your disk. No script uploads data. A run still pulls images and fetches a few public files; Why run locally? lists each one and what the HTML reports load when you open them.
  • The example outputs on these pages are invented or use placeholders. No page shows a real person's genotype, haplogroup, HLA type, score or sample name.
  • Why run locally? compares the cost and the exposure of the alternatives.

Guides and reference

Getting help

License

Personal Genome Pipeline is released under the GPL-3.0-or-later license. It is for educational and research use; it is not a medical device.