Troubleshooting Guide¶
Definitive reference for diagnosing and fixing problems with the Personal Genome Pipeline. Organized by symptom so you can Ctrl+F your error message and find the fix.
Before troubleshooting: Run docker info to confirm Docker is running, and check docker stats for resource usage. Most failures are either Docker resource limits or wrong file paths inside containers.
Table of Contents¶
- Docker Issues
- Input Data Issues
- Per-Step Troubleshooting
- Performance Issues
- Output Issues
- Getting Help
Docker Issues¶
Container exits immediately with no output¶
Symptom: docker run returns instantly. No error message. Exit code 137.
Cause: The container was OOM-killed (Out Of Memory). Docker silently kills containers that exceed their --memory limit. Exit code 137 = SIGKILL from the kernel OOM killer.
How to confirm:
# Check the last container's exit status
docker ps -a --latest --format "{{.Status}}"
# Output like "Exited (137)" confirms OOM
# Check system logs for OOM events
dmesg | grep -i "oom\|killed" | tail -10
# On Docker Desktop (Mac/Windows)
# Check Docker Desktop > Troubleshoot > Logs
Fix:
1. Increase the --memory flag in the script that failed. With run-all.sh you do not edit a script: the pipeline gives each task a memory request and doubles it on the one retry after an out-of-memory exit, up to --max_memory
2. Reduce parallelism (run fewer steps simultaneously). With run-all.sh, set --max_memory to what Docker may use, e.g. ./scripts/run-all.sh <sample> <sex> --max_memory 24.GB (without it, the machine's RAM), and THREADS=N to cap the CPUs of each task (--max_cpus, by default the machine's CPU count); a rerun reuses the steps that finished
3. Increase Docker Desktop memory allocation (see Docker Desktop not enough memory)
4. For DeepVariant, reduce --num_shards (each shard needs ~2-4 GB)
Container exits with code 1 but no error message¶
Symptom: Script fails, exit code 1, but the container printed nothing useful.
Cause: Most bioinformatics tools print errors to stderr, which may not be captured depending on how you ran the script.
Fix:
source versions.env # from the repository root
# Re-run with full debug output
bash -x scripts/03-deepvariant.sh your_name 2>&1 | tee debug.log
# Or run the docker command manually with -it for interactive output
docker run --rm -it \
--cpus 4 --memory 8g \
-v "${GENOME_DIR}:/genome" \
"${DEEPVARIANT_IMAGE}" \
/bin/bash
# Then run the command inside the container to see the error
"No such image" / "manifest for ... not found"¶
Symptom: Error response from daemon: manifest for quay.io/biocontainers/toolname:tag not found
Cause: Biocontainer image tags change frequently. The hash suffix (e.g., --hd03093a_1) encodes the conda build environment and breaks between releases.
Fix: 1. Check the current tag at the container registry:
# For quay.io images
# Visit: https://quay.io/repository/biocontainers/TOOLNAME?tab=tags
# For Docker Hub images
docker search toolname
| Tool | Wrong image | Use instead |
|---|---|---|
| Delly | an old delly 1.2.9 build tag |
DELLY_IMAGE |
| ExpansionHunter | weisburd/expansionhunter (v2.5.5, a different CLI) |
EXPANSIONHUNTER_IMAGE |
| AnnotSV | bioinfochrustrasbourg/annotsv |
ANNOTSV_IMAGE |
| CPSR | sigven/cpsr (does not exist) |
PCGR_IMAGE (bundles both) |
| MToolBox | robertopreste/mtoolbox (does not exist) |
GATK_IMAGE (step 20 uses Mutect2) |
- If an image disappears entirely, check if the tool has an official Docker image on GitHub Container Registry (
ghcr.io), Docker Hub, or the tool's documentation.
"Permission denied" writing to mounted volumes¶
Symptom: OSError: [Errno 13] Permission denied: '/genome/sample/output' or similar.
Cause: Most bioinformatics Docker images run as a non-root user (e.g., UID 1000). When writing to bind-mounted host directories, the container user may not have write access.
Fix: Add --user root to the docker run command. All pipeline scripts already include this flag, but if you are running commands manually:
Note: Files created with --user root will be owned by root on the host. To fix ownership afterward:
Docker Desktop not enough memory (Mac/Windows)¶
Symptom: Containers fail with exit code 137 even though your machine has enough RAM.
Cause: Docker Desktop runs inside a virtual machine with a fixed memory allocation. Default is often 2-4 GB, which is far too little for genomics.
Fix for macOS: 1. Open Docker Desktop > Settings > Resources 2. Set Memory to at least 16 GB (32 GB recommended) 3. Set Swap to at least 4 GB 4. Click Apply & Restart
Fix for Windows (WSL2):
Create or edit %UserProfile%\.wslconfig:
wsl --shutdown and reopen your terminal.
Verification:
Containers stuck / running forever¶
Symptom: A container has been running for much longer than expected. No output. Not using CPU.
How to diagnose:
# Check what's running
docker ps
# Check resource usage — is the container actually doing work?
docker stats --no-stream
# Check container logs
docker logs <container_id>
# If CPU is at 0% for a long time, the container may be stuck
Common causes:
1. Waiting for stdin: Some tools expect interactive input. Fix: add -i flag or ensure all parameters are passed on the command line.
2. I/O blocked on slow storage: NFS/SMB/USB mounts can stall Docker I/O. Check docker stats — if CPU is 0% and memory is stable, the container is likely waiting for I/O.
3. Java heap space: Tools like PharmCAT or GATK may be stuck in garbage collection. Fix: increase --memory.
Fix:
# Kill the stuck container
docker kill <container_id>
# Or if you started with run-all.sh, stop Nextflow with Ctrl-C: it stops its tasks'
# containers. The same command later reruns only the tasks that did not finish.
"No space left on device" mid-analysis¶
Symptom: Container crashes with "No space left on device", "ENOSPC", or "write failed: No space left".
Cause: Docker needs free space in TWO places:
1. Your data directory (GENOME_DIR) for output files
2. Docker's storage driver (usually /var/lib/docker) for container layers and temp files
Diagnosis:
# Check host filesystem
df -h ${GENOME_DIR}
df -h /var/lib/docker # Linux
# Docker Desktop stores data differently — check Docker Desktop settings
# Check Docker disk usage
docker system df
Fix:
# Clean up Docker (removes unused containers, images, build cache)
docker system prune -a
# WARNING: This removes ALL unused images. You'll need to re-pull them.
# Safer: only remove stopped containers and dangling images
docker container prune
docker image prune
# Delete intermediate files from completed steps
rm -f ${GENOME_DIR}/${SAMPLE}/cnvpytor/*.pytor # a few GB each
rm -f ${GENOME_DIR}/${SAMPLE}/delly/*.bcf # After VCF conversion
Space requirements per sample:
| Phase | Space Needed | Cumulative |
|---|---|---|
| FASTQ input | 60-90 GB | 90 GB |
| BAM (step 2) | 80-120 GB | 210 GB |
| VCF + analyses | 10-30 GB | 240 GB |
| VEP output | 2-5 GB | 245 GB |
| CNVpytor .pytor file | 2-5 GB | 250 GB |
| Recommended free | 500 GB |
Input Data Issues¶
Wrong genome build (hg19 vs hg38)¶
Symptom: Variant calling produces far fewer variants than expected (e.g., 100K instead of 4-5 million). Or ClinVar screen finds 0 hits.
How to detect:
source versions.env # from the repository root
# Check BAM header for chromosome naming and lengths
docker run --rm -v "${GENOME_DIR}:/genome" "${SAMTOOLS_IMAGE}" \
samtools view -H /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam | grep "^@SQ" | head -3
# GRCh38 (hg38): SN:chr1 LN:248956422
# GRCh37 (hg19): SN:1 LN:249250621 (or SN:chr1 LN:249250621)
# For VCF files
docker run --rm -v "${GENOME_DIR}:/genome" "${BCFTOOLS_IMAGE}" \
bcftools view -h /genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz | grep "^##contig" | head -3
Key differences:
| Feature | GRCh37 / hg19 | GRCh38 / hg38 |
|---|---|---|
| chr1 length | 249,250,621 | 248,956,422 |
| Chromosome prefix | Often no chr |
chr in UCSC-style files (what the pipeline needs); none in Ensembl-style files |
| Mitochondria name | MT (or chrM) |
chrM (UCSC style) or MT (Ensembl style) |
| ALT contigs | No | Yes in the full assembly; the pipeline's reference (the no-ALT analysis set) leaves them out |
The prefix alone does not tell the build: some providers deliver GRCh38 with Ensembl names (1, MT). Check chr1's length. A GRCh38 VCF with Ensembl names needs only a rename, which the Nextflow pipeline prints when it stops on it: see Starting from a Vendor VCF.
Fix (GRCh37 data): Extract FASTQ from BAM and re-align to GRCh38:
source versions.env # from the repository root
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
bash -c "samtools sort -n /genome/${SAMPLE}/aligned/old_hg19.bam | \
samtools fastq -1 /genome/${SAMPLE}/fastq/${SAMPLE}_R1.fastq.gz \
-2 /genome/${SAMPLE}/fastq/${SAMPLE}_R2.fastq.gz -"
# Then run alignment (step 2) which uses GRCh38
./scripts/02-alignment.sh $SAMPLE
Do not use LiftOver on BAM files. Re-alignment from FASTQ is cleaner and avoids coordinate translation artifacts.
CRAM files instead of BAM¶
Symptom: Pipeline scripts expect .bam files but you have .cram files.
Fix: Convert CRAM to BAM (requires the reference genome used for encoding, which is usually GRCh38):
source versions.env # from the repository root
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
docker run --rm --user root \
-v ${GENOME_DIR}:/genome \
"${SAMTOOLS_IMAGE}" \
samtools view -b \
-T "/genome/${REF_FASTA}" \
-o /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam \
/genome/${SAMPLE}/aligned/${SAMPLE}.cram
docker run --rm --user root \
-v ${GENOME_DIR}:/genome \
"${SAMTOOLS_IMAGE}" \
samtools index /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam
If you get an error about mismatched references: The CRAM was encoded against a different reference build. You need the exact FASTA used during encoding, or extract FASTQ and re-align:
source versions.env # from the repository root
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
bash -c "samtools sort -n /genome/${SAMPLE}/aligned/${SAMPLE}.cram | \
samtools fastq -1 /genome/${SAMPLE}/fastq/${SAMPLE}_R1.fastq.gz \
-2 /genome/${SAMPLE}/fastq/${SAMPLE}_R2.fastq.gz -"
Paired FASTQ naming issues¶
Symptom: Step 2 (alignment) fails with "ERROR: File not found" for R1 or R2.
Cause: The pipeline expects FASTQ files named exactly:
${GENOME_DIR}/${SAMPLE}/fastq/${SAMPLE}_R1.fastq.gz
${GENOME_DIR}/${SAMPLE}/fastq/${SAMPLE}_R2.fastq.gz
Vendors use different naming conventions:
# Illumina standard
Sample_S1_L001_R1_001.fastq.gz / Sample_S1_L001_R2_001.fastq.gz
# BGI/DNBSEQ
V350012345_L01_1.fq.gz / V350012345_L01_2.fq.gz
# Other
sample.1.fastq.gz / sample.2.fastq.gz
Fix: Create symlinks or rename to match the expected pattern:
cd ${GENOME_DIR}/${SAMPLE}/fastq/
# If you have multiple lanes, concatenate first:
cat *_R1_*.fastq.gz > ${SAMPLE}_R1.fastq.gz
cat *_R2_*.fastq.gz > ${SAMPLE}_R2.fastq.gz
# Or create symlinks for single files:
ln -s original_name_R1.fastq.gz ${SAMPLE}_R1.fastq.gz
ln -s original_name_R2.fastq.gz ${SAMPLE}_R2.fastq.gz
Warning: If concatenating multi-lane FASTQ, ensure R1 and R2 files are concatenated in the same lane order. Mismatched pairs will produce corrupt alignments.
Corrupt or truncated downloads¶
Symptom: Tools crash with "unexpected end of file", "truncated file", or "not in gzip format".
How to detect:
source versions.env # from the repository root
# Test gzip integrity (for FASTQ.gz, VCF.gz)
gzip -t ${GENOME_DIR}/${SAMPLE}/fastq/${SAMPLE}_R1.fastq.gz
# If corrupt: "unexpected end of file" or "invalid compressed data"
# Test BAM integrity
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
samtools quickcheck /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam
# No output = OK. Error message = corrupt.
# Check VCF can be read
docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools view -h /genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz > /dev/null
Fix: Re-download the file. For large files, use wget -c (supports resume):
For reference data downloads that fail repeatedly, see the per-step sections for VEP cache and PCGR data bundle.
BAM not sorted or not indexed¶
Symptom: Tools fail with "file is not sorted", "index file not found", or "[E::hts_idx_load3] Could not load index".
How to detect:
source versions.env # from the repository root
# Check if BAM is sorted
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
samtools view -H /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam | grep "^@HD"
# Should show: SO:coordinate
# Check if index exists
ls -la ${GENOME_DIR}/${SAMPLE}/aligned/${SAMPLE}_sorted.bam.bai
Fix:
source versions.env # from the repository root
# Sort the BAM (if not already sorted)
docker run --rm --user root \
--cpus 4 --memory 8g \
-v ${GENOME_DIR}:/genome \
"${SAMTOOLS_IMAGE}" \
samtools sort -@ 4 \
-o /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam \
/genome/${SAMPLE}/aligned/${SAMPLE}_unsorted.bam
# Create index (always needed)
docker run --rm --user root \
-v ${GENOME_DIR}:/genome \
"${SAMTOOLS_IMAGE}" \
samtools index /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam
Note: The BAM index file must have the exact same base name as the BAM. If your BAM is sample_sorted.bam, the index must be sample_sorted.bam.bai (not sample_sorted.bai).
Per-Step Troubleshooting¶
Step 2: Alignment (minimap2 + samtools)¶
Problem: "Cannot find FASTQ" but files exist
The script looks for ${SAMPLE}_R1.fastq.gz and ${SAMPLE}_R2.fastq.gz. See Paired FASTQ naming issues.
Problem: Alignment produces 0-byte BAM
Usually means minimap2 failed silently. Run with bash -x to see the actual error. Common causes:
- Wrong reference genome path
- FASTQ files are corrupt (run gzip -t on both)
- Docker --memory too low (minimap2 needs ~6-10 GB for the GRCh38 index)
Problem: minimap2 index build takes forever
The .mmi index build is a one-time step (~30 minutes). If it seems stuck, check that the reference FASTA is not corrupt and the output path is writable.
Step 3: DeepVariant GPU acceleration (not worth it)¶
Symptom: You want to use GPU acceleration for DeepVariant to speed it up.
Reality: DeepVariant's GPU Docker image is built against one specific CUDA version, so it needs an NVIDIA driver that supports that version. Check the release notes of the DeepVariant version you run before trying it.
Why it's not worth the hassle:
1. The GPU image (the same tag as DEEPVARIANT_IMAGE with -gpu appended) requires the NVIDIA container runtime and a driver that matches its CUDA version
2. Only call_variants uses the GPU. In the DeepVariant 1.10 runtime metrics for a 30x WGS on 96 CPU cores, make_examples takes 46 min, call_variants 16 min and postprocess_variants 7 min, so a GPU speeds up about a quarter of the run
3. On a 16-core CPU, DeepVariant finishes in 2-4 hours — GPU saves maybe 30-60 minutes
Recommendation: Use the CPU image (DEEPVARIANT_IMAGE) with --num_shards set to your core count. If you need it faster, run on a cloud instance with more CPU cores rather than fighting CUDA compatibility.
Step 3: DeepVariant crashes on Mac (amd64 emulation)¶
Symptom: DeepVariant container exits with code 137 or crashes with memory errors on Apple Silicon Mac.
Cause: DeepVariant is an amd64-only image. On Apple Silicon Macs, Docker runs it through Rosetta 2 emulation, which adds ~2x memory overhead and 3-5x CPU overhead. The default --memory 32g and --cpus 8 settings are too aggressive for emulation.
Fix:
source versions.env # from the repository root
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
# scripts/03-deepvariant.sh has fixed limits (--cpus 8 --memory 32g), so run step 3
# manually with reduced resources:
docker run --rm \
--cpus 2 --memory 12g \
-v "${GENOME_DIR}:/genome" \
"${DEEPVARIANT_IMAGE}" \
/opt/deepvariant/bin/run_deepvariant \
--model_type=WGS \
--ref="/genome/${REF_FASTA}" \
--reads="/genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam" \
--output_vcf="/genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz" \
--num_shards=2
Expected runtime on Mac: 8-16 hours (vs 2-4 hours on native Linux amd64).
Alternative: If DeepVariant repeatedly crashes on your Mac, run ONLY step 3 on a cloud Linux instance (e.g., Hetzner CCX33 for ~$0.18/hr) and download the VCF. All other steps can proceed on Mac.
Step 4: Manta produces no variants¶
Symptom: diploidSV.vcf.gz is empty or has only header lines.
Common causes:
1. BAM has too few reads (low-coverage WGS below ~10X)
2. BAM was not indexed (Manta requires .bai)
3. Reference FASTA mismatch — chromosome names must match between BAM and FASTA
Diagnosis:
source versions.env # from the repository root
# Check BAM read count
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
samtools flagstat /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam
# "mapped" count should be 500M+ for 30X WGS
# Check chromosome naming consistency
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
samtools view -H /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam | grep "^@SQ" | head -1
# Must use "chr" prefix (SN:chr1) to match GRCh38 reference
Step 5: AnnotSV produces no output¶
Symptom: AnnotSV runs but the output TSV is empty or missing.
Common causes:
1. Input VCF path wrong: AnnotSV expects the Manta diploidSV.vcf.gz. The script checks both manta/ and manta2/ directories.
2. No PASS variants: If Manta flagged all SVs as filtered (no PASS), AnnotSV may produce empty output.
3. Image version mismatch: Some AnnotSV versions require specific internal database versions.
Fix:
source versions.env # from the repository root
# Verify the Manta VCF has PASS variants
docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools view -f PASS /genome/${SAMPLE}/manta/results/variants/diploidSV.vcf.gz | grep -c -v "^#"
# Should be > 0. Typical: 7,000-9,000 for 30X WGS.
If the Manta VCF is valid but AnnotSV still produces nothing, try running with verbose output:
source versions.env # from the repository root
docker run --rm --user root \
-v "${GENOME_DIR}:/genome" \
"${ANNOTSV_IMAGE}" \
AnnotSV \
-SVinputFile "/genome/${SAMPLE}/manta/results/variants/diploidSV.vcf.gz" \
-outputFile "/genome/${SAMPLE}/annotsv/${SAMPLE}_test.tsv" \
-genomeBuild GRCh38 \
-annotationMode both 2>&1 | tee annotsv_debug.log
Step 7: PharmCAT VCF format requirements¶
Symptom: PharmCAT exits with "Input VCF does not meet requirements" or produces empty report with no gene calls.
Cause: PharmCAT is strict about VCF format:
- Must be block-gzipped (.vcf.gz, not plain .vcf)
- Must have a tabix index (.vcf.gz.tbi)
- Must be aligned to GRCh38
- Must contain a GT (genotype) FORMAT field
- Multi-allelic sites must be decomposed
Fix:
source versions.env # from the repository root
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
# Verify your VCF has the GT field
docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools query -f '[%GT]\n' /genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz | head -1
# Should show something like "0/1" or "1/1"
# If your VCF is plain text, compress and index it:
docker run --rm --user root -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bash -c "bcftools view /genome/${SAMPLE}/vcf/${SAMPLE}.vcf -Oz \
-o /genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz && \
bcftools index -t /genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz"
# If multi-allelic sites are an issue, normalize:
docker run --rm --user root -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools norm -m -both \
-f "/genome/${REF_FASTA}" \
/genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz \
-Oz -o /genome/${SAMPLE}/vcf/${SAMPLE}_norm.vcf.gz
Note: PharmCAT 3.2.0 includes a VCF preprocessor. If direct input fails, try the preprocessor first:
source versions.env # from the repository root
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
docker run --rm \
--cpus 2 --memory 4g \
-v "${GENOME_DIR}/${SAMPLE}/vcf:/data" \
-v "${GENOME_DIR}:/genome" \
"${PHARMCAT_IMAGE}" \
python3 /pharmcat/pharmcat_vcf_preprocessor \
-vcf "/data/${SAMPLE}.vcf.gz" \
-refFna "/genome/${REF_FASTA}" \
-o /data/ \
-bf "$SAMPLE"
Step 9: ExpansionHunter fails immediately¶
Symptom: ExpansionHunter crashes with model or catalog errors.
Cause: The pipeline uses ExpansionHunter v5.0.0 (EXPANSIONHUNTER_IMAGE). The v5 CLI is different from the old v2.5.5 in the weisburd/expansionhunter image.
Key v5 CLI changes from v2.5.5:
- --bam → --reads
- --ref-fasta → --reference
- --repeat-specs <directory> → --variant-catalog <file.json>
- --vcf/--json/--log → --output-prefix (auto-generates .vcf, .json)
- The variant catalog (31 pathogenic GRCh38 loci) is bundled inside the container at /usr/local/share/ExpansionHunter/variant_catalog/grch38/variant_catalog.json
Fix: If running manually, use the v5 syntax:
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
ExpansionHunter \
--reads /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam \
--reference "/genome/${REF_FASTA}" \
--variant-catalog /usr/local/share/ExpansionHunter/variant_catalog/grch38/variant_catalog.json \
--output-prefix /genome/${SAMPLE}/expansion_hunter/${SAMPLE}_eh \
--threads 4 \
--sex male
Step 13: bcftools +split-vep and INFO/CSQ field parsing¶
Symptom: You're trying to extract VEP annotations from the annotated VCF using bcftools query or bcftools +split-vep, but the output is empty or garbled.
Cause: VEP stores all annotations in a single INFO/CSQ field as a pipe-delimited string. This is not a standard VCF INFO field — it is a compound annotation that requires special handling.
Common mistakes:
-
Using
bcftools query -f '%INFO/CSQ'directly: This dumps the raw pipe-delimited string, which is unreadable. Usebcftools +split-vepinstead. -
Wrong column names in +split-vep: The column names in CSQ depend on your VEP command-line options. Check the VCF header for the actual field order:
-
Forgetting to specify -f (format) in +split-vep: Without
-f, it outputs all fields. Specify the ones you want, and choose how multiple transcripts are reported:source versions.env # from the repository root docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \ bcftools +split-vep \ /genome/${SAMPLE}/vep/${SAMPLE}_vep.vcf.gz \ -f '%CHROM %POS %Consequence %SYMBOL %SIFT %PolyPhen %gnomADe_AF\n' \ -s worst # -s worst keeps only the most severe consequence per variant. # -d selects nothing: it puts each kept consequence on its own line, so without -s # it prints every transcript's consequence. -
gnomAD field naming confusion: Depending on VEP version and cache, the gnomAD frequency field may be named
gnomADe_AF,gnomAD_AF,AF, orMAX_AF. Always check the CSQ header first.
Step 13: VEP cache download failures¶
Symptom: VEP cache download times out, produces a partial file, or extraction fails.
Cause: The VEP cache is ~26 GB, hosted on Ensembl FTP servers that can be slow. The built-in INSTALL.pl downloader does not support resume and may fail silently.
Fix — manual download with resume:
mkdir -p ${GENOME_DIR}/vep_cache/tmp
cd ${GENOME_DIR}/vep_cache/tmp
# wget -c resumes interrupted downloads
# Release 116 matches the pinned VEP image (VEP_IMAGE in versions.env)
wget -c https://ftp.ensembl.org/pub/release-116/variation/indexed_vep_cache/homo_sapiens_vep_116_GRCh38.tar.gz
# Verify download size (~26 GB)
ls -lh homo_sapiens_vep_116_GRCh38.tar.gz
# Extract (takes 10-20 minutes, expands to ~30 GB)
cd ${GENOME_DIR}/vep_cache
tar xzf tmp/homo_sapiens_vep_116_GRCh38.tar.gz
# Verify extraction
ls ${GENOME_DIR}/vep_cache/homo_sapiens/116_GRCh38/
# Should contain info.txt, variation_set_*.gz, and many other files
Symptom: VEP INSTALL.pl fails with "Cannot open Local file" permission error.
Cause: The VEP container runs as a non-root user who cannot write to /opt/vep/.vep/tmp/.
Fix: Use --user root when running VEP and pre-create the temp directory. Or (recommended) download the cache manually as shown above.
Step 17: CPSR data directory or CLI issues¶
Symptom (PCGR 2.x): refdata_dir not found or similar errors about missing bundle paths.
Cause: PCGR 2.x uses a completely different CLI and volume mount structure from 1.x. The old --pcgr_dir flag no longer exists. You now need --refdata_dir and --vep_dir as separate arguments, each with its own Docker volume mount.
Fix: Ensure you have the correct data bundle and mount structure:
source versions.env # from the repository root
# PCGR 2.x requires four separate volume mounts:
docker run --rm --user root \
-v ${GENOME_DIR}/vep_cache:/mnt/.vep \
-v ${GENOME_DIR}/pcgr_data/${PCGR_DATA_BUNDLE}:/mnt/bundle \
-v ${GENOME_DIR}/${SAMPLE}/vcf:/mnt/inputs \
-v ${GENOME_DIR}/${SAMPLE}/cpsr:/mnt/outputs \
"${PCGR_IMAGE}" cpsr \
--refdata_dir /mnt/bundle \
--vep_dir /mnt/.vep \
...
If you have the old 1.x data bundle (pcgr.databundle.grch38.20220203.tgz), it is not compatible with PCGR 2.x. Download the new bundle:
cd ${GENOME_DIR}/pcgr_data
wget -c https://insilico.hpc.uio.no/pcgr/pcgr_ref_data.20260620.grch38.tgz
tar xzf pcgr_ref_data.20260620.grch38.tgz
mkdir -p 20260620 && mv data/ 20260620/
Additional CPSR issue -- wrong Docker image: CPSR does not have its own Docker image. It is bundled inside the PCGR image:
source versions.env # from the repository root
# CORRECT:
docker run ... "${PCGR_IMAGE}" cpsr ...
# WRONG (there is no sigven/cpsr image):
docker run ... sigven/cpsr ...
Step 18: CNVpytor missing resources or empty output¶
Symptom: cnvpytor -his aborts with "Some reference genome resource files are missing", or the final output has 0 CNV calls.
Common causes:
1. GC/mask resources not staged. The 1.3.2 biocontainer ships without them and its -download is broken, so the pinned files must be present at ${GENOME_DIR}/reference/cnvpytor/ and bind-mounted (see 00-reference-setup.md). All seven files must exist or the resource check aborts.
2. .pytor not writable / data dir read-only. Add --user root; the step needs a writable container FS (docker).
3. BAM chromosome naming mismatch. CNVpytor inherits chromosome names from the BAM header — an hg38 BAM must use chr1-style names.
Diagnosis:
source versions.env # from the repository root
# Confirm all 7 pinned resource files are present
ls -lh ${GENOME_DIR}/reference/cnvpytor/
# Check the .pytor size (should be several GB for 30X WGS)
ls -lh ${GENOME_DIR}/${SAMPLE}/cnvpytor/${SAMPLE}.pytor
# Re-verify the container data path for the pinned image (mount target)
docker run --rm "${CNVPYTOR_IMAGE}" \
python -c 'import cnvpytor, os; print(os.path.dirname(cnvpytor.__file__) + "/data")'
Fix: Download the pinned resource files (docs/00-reference-setup.md), then re-run. Confirm the BAM path resolves to the containerized /genome/... path.
Step 19: Delly long runtime / high memory¶
Symptom: Delly has been running for 8+ hours on a 30X genome.
Cause: Delly examines every read pair across the genome for split-read and paired-end SV evidence. On high-coverage samples or with slow storage, this takes longer.
Expected runtimes:
| Coverage | CPU Cores | Expected Runtime |
|---|---|---|
| 30X | 4 | 2-4 hours |
| 30X | 8 | 1.5-3 hours |
| 50X+ | 4 | 4-8 hours |
If Delly exceeds 12 hours:
1. Check docker stats to verify it is still using CPU (not stuck)
2. Confirm storage is not a bottleneck (NFS/SMB mounts are much slower)
3. Consider running only on specific chromosomes to reduce scope:
source versions.env # from the repository root
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
# Delly on the sample's BAM without the exclude map step 19 passes
docker run --rm --user root \
--cpus 4 --memory 8g \
-v "${GENOME_DIR}:/genome" \
"${DELLY_IMAGE}" \
delly sr \
-g "/genome/${REF_FASTA}" \
-o /genome/${SAMPLE}/delly/${SAMPLE}_sv.bcf \
/genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam
# Delly processes every contig of the reference; slowness usually comes
# from centromeres, telomeres and the unplaced scaffolds. Step 19 skips
# them with Delly's exclude map (-x reference/delly_human.hg38.excl.tsv,
# installed by setup.sh); add it here the same way.
Step 20: GATK Mutect2 mitochondrial mode¶
Problem: "Cannot create sequence dictionary"
The script tries to create the reference's sequence dictionary (GRCh38_no_alt_analysis_set.dict) if it does not exist. If you get permission errors, create it manually:
source versions.env # from the repository root
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
docker run --rm --user root \
-v "${GENOME_DIR}:/genome" \
"${GATK_IMAGE}" \
gatk CreateSequenceDictionary \
-R "/genome/${REF_FASTA}" \
-O "/genome/${REF_FASTA%.fasta}.dict"
Problem: 0 mitochondrial variants called
Possible causes:
- BAM has no chrM reads (check with samtools idxstats | grep chrM)
- Mitochondrial chromosome name mismatch (MT vs chrM). GRCh38 uses chrM.
HLA Typing (Step 8): Known difficulties¶
HLA typing from WGS data is unreliable in Docker. The two main tools have unresolved issues:
HLA-LA: Requires a pre-serialized 40 GB graph. Most Docker images do not include it. The one image that does (jiachenzdocker/hla-la, pinned by digest in step 8) crashes during graph alignment. Status: UNSOLVED.
T1K: Works partially, but the coordinate file build step requires the full reference FASTA (not the .fai index). With correctly built coordinates, T1K can call some HLA alleles but may have ~50% unmapped alleles.
Recommendation: For clinical HLA typing, rely on dedicated lab assays. Short-read WGS is not ideal for this highly polymorphic region.
Performance Issues¶
Everything is slow on macOS Apple Silicon¶
Cause: All bioinformatics Docker images are amd64/x86_64. On Apple Silicon Macs (M1-M4), Docker Desktop runs them through Rosetta 2 emulation. This adds: - 2-5x CPU overhead (emulation is compute-expensive) - ~2x memory overhead (emulation runtime consumes extra RAM) - I/O overhead from the Docker Desktop VM
Expected slowdown by step:
| Step | Native Linux | Mac (Rosetta 2) | Slowdown |
|---|---|---|---|
| DeepVariant | 3-5 hr | 9-25 hr | 3-5x |
| minimap2 alignment | 1-2 hr | 3-6 hr | 3x |
| VEP annotation | 2-4 hr | 4-8 hr | 2x |
| Manta | 20 min | 1-2 hr | 3-4x |
| CNVpytor | 1-3 hr | 3-8 hr | 3x |
| bcftools steps | 1-5 min | 2-10 min | 2x |
Mitigation strategies:
1. Reduce parallelism. Run one heavy step at a time instead of many in parallel.
2. Reduce resource limits. Use --cpus 2 --memory 8g instead of --cpus 8 --memory 32g to avoid emulation thrashing.
3. Offload the heaviest steps. Run steps 2, 3, 18, 19 on a remote Linux machine and bring back the outputs. Everything after step 3 needs only the VCF or BAM.
4. Use a cloud instance. A Hetzner CCX33 (8 vCPU, 32 GB, ~$0.18/hr) will run the full pipeline in 6-10 hours for about $1-2.
WSL2 file I/O is extremely slow¶
Symptom: Steps that read large files (alignment, DeepVariant, VEP) take 10-50x longer than expected.
Cause: If your data is on a Windows drive (/mnt/c/, /mnt/d/), WSL2 accesses it through the 9P filesystem protocol, which is catastrophically slow for large random-access files.
Fix: Move ALL genomics data to the Linux filesystem:
# Inside WSL2
mkdir -p ~/genome_data
# Move or copy data from /mnt/c/ to ~/genome_data/ (replace <windows-username>)
cp -r "/mnt/c/Users/<windows-username>/genome_data/"* ~/genome_data/
export GENOME_DIR=~/genome_data
Data on the Linux filesystem (~/, /home/, /tmp/) uses ext4 natively and performs at full speed.
Monitoring container resources¶
Use docker stats to see real-time resource usage:
# Live dashboard (updates every second)
docker stats
# One-time snapshot
docker stats --no-stream
# Output columns:
# CONTAINER CPU% MEM USAGE/LIMIT NET I/O BLOCK I/O
What to look for:
- CPU at 0% for a long time: Container may be stuck or waiting for I/O
- MEM USAGE near LIMIT: Container is about to be OOM-killed. Increase --memory.
- NET I/O > 0: Container is downloading something (may be slow due to network)
- BLOCK I/O very high: I/O-bound step. Faster storage would help.
When to reduce --cpus or --memory¶
With run-all.sh you do not edit the scripts' flags: THREADS=N caps every task's CPUs (--max_cpus), and --max_memory 24.GB after the sex caps every task's memory. Nextflow starts a task only when its CPUs and memory fit in what the machine has free.
Reduce --cpus when:
- Running multiple steps in parallel on a machine with limited cores
- Running on macOS with Rosetta 2 emulation (more shards = more overhead)
- Your machine has fewer than 8 cores
Reduce --memory when:
- Docker Desktop memory allocation is less than the script's --memory flag
- Running multiple containers simultaneously
- System becomes unresponsive during analysis
Safe minimum values per step:
| Step | Min --cpus | Min --memory | Notes |
|---|---|---|---|
| 2 (minimap2) | 2 | 8g | Slower but works |
| 3 (DeepVariant) | 2 | 8g | Set --num_shards=2 to match |
| 4 (Manta) | 2 | 4g | |
| 6 (ClinVar) | 1 | 1g | Very light |
| 7 (PharmCAT) | 1 | 2g | |
| 9 (ExpansionHunter) | 2 | 2g | |
| 13 (VEP) | 2 | 4g | Set --fork 2 to match |
| 17 (CPSR) | 2 | 4g | |
| 18 (CNVpytor) | 4 | 8g | |
| 19 (Delly) | 2 | 4g |
Steps that can run in parallel vs must be sequential¶
SEQUENTIAL (must run in order):
Step 2 (alignment) ──> Step 3 (variant calling)
Step 4 (Manta) ──> Step 5 (AnnotSV)
Step 4 (Manta) ──> Step 15 (duphold)
PARALLEL after Step 3 completes (all independent):
┌─ Step 4 (Manta) ← needs BAM
├─ Step 6 (ClinVar) ← needs VCF
├─ Step 7 (PharmCAT) ← needs VCF
├─ Step 9 (ExpansionHunter) ← needs BAM
├─ Step 10 (TelomereHunter) ← needs BAM
├─ Step 11 (ROH) ← needs VCF
├─ Step 12 (haplogrep3) ← needs VCF
├─ Step 13 (VEP) ← needs VCF
├─ Step 14 (imputation) ← needs VCF
├─ Step 16 (indexcov) ← needs BAM index only
├─ Step 17 (CPSR) ← needs VCF
├─ Step 18 (CNVpytor) ← needs BAM
├─ Step 19 (Delly) ← needs BAM
└─ Step 20 (Mutect2 mito) ← needs BAM
RESOURCE CAUTION:
Steps 3, 18, 19 are the most CPU/RAM intensive.
Do not run more than 2 of these simultaneously on a 32 GB machine.
Output Issues¶
0-byte output files¶
Symptom: Output file exists but is 0 bytes.
Most common cause: The input path was wrong inside the Docker container. Docker mounts the host path to /genome, so all paths inside the container must start with /genome/, not the host path.
Diagnosis checklist: 1. Check that the input file exists at the expected path:
2. Run the script with debug output: 3. Inspect Docker mount mapping. The script mounts${GENOME_DIR}:/genome. Inside the container, ${GENOME_DIR}/sample/vcf/sample.vcf.gz becomes /genome/sample/vcf/sample.vcf.gz.
4. Check for bgzip/tabix path issues. The bcftools image (BCFTOOLS_IMAGE) does not include bgzip or tabix in $PATH. Use bcftools view -Oz -o instead of piping to bgzip. See lessons-learned.md.
Other causes:
- Tool crashed silently before writing output (check exit code and stderr)
- Disk full (check df -h)
- Permission denied in output directory (add --user root)
Wrong number of variants (too few)¶
Symptom: DeepVariant VCF has far fewer variants than expected.
Expected variant counts for 30X WGS (GRCh38):
| Metric | Expected Range | Notes |
|---|---|---|
| Total variants (SNP + indel) | 4.5-5.5 million | Total rows in VCF |
| PASS variants | 4.0-4.8 million | After DeepVariant quality filtering |
| SNPs | 3.8-4.5 million | ~85% of total |
| Indels | 600K-900K | ~15% of total |
| Heterozygous | 2.5-3.2 million | ~55-60% of total |
| Homozygous ALT | 1.8-2.3 million | ~40-45% of total |
If total variants < 3 million: 1. Genome build mismatch. Your BAM may be hg19, but the reference is hg38. See Wrong genome build. 2. Low coverage. Check average depth:
source versions.env # from the repository root
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
samtools depth -a /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam | \
awk '{sum+=$3; n++} END {print "Average depth:", sum/n}'
source versions.env # from the repository root
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
samtools flagstat /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam
If total variants > 7 million: 1. Likely includes many false positives. Filter to PASS only:
source versions.env # from the repository root
docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools view -f PASS /genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz | \
bcftools stats | grep "number of records"
source versions.env # from the repository root
REF_FASTA=reference/GRCh38_no_alt_analysis_set.fasta # see 00-reference-setup.md#the-reference-path-on-every-page
docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools norm -m -both \
-f "/genome/${REF_FASTA}" \
/genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz | bcftools stats | grep "number of records"
Wrong number of variants (too many ClinVar hits)¶
Symptom: ClinVar screen returns hundreds of "pathogenic" hits, which seems too many.
Cause: The ClinVar pathogenic filter includes CLNSIG~"Pathogenic", which matches variants where pathogenicity depends on context. Some variants are "Pathogenic" for one condition but "Benign" for another (compound CLNSIG fields like Pathogenic/Likely_benign).
Expected ClinVar hit counts:
| Category | Expected Count |
|---|---|
| Total ClinVar matches (any significance) | 50-150 |
| Pathogenic + Likely pathogenic | 0-10 |
| True actionable findings | 0-3 |
If you see more than ~15 pathogenic hits: The ClinVar VCF may not be filtered correctly, or chromosome naming may be mismatched (causing position-only matches without allele verification). Re-run the ClinVar screen from step 0 setup.
How to verify outputs are correct¶
Quick sanity checks for each major output:
source versions.env # from the repository root
SAMPLE=your_name
# VCF: check variant count and type distribution
docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools stats /genome/${SAMPLE}/vcf/${SAMPLE}.vcf.gz | grep "^SN"
# BAM: check alignment rate and depth
docker run --rm -v ${GENOME_DIR}:/genome "${SAMTOOLS_IMAGE}" \
samtools flagstat /genome/${SAMPLE}/aligned/${SAMPLE}_sorted.bam
# Manta SV count (expect 7,000-9,000 total)
docker run --rm -v ${GENOME_DIR}:/genome "${BCFTOOLS_IMAGE}" \
bcftools view /genome/${SAMPLE}/manta/results/variants/diploidSV.vcf.gz | grep -c -v "^#"
# PharmCAT: check the HTML report exists and has content
ls -la ${GENOME_DIR}/${SAMPLE}/vcf/${SAMPLE}.report.html
# Should be > 50 KB
# CPSR: check the HTML report exists
ls -la ${GENOME_DIR}/${SAMPLE}/cpsr/${SAMPLE}.cpsr.grch38.html
# Should be > 100 KB
# VEP: check annotated VCF exists and has annotations
head -50 ${GENOME_DIR}/${SAMPLE}/vep/${SAMPLE}_vep.vcf | grep "CSQ="
# Should see consequence annotations
Expected output sizes per step¶
| Step | Output File(s) | Expected Size |
|---|---|---|
| 2 | *_sorted.bam + .bai |
80-120 GB + 5-8 MB |
| 3 | *.vcf.gz + .tbi |
80-200 MB + 1-2 MB |
| 4 | diploidSV.vcf.gz |
1-5 MB |
| 5 | *_sv_annotated.tsv |
25-35 MB |
| 6 | *_clinvar_hits.vcf |
< 1 MB |
| 7 | *.report.html |
50-200 KB |
| 9 | *_eh.json + *_eh.vcf |
< 1 MB each |
| 10 | *_summary.tsv + plots |
50-200 MB total |
| 11 | *_roh.txt |
1-5 MB |
| 12 | *_haplogroup.txt |
< 1 KB |
| 13 | *_vep.vcf (uncompressed) |
2-5 GB |
| 17 | *.cpsr.grch38.html + TSV |
50-200 MB total |
| 18 | *_cnvs.txt + .root |
Text: < 1 MB, ROOT: 5-15 GB |
| 19 | *_sv.vcf.gz |
5-20 MB |
| 20 | *_chrM_filtered.vcf.gz |
< 1 MB |
If any output is significantly smaller than expected (especially 0 bytes), see 0-byte output files.
Getting Help¶
Before opening an issue¶
- Check this guide first. Ctrl+F your error message.
- Check docs/lessons-learned.md for known Docker image issues and tool quirks.
- Run with debug output:
bash -x scripts/XX-stepname.sh your_name 2>&1 | tee debug.log - Collect diagnostic info:
Opening a GitHub issue¶
File an issue at: github.com/GeiserX/Personal-Genome-Pipeline/issues
Include:
- Which step failed (step number and script name)
- Full error message (copy-paste, not screenshot)
- Your platform (Linux distro, macOS version, WSL2 version)
- Docker version (docker version)
- Available RAM and disk space
- Input data type (FASTQ, BAM, or VCF) and approximate size
- Whether you are using native Linux or Docker Desktop
Useful diagnostic commands¶
# System info
uname -a # OS and architecture
docker version # Docker version
docker info | head -30 # Docker config summary
docker system df # Docker disk usage
# Container debugging
docker ps -a --latest # Last container status
docker logs <container_id> # Container output
docker inspect <container_id> # Full container metadata
# File validation
samtools quickcheck file.bam # BAM integrity (via Docker)
gzip -t file.vcf.gz # Gzip integrity
bcftools view -h file.vcf.gz # VCF header readability
# Resource monitoring
docker stats --no-stream # Current container resource usage
df -h # Disk space
free -h # RAM (Linux only)