Step 26: Ancestry (projection onto a reference panel)¶
What This Does¶
Places your genome among the samples of a reference panel whose populations are known, and names the panel population your genetic ancestry is most similar to. It uses pgsc_calc's ancestry projection: the panel's principal components are computed from the panel's own samples, and your genotypes are projected onto them, so one sample is enough. The same pgsc_calc run gives step 25 its percentiles.
Without the panel the step prints one line and exits 0:
Step 26 skipped: no ancestry reference panel at <genome_dir>/reference/pgsc_calc/pgsc_1000G_v1.tar.zst; install it with scripts/setup.sh --ancestry-panel <genome_dir>
Why¶
- PRS interpretation: a polygenic score (step 25) only means something next to people of similar genetic ancestry. The population found here is the group step 25 compares your score with.
- Context for other results: population frequencies and some risk estimates differ between ancestries.
A principal component analysis of one genome alone cannot work, since the axes come from the variation between many people. Projection onto axes that a reference panel defines is the method that works on one sample.
Tool¶
- pgsc_calc (
PGSC_CALC_VERSIONinversions.env), run by step 25: FRAPOSA's online augmentation, decomposition and Procrustes projection (--projection_method oadp), then a random forest on the first principal components to assign the most similar population. - Reference panel: pgsc_calc's 1000 Genomes database
pgsc_1000G_v1(PGSC_PANELinversions.env), 3,202 samples (2,583 founders) of the five 1000 Genomes super-populations (AFR, AMR, EAS, EUR, SAS), published by the PGS Catalog.
Docker Images¶
The images of step 25, all pinned in versions.env: PGSC_UTILS_IMAGE, PLINK2_IMAGE, PGSC_FRAPOSA_IMAGE, PGSC_ZSTD_IMAGE, PGSC_PYYAML_IMAGE, PGSC_REPORT_IMAGE, with PYTHON_IMAGE and BCFTOOLS_IMAGE. Image versions lists the tags.
Input¶
- VCF from DeepVariant (step 3), and its gVCF beside it: the panel's SNVs are genotyped from the gVCF, so the sites where you match the reference count in the projection. Without a gVCF only your variant sites are projected, which weakens it.
- The panel and its site list, installed once (7.4 GB, and about 23 GB of disk while pgsc_calc runs):
./scripts/setup.sh --ancestry-panel <genome_dir>. It is opt-in for that reason; see the measured numbers in step 25. - Java 17+ and Nextflow, as for step 25.
Command¶
run-all.sh runs it inside the pipeline's PRS run (step 25) whenever the panel and a PRS scoring file (prs_scores/, step 25) are installed, so a plain run includes it; without a scoring file, PRS is skipped and step 26 is listed as skipped (needs PRS):
ANCESTRY_PANEL=/path/to/panel.tar.zst points at another pgsc_calc panel (its _GRCh38_sites.tsv must be beside it); ANCESTRY_PANEL=none skips the step.
What the Script Does Internally¶
- Without the panel: prints the one line above and exits 0.
- Runs step 25, which uses the panel when it is installed: the score and panel positions are genotyped from the gVCF, and pgsc_calc runs with
--run_ancestry. pgsc_calc intersects your genotypes with the panel, keeps unrelated panel samples and common, LD-thinned SNVs, computes the panel's principal components, projects you onto them, and assigns the population with a random forest trained on the panel's labels. - Prints the population and checks that the ancestry table was written.
In the Nextflow pipeline the same happens inside prs when --ancestry_ref is set; ancestry in --tools needs prs and --ancestry_ref.
Output¶
| File | Contents |
|---|---|
${SAMPLE}_ancestry.tsv |
key and value rows: sample, reference_panel, population (the most similar panel population), population_low_confidence, probability_<POP> for each panel population, and PC1 to PC10 |
Written to ${GENOME_DIR}/${SAMPLE}/ancestry/. pgsc_calc's own files, including the principal components of every panel sample, are in ${GENOME_DIR}/${SAMPLE}/prs/pgsc_calc/results/sample/score/ (sample_popsimilarity.txt.gz).
Runtime¶
The panel's extraction, QC and PCA come on top of step 25's run; see the measured numbers in step 25.
Interpreting Results¶
- population is the reference group whose genetic ancestry is most similar to yours. It is a statement about similarity to five continental groups of the 1000 Genomes Project, not about identity, nationality or ethnicity.
- probability_
are the random forest's probabilities. A low top probability ( population_low_confidenceisTrue) means you sit between groups or far from all of them; the step 25 percentile is then less reliable. - PC1 to PC10 are your coordinates on the panel's principal components. They can be plotted against the panel samples in
sample_popsimilarity.txt.gz.
Limitations¶
- Five continental groups only. Mixed ancestry is shown as the single most similar group, with its probabilities; this step does not estimate admixture fractions or local ancestry.
- Groups the 1000 Genomes Project does not sample well are placed less reliably.
- Fine-grained ancestry (for example Spanish versus Italian) needs other panels and tools.
Notes¶
- The panel is one file kept as downloaded; pgsc_calc unpacks the GRCh38 part into its work folder on each run. The site list beside it (
pgsc_1000G_v1_GRCh38_sites.tsv) is made bysetup.shwith plink2 from the panel's own genotypes: the biallelic autosomal SNVs with a panel frequency of 5% or more, the threshold pgsc_calc's projection applies (maf_ref): 6,966,553 of the 1000 Genomes panel's 61.6 million. - pgsc_calc's synthetic HAPNEST panel (
GRCh38_HAPNEST_reference, 268 MB) is what the pull-request tests project onto;ANCESTRY_PANEL_NAME=GRCh38_HAPNEST_reference ./scripts/setup.sh --ancestry-panel <genome_dir>installs it, but its populations are simulated and say nothing about you.