How GenomeNarrator Analyzes Your Genome
GenomeNarrator combines clinical databases, population genetics, and pharmacogenomics guidelines to turn raw DNA data into evidence-graded insights. No AI or large language model generates any finding — every result is a deterministic match against peer-reviewed data. Key figures: 248,544 rsids indexed against gnomAD v4.1, 41 pharmacogenes across 110 CPIC drug-gene pairs, all 81 ACMG SF v3.2 genes, and a population reference cohort of 807,000+ individuals.
This page describes that methodology in full technical detail for anyone who wants it. No biology or genetics background is needed to read your actual report — every finding is explained in plain English first, with this level of detail available underneath for verification.
The Analysis Pipeline — Seven Stages
- Detect & normalize — Raw genome files (VCF — including whole-genome/whole-exome sequencing VCFs up to 2 GB, streamed rather than loaded whole — 23andMe, AncestryDNA) are parsed and auto-detected as GRCh37 or GRCh38, normalizing raw genotype calls to a single build, then matched against rsID lookups spanning 248,544 positions in a curated gnomAD v4.1 index.
- Cross-reference — Normalized variants are checked against SNPedia, the NHGRI-EBI GWAS Catalog, and a 353,242-variant ClinVar index, using a two-pass matching algorithm with curated overrides.
- Dedupe & gate — Matches are deduplicated and LD-pruned so linked variants aren't double-counted, then screened through a rare-Mendelian gate separating high-impact single-gene findings from common polygenic background.
- Classify — Every finding gets an ACMG/AMP evidence tier, including PP3/BP4 in-silico evidence from AlphaMissense's missense pathogenicity predictions. All 81 ACMG SF v3.2 actionable genes are screened across 18 clinical categories. Carrier status and compound-heterozygous candidates are flagged, with a note that SNP array data can't confirm cis/trans phase.
- Pharmacogenomics — CPIC-guideline diplotypes are called across 37 pharmacogenes (CYP2D6, CYP2C19, TPMT, DPYD, and others) plus 4 HLA hypersensitivity risk alleles — 41 genes and 110 drug-gene pairs total.
- Population-match — Polygenic risk scores from multi-SNP models are Bayesian-adjusted against gnomAD v4.1 allele frequencies (807,000+ individuals), bias-corrected to your auto-inferred ancestry.
- Score & report — Bayesian absolute risk is calculated across 63 diseases with a full evidence trail to source studies, calibrated against whichever of eight gnomAD super-populations is inferred from your genotypes.
Ancestry Inference — Composition, Chromosome Painting, and Haplogroups
Ancestry is auto-detected from ancestry-informative markers (AIMs) — no manual population picker is shown. Global ancestry composition is inferred across all eight gnomAD super-populations (African, Amerindigenous, Ashkenazi Jewish, East Asian, European, Finnish, Middle Eastern, South Asian) with a confidence score and 95% bootstrap confidence interval per population, plus a finer regional sub-population call (e.g. Han Chinese vs. Japanese within East Asian) when marker coverage supports it.
A second, independent model — local ancestry inference, shown as Chromosome Painting — paints ancestry per-segment along each chromosome using a hidden Markov model with Viterbi decoding, rather than a single genome-wide average. It's validated by seeded simulation against real 1000 Genomes reference haplotypes (not synthetic data): mean per-window classification accuracy runs 89–94% for high-divergence population pairs (e.g. African-ancestry vs. others) and 70–84% for closely related pairs (e.g. East Asian/European/Finnish), which tracks real population genetics — closely related populations are inherently harder to distinguish per-window, not a modeling shortfall. Run npm run validate:lai to reproduce this validation against the production code directly.
Maternal (mtDNA) and paternal (Y-DNA) haplogroups are called separately from a phylogenetic marker panel — a PhyloTree Build 17 mtDNA reference and 32 curated ISOGG Y-chromosome-defining SNPs — each with its own confidence score, ancestral lineage path, and geographic migration path.
Ancestry composition also drives targeted screening: 8 ancestry-specific founder-mutation panels (Ashkenazi Jewish, Finnish, African, South Asian, Middle Eastern/North African, East/Southeast Asian, European, and Americas/Latino), covering 56 curated founder variants, are automatically screened once your inferred ancestry for that population reaches 5% or higher.
Data Sources
Every annotation traces back to a publicly available, peer-reviewed database:
- gnomAD v4.1 — 730,947 exomes + 76,215 genomes; the largest population allele-frequency reference available. gnomad.broadinstitute.org
- ClinVar (NCBI) — NIH's curated archive of variant-disease relationships, full monthly release with our own quality filters. ncbi.nlm.nih.gov/clinvar
- NHGRI-EBI GWAS Catalog — ~700K variant-trait associations from ~6,000 published studies, rebuilt monthly. ebi.ac.uk/gwas
- PGS Catalog — polygenic score models converted to population percentiles per ancestry. pgscatalog.org
- CPIC Guidelines — peer-reviewed pharmacogenomics guidelines used in hospital systems worldwide. cpicpgx.org
- ACMG SF v3.2 — 81 actionable genes recommended for return of secondary findings. acmg.net
- ClinGen — the NIH-funded Clinical Genome Resource's gene-disease actionability curation, used to attach surveillance, preventive, and treatment guidance to every actionable ACMG SF finding. clinicalgenome.org
- SNPedia — community-curated SNP knowledge base for trait and phenotype annotation. snpedia.com
- AlphaMissense (Cheng et al., Google DeepMind) — 71M precomputed missense-variant pathogenicity predictions, joined against our ClinVar index for PP3/BP4 in-silico evidence. github.com/google-deepmind/alphamissense
- Reactome — curated human biological pathways, shown for informational context only, never as scoring input. reactome.org
- GTEx — population-level gene expression across human tissues, shown for context only, never as scoring input. gtexportal.org
Evidence Quality Tiers
Every condition in your report is assigned an evidence tier based on the strength and source of the underlying research:
- Tier 4 — Clinical Grade (ClinVar P/LP): Variant classified Pathogenic or Likely Pathogenic by ClinVar, the NIH gold standard.
- Tier 3 — Strong Evidence (replicated risk studies): Genotype confirmed in SNPedia or peer-reviewed risk-factor literature with a known effect size, replicated across independent cohorts.
- Tier 2 — Moderate Evidence (GWAS / VUS): A GWAS signal or variant of uncertain significance — statistically associated but with modest effect size or incomplete replication.
- Tier 1 — Limited Evidence (general annotation): Known annotation with incomplete genotype-specific evidence, or based on a single study.
- Tier 0 — Reference Only (no personal variant): Condition is in our reference panel but no matching variant was detected — included for completeness; a negative result doesn't rule out rare or structural variants outside SNP-array coverage.
Known Limitations
- SNP arrays vs. whole-genome sequencing — consumer arrays genotype ~700K-1M pre-selected positions; rare variants (MAF < 0.1%) and off-panel positions aren't assessed. This limitation is specific to 23andMe/AncestryDNA-style array uploads — GenomeNarrator also accepts WGS/WES VCF files directly (up to 2 GB) for deeper coverage beyond the array panel.
- Structural variants not detected — large deletions, duplications, inversions, and CNVs are invisible to SNP arrays; for BRCA1/BRCA2, an array-negative result does not exclude a structural variant.
- PGx star-allele limitations — star alleles are inferred from tagged SNPs; complex CYP2D6 structural rearrangements may not be correctly assigned.
- Ancestry inference accuracy — global ancestry is auto-inferred from ancestry-informative markers across 8 gnomAD super-populations; admixed individuals may not map cleanly to a single population, and local-ancestry (chromosome painting) window accuracy is lower for closely related population pairs (70–84%) than for high-divergence pairs (89–94%).
- Population representation — most large-scale GWAS studies have historically over-represented European-ancestry populations, which can reduce PRS accuracy for other ancestries.
- Not a clinical diagnostic test — GenomeNarrator is a health-education tool, not an FDA-cleared diagnostic; findings require confirmation by a physician, genetic counselor, or clinical lab.
The Technical Ceiling of Array-Based Pharmacogenomics
Every consumer-genomics PGx report — ours included — is built on tag-SNP proxy genotyping, not direct haplotype sequencing. A microarray reads individual biallelic positions; it does not physically observe which alleles sit together on the same chromosome. Star-allele diplotypes are therefore inferred from marker SNPs known to co-segregate with a haplotype at high, but never perfect, linkage disequilibrium. This has three specific, well-documented failure points:
- Phase ambiguity in compound heterozygotes — array data cannot resolve whether two heterozygous marker SNPs sit in cis (one functional haplotype) or in trans (two disrupted alleles). GenomeNarrator defaults to the more conservative, higher-risk interpretation and marks the result unphased rather than assuming cis-favorable phase.
- Copy-number and structural variation invisible to biallelic calls — whole-gene deletions/duplications like CYP2D6*5 or CYP2D6*1×N read identically to a single wildtype copy on a genotyping array. A Normal or Intermediate Metabolizer call cannot, on array data alone, rule out a true Ultrarapid or Poor Metabolizer phenotype masked by an undetected structural variant. GenomeNarrator runs CYP2D6 structural-variant caveats unconditionally on every Normal/Intermediate call.
- Ancestry-dependent tag-SNP transferability — a tag SNP validated in a European-ancestry cohort can perform substantially worse in another ancestry group. Published validation cohorts (PMC8350439, HLA tag-SNP genotyping) report single-marker sensitivity for HLA-B*15:02 as low as 31.8% depending on tag SNP and population — a single-marker design can silently miss roughly two-thirds of true carriers of a CPIC Level A, Stevens-Johnson-syndrome-associated allele. GenomeNarrator combines multiple independently-validated tag SNPs per HLA allele under OR logic to raise sensitivity.
None of these three constraints can be engineered away at the software layer — they are physical limits of what a biallelic SNP array measures. A confirmatory clinical PGx panel remains the standard for any result that will change a prescribing decision.
Who's Behind This
GenomeNarrator is an independently built project, not a hospital lab or a venture-backed health startup. Every claim on this page cites its source (ClinVar, CPIC, the GWAS Catalog, PGS Catalog, peer-reviewed literature), every finding in your report links back to the record it came from, and the scoring logic runs against an automated regression suite before any change ships. ClinVar and the GWAS Catalog are rebuilt monthly, with gnomAD, PGS Catalog, and CPIC refreshed each major release, and nothing about your genome ever reaches a server — verify this yourself by opening your browser's network tab during an analysis. This is a tool for informing a conversation with a real clinician, not replacing one.
See pricing, start your genomic analysis, or read our GWAS guide and pharmacogenomics guide.