AlphaGenome Atlas

RawGraph

AlphaGenome Atlas is a precomputed catalog of molecular-effect predictions for genetic variants, released by Google DeepMind on September 8, 2026. It applies the AlphaGenome sequence-to-function model across the GRCh38 human reference genome and includes scores for about 9 billion single-nucleotide variants (SNVs). Google describes the full resource as a 1-petabyte dataset. The Atlas is a prediction database and access platform, not a new foundation model or a collection of experimental measurements.[1][2]

The 9 billion figure represents the three possible alternative A, C, G, or T bases at each usable reference position. It does not include every possible indel, structural variant, combination of variants, or personal genome. The research pipeline also scored more than 100 million observed insertions and deletions, but the preprint's data-availability statement said the public Atlas datasets at launch were limited to valid genome-wide SNVs, with a possible expansion to indels after final publication.[2]

The release also introduced AlphaGenome Variant Impact (AVI), a genome-wide ranking score that combines AlphaGenome predictions with AlphaMissense, coding-consequence indicators, and evolutionary conservation. A linked feature-attribution layer shows which inputs contribute to each AVI score, and a motif compendium maps recurrent regulatory sequence patterns predicted by AlphaGenome.[1][2]

FieldDetail
DeveloperGoogle DeepMind[1]
ReleasedSeptember 8, 2026[1]
Reference assemblyGRCh38, also called hg38[2]
Gene annotationGENCODE v46[2]
Main coverageAbout 9 billion possible SNVs at non-N reference bases[2]
Reported full scale1 petabyte[1]
Predictions per variantAbout 27,000 experiment-specific scalar scores on average, or about 15,000 without active-allele scores[2]
Main interfacesWeb portal, Python API, static downloads for selected data, and an AlphaGenome skill in Google Antigravity[1][2]
Launch status of evidenceAtlas technical report was a preprint; the underlying AlphaGenome model had a peer-reviewed Nature paper[2][3]
Intended useMolecular prediction and research, not diagnosis or medical decision-making[1][4]

Relationship to AlphaGenome

AlphaGenome Atlas is built from AlphaGenome rather than replacing it. The underlying deep learning model takes as much as 1 megabase of DNA sequence and predicts thousands of functional-genomics tracks across gene expression, transcription initiation, chromatin accessibility, transcription-factor binding, histone modifications, splicing, and chromatin contacts. Its peer-reviewed paper reports 5,930 human and 1,128 mouse tracks across 11 broad modalities.[3]

AlphaGenome normally computes predictions when a researcher submits a sequence interval or variant. The Atlas precomputes and indexes scalar variant-effect scores across the whole human reference genome. This makes it better suited to lookup and ranking across large SNV sets, while the underlying model remains useful when a researcher needs full reference-versus-alternate tracks for selected regions or variants. The official API documentation says on-demand prediction is intended for limited regions and thousands of predictions, and is likely unsuitable for analyses requiring more than one million predictions.[4]

The Atlas was generated with AlphaGenome's distilled model, which is the smaller inference model used for variant scoring in the original study. Model implementation code and weights are available separately through the AlphaGenome research repository and Hugging Face under AlphaGenome's model terms. Those artifacts are distinct from the Atlas dataset and its access terms.[3][5][6]

How the Atlas was constructed

The authors divided GRCh38 into non-overlapping 128-base-pair windows. At every reference position that was not marked N, they enumerated the three alternative DNA bases. Each variant was placed in a 1-megabase AlphaGenome input centered on its 128-base-pair window. The reference prediction was cached for variants in the same window, and annotation masks were reused where possible. The team reported that positioning a variant by window rather than exactly at the interval center did not produce a statistically significant change in their re-run of the original AlphaGenome evaluations.[2]

Insertions and deletions require a different procedure because changing sequence length can misalign downstream representations in a convolutional model. The preprint describes a shift-augmentation and representation-stitching method intended to reduce those artifacts. The indel set was the union of variants observed in gnomAD v4.1, UK Biobank, All of Us, and the study's evaluation datasets. It was not an enumeration of every possible insertion or deletion.[2]

For each variant, the Atlas reduces reference-versus-alternate model outputs to scalar values using assay-specific rules. For example, accessibility and transcription-factor binding scores summarize changes in a local window, while gene-expression and splicing scores use gene annotations and predicted transcript structure. The report separates 12 scorer families derived from AlphaGenome's 11 broader output modalities.[2][3]

Atlas scorer familyPredicted molecular signal
DNase and ATACChromatin accessibility
TF ChIPTranscription-factor binding
Histone ChIPHistone modifications
CAGE and PRO-capTranscription initiation
RNA-seqGene-expression coverage
PolyadenylationChanges in polyadenylation-site use
Splice sitesDonor and acceptor changes
Splice-site usageRelative use of predicted splice sites
Splice junctionsJunction coordinates and strength
Contact mapsPredicted three-dimensional chromatin contacts

The reported average of about 27,000 scores per variant includes active-allele scores, which describe the larger predicted activity of the reference or alternate allele rather than their difference. Active-allele values were computed for ATAC, DNase, TF ChIP, histone ChIP, CAGE, PRO-cap, and RNA-seq, but not for polyadenylation, the three splicing families, or contact maps.[2]

Atlas records are indexed by chromosome, position, and alternate allele, with the reference allele implied by GRCh38. A query using a different reference allele, swapping the reference and alternate alleles, or asking for a non-reference-to-non-reference substitution will not match the precomputed SNV record.[2]

AlphaGenome Variant Impact score

AVI condenses 18 inputs into a single ranking score. Ten inputs are maximum effects aggregated across tissues for AlphaGenome modalities: DNase, ATAC, TF ChIP, histone ChIP, CAGE, PRO-cap, RNA-seq, polyadenylation, merged splicing, and contact maps. Four are coding features: AlphaMissense plus indicators for protein termination or frameshift, start loss, and stop loss. Two are conservation measures, PhastCons 470-way and Zoonomia Cactus 241-way. The remaining two mark insertions and deletions.[2]

The model was trained on variants from gnomAD v4.1 using allele frequency as a proxy target. Variants with a group-maximum filtering allele-frequency lower bound of at least 0.001 and below 0.999 were labeled proxy neutral, while variants below 0.001 were labeled proxy impactful. These labels reflect population rarity, not experimentally established pathogenicity. AVI therefore ranks variants for follow-up; it does not estimate the probability that a variant causes a disease.[2]

Raw AVI predictions are converted to PHRED-scaled genome-wide ranks. An AVI PHRED score of 10 marks the top 10 percent of ranked predictions, 20 marks the top 1 percent, and 30 marks the top 0.1 percent. The values should not be read as probabilities or compared with an assay-specific AlphaGenome score as if they used the same scale.[2]

AVI also provides SHAP feature attributions. The 18 signed contributions sum to the raw AVI score and indicate whether an input pushes that score up or down. If splicing is the main positive contributor, for example, a researcher can inspect the Atlas's precomputed splicing scores and then request full AlphaGenome tracks for relevant cell types. Conservation can improve ranking performance but may dominate an attribution without identifying a specific molecular mechanism.[2]

Regulatory motif maps

The motif resource was derived from genome-wide in silico mutagenesis contributions rather than copied from an existing motif database. The authors ran TF-MoDISco in selected genomic regions, producing roughly 900 million patterns, and used MotifCompendium to filter and cluster them into 2,601 motifs. Manual curation assigned broad labels covering 94 main transcription factors and 122 zinc-finger transcription factors. The analysis also reported 464 composite patterns representing 104 motif combinations.[2]

The team used FiNeMO to map predicted motif instances across the genome. Comparisons with JASPAR and ENCODE GRAMMAR found similarities to known motifs. ChIP-nexus experiments in HepG2 cells supported mapped instances for three tested transcription factors, and a CTCF analysis tested cell-type specificity. These experiments validate selected examples, not every cluster or genome-wide motif instance. The report cautions that motif calls remain dependent on AlphaGenome's fidelity and do not by themselves establish causal binding in living cells.[2]

Evaluation

The Atlas report evaluated AVI on clinical databases, fine-mapping and trait benchmarks, and saturation genome-editing screens. The results below are from the launch preprint and had not been independently reproduced for this article. AUPRC depends on the benchmark's case mix and negative set, so values from different rows are not directly interchangeable.[2]

EvaluationAVI AUPRCReported comparison
ClinVar intronic SNVs0.76GPN-Star-V, 0.44
ClinVar synonymous SNVs0.57CADD v1.7, 0.35
ClinVar 3' UTR SNVs0.50GPN-Star-M, 0.18
ClinVar 5' UTR SNVs0.26GPN-Star-V, 0.27
TraitGym non-coding Mendelian variants0.76GPN-Star-M, 0.77
Ten pooled saturation genome-editing screens, SNVs only0.668CADD v1.7, 0.647
The same screens with indels0.759CADD v1.7, 0.653

AVI led the reported intronic, synonymous, and 3' UTR ClinVar comparisons, but it did not lead every evaluation. GPN-Star-V was slightly higher on 5' UTR ClinVar SNVs, and GPN-Star-M was slightly higher on the TraitGym Mendelian set. Across ten held-out saturation genome-editing screens, AVI had the highest reported Spearman correlation in eight.[2]

These Atlas evaluations are separate from the earlier peer-reviewed evaluation of AlphaGenome itself. The Nature paper reported that AlphaGenome matched or exceeded the strongest external model on 25 of 26 variant-effect evaluations and 22 of 24 genome-track evaluations. That result supports the underlying model, but it should not be treated as independent validation of AVI, its training labels, or the Atlas access platform.[3]

Research case studies

DNM1 rare-disease analysis

In a retrospective analysis of 112 likely pathogenic or pathogenic variants from solved GREGoR cases, the preprint reports that AVI placed 29.5 percent within the top 50 variants per person, compared with 12.5 percent for CADD v1.7. Restricting candidates with a gnomAD allele-frequency threshold of 0.001 raised those figures to 74.3 percent and 61 percent. The team then ranked small de novo variants in 814 unsolved individuals with parent and proband genomes.[2]

For one proband with epileptic encephalopathy, the top AVI candidate was the deep-intronic variant chr9:128225994:G>A in DNM1. Its AVI PHRED score was 24.7, with 69 percent of the attribution assigned to splicing. AlphaGenome predicted a brain-specific cryptic splice acceptor and a 13-amino-acid in-frame extension of exon 10a. A minigene screen mutating 265 nucleotides upstream of the exon tested 796 variants across five cell lines and found 12 variants that caused in-frame extensions, including the candidate and previously reported nearby variants.[2]

The combined genetic, literature, prediction, and experimental evidence led the study team to recommend a Likely Pathogenic classification for the DNM1 variant. The report did not present AVI alone as sufficient for that classification. This distinction is consistent with the product's explicit statement that Atlas predictions are not validated or approved for clinical use.[1][2]

Population association analysis

The preprint also used Atlas scores to filter rare non-coding variants in analyses of 2,028 circulating proteins from 54,189 UK Biobank participants. Compared with regional annotation approaches without Atlas filtering, the authors reported a 22 percent increase in conditionally independent rare non-coding variant-aggregate discoveries. The filtering was intended to group variants predicted to affect the same gene through a more consistent molecular mechanism.[2]

Replication was limited. Of 31 reported associations, 25 had enough variants for evaluation in the All of Us version 8 whole-genome cohort. Four were nominally significant at P below 0.05, and none passed Bonferroni correction. The UK Biobank analysis also used a slightly older Atlas release because of temporary platform downtime. These details make the result a useful research demonstration, not settled evidence that Atlas filtering improves association studies across cohorts or traits.[2]

Access and licensing

At launch, researchers could use a no-code web portal, request precomputed values through the AlphaGenome Python API, or obtain selected static archives. Access rights differ by artifact. The technical report lists the static AVI SNV download as permissive for commercial and non-commercial use, while API access to AVI, splicing data, AVI feature attributions, and raw Atlas features is non-commercial. Raw Atlas features were API-only in that table.[2][4][7]

The official launch catalog listed an 88.5 GB compressed AVI SNV archive, a 20.6 GB merged-splicing SNV archive, and a 283.9 GB AVI feature-importance archive. Those files are subsets or derived products and should not be confused with the full 1-petabyte resource. The report said bulk motif downloads were planned, while motif browsing was available through the portal.[2][7]

The API repository is Apache-2.0 software, and its examples and documentation are CC BY 4.0. Those licenses do not grant equivalent rights to model outputs or Atlas data. Except for artifacts expressly marked as permissive, the official terms restrict AlphaGenome outputs and Atlas information to non-commercial use and prohibit using them to train other machine-learning models. The separately released AlphaGenome weights also use non-commercial model terms. Google offers commercial access to the base AlphaGenome model through Google Cloud, and the Atlas launch article said commercial Atlas access on Google Cloud was planned rather than already generally available.[1][4][5][6]

Google Antigravity integration

Google's launch article describes AlphaGenome Atlas as available through an AlphaGenome skill in Google Antigravity. The linked Science Skills collection includes a public alphagenome-single-variant-analysis skill that calls the AlphaGenome API and provides a workflow for examining gene expression, splicing, chromatin accessibility, histone marks, and transcription-factor predictions for a specified variant. It requires an AlphaGenome API key and applies the same AlphaGenome terms as direct API use.[1][8][9][10]

The skill is an agent workflow around AlphaGenome and Atlas access, not a separate genomics model or an independent validator. The Science Skills repository can also be installed outside Antigravity, but the repository states that its software and instructions have their own Apache-2.0 and CC BY licenses while third-party data sources retain separate terms.[9][10]

Limitations

  • Model predictions are not measurements. Atlas entries are AlphaGenome counterfactual predictions for reference and alternate sequences. Laboratory assays and independent data are still needed to establish molecular effects.[2][11]
  • Reference-genome scope. The launch SNV set is indexed against GRCh38 and does not account for a person's complete haplotype, other nearby or distant variants, or gene-environment interactions. Independent researchers quoted by Nature emphasized that personal genomic context can change a focal variant's effect.[2][11]
  • Training-data gaps. Standard poly(A)-selected RNA-seq omits non-polyadenylated RNAs, many relevant cell types are absent, and assay coverage is uneven across biosamples.[2]
  • Cis rather than trans regulation. AlphaGenome learns sequence-linked cis-regulatory effects but does not directly model trans-acting changes such as altered transcription-factor expression.[2]
  • AVI target and calibration. AVI uses population rarity as a proxy during training. Its PHRED value is a rank, not a calibrated pathogenicity probability, and the preprint calls for calibration standards and evaluation across diverse genetic ancestries.[2]
  • Interpretation can be incomplete. Conservation can dominate AVI attributions without specifying a molecular mechanism. Motif maps are sensitive to model quality, discovery parameters, and broad cell-type-agnostic labels, and predicted instances need not be causal in vivo.[2]
  • Early evidence. The Atlas report was a launch preprint. Its benchmark and application results came from the developing team and collaborators, with weak replication in the population-association example and experimental validation covering only selected variants and motifs.[2][11]

Google DeepMind states that AlphaGenome has not been validated for, and is not approved for, clinical use. Atlas and AVI should be used as research prioritization and interpretation aids, not as substitutes for diagnostic procedures, professional medical advice, or treatment decisions.[1][4]

References

  1. ^Google DeepMind. "AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome." September 8, 2026. deepmind.google/...tter-change-in-the-human-genome
  2. ^Jun Cheng, Kyle R. Taylor, Lauren Nicolaisen, et al. "AlphaGenome Atlas: in silico mutagenesis of the entire human genome improves prioritization and interpretation of non-coding variants." Technical report and preprint. September 2026. storage.googleapis.com/...alphagenome-atlas.pdf
  3. ^Ziga Avsec, Natasha Latysheva, Jun Cheng, et al. "Advancing regulatory variant effect prediction with AlphaGenome." Nature 649, 1206-1218. January 28, 2026. doi.org/...s41586-025-10014-0
  4. ^Google DeepMind. "AlphaGenome API." GitHub repository and documentation. Accessed September 9, 2026. github.com/...alphagenome
  5. ^Google DeepMind. "AlphaGenome Research." Model implementation, data loaders, and research code. Accessed September 9, 2026. github.com/...alphagenome_research
  6. ^Google. "AlphaGenome all folds." Hugging Face model card and weights. Accessed September 9, 2026. huggingface.co/...alphagenome-all-folds
  7. ^Google DeepMind. "AlphaGenome terms and downloads." Accessed September 9, 2026. deepmind.google.com/...terms
  8. ^Google Antigravity. "Why choose Google Antigravity for science." Accessed September 9, 2026. antigravity.google/...science
  9. ^Google DeepMind. "Science Skills." GitHub repository. Accessed September 9, 2026. github.com/...science-skills
  10. ^Google DeepMind. "AlphaGenome single variant analysis skill." Accessed September 9, 2026. github.com/...SKILL.md
  11. ^Ewen Callaway. "DeepMind's new genome 'atlas' charts effects of all nine billion human gene mutations." Nature. September 9, 2026. nature.com/...d41586-026-02835-4

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 2,805 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independently checked against cited and current primary sources on 2026-09-09.

Cite this page: AI Wiki. "AlphaGenome Atlas." aiwiki.ai, updated 9 Sept 2026, fact-checked 9 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/alphagenome_atlas

Suggest edit