Quick Answer:
AlphaGenome is DeepMind's AI model that analyses DNA sequences to predict biological function. Using a hybrid convolutional and Transformer architecture, it processes up to 1 million base pairs at single-letter resolution, and DeepMind reports it beats prior specialised models on the large majority of the sequence-prediction and variant-effect tasks it was tested on. It is a research tool for scientists, not a consumer product.
Following AlphaFold's disruption of protein structure prediction, Google DeepMind has turned its attention to the source code of life itself. AlphaGenome is an AI model that analyses DNA sequences to predict their function, with DeepMind positioning it as a step towards better understanding of genetic disease and, eventually, more precise medicine.
It is not just about reading the genetic code — it is about predicting what that code does across different cell types and tissues, something the human genome's 98% non-coding portion has stubbornly resisted for decades.
DeepMind's own breakdown of how AlphaGenome maps gene regulation.
Decoding DNA Function
The human genome is vast, and while we have sequenced it in full, we still do not fully understand how the non-coding regions — roughly 98% of our DNA — influence gene expression. This matters more than it might sound: genome-wide association studies have repeatedly found that a large share of the genetic variants linked to common diseases sit in this non-coding "dark matter", far from any protein-coding gene, which makes them extremely difficult to interpret using traditional genetics alone. AlphaGenome is DeepMind's attempt to close that gap by predicting, directly from raw sequence, what a given stretch of DNA actually regulates.
Capable of analysing sequences up to 1 million base pairs long at single-letter resolution, AlphaGenome predicts several distinct layers of genomic activity from one model: where genes start and end across different cell types, RNA splicing patterns, overall RNA production levels, chromatin accessibility (how physically "open" a stretch of DNA is to regulatory proteins), and protein-DNA binding sites. Crucially, it can score the effect of a variant by comparing its predictions for the mutated sequence against the unmutated original, which is what makes it useful for interpreting a specific patient's mutation rather than only describing the reference genome in general.
This variant-scoring approach is what separates AlphaGenome from a purely descriptive annotation tool. A researcher can take a single-letter change found in a patient's genome, run both the reference and mutated sequence through the model, and read off a predicted difference across each modality — for example, whether the mutation is likely to reduce RNA production in a specific tissue, disrupt a splice junction, or open up a previously inaccessible stretch of chromatin to transcription factors. That per-modality, per-tissue granularity is what makes the output genuinely actionable for a lab trying to prioritise which of dozens of candidate variants to investigate experimentally.
From AlphaFold to AlphaGenome
AlphaGenome sits inside a specific lineage of DeepMind genomics and biology models, each covering a different slice of the same underlying problem:
- AlphaFold predicts the 3D structure a protein folds into from its amino acid sequence — a downstream question of what a gene's product looks like, not what switches the gene on.
- AlphaMissense focuses specifically on missense mutations within the roughly 2% of the genome that codes for proteins, scoring how likely a given amino-acid change is to be pathogenic.
- Enformer, an earlier DeepMind/Google Research genomics model, established transformer-based long-range sequence modelling for gene expression prediction. AlphaGenome builds directly on this foundation.
- AlphaGenome covers the remaining 98% of the genome — the non-coding, regulatory portion — where a large share of disease-associated variants identified by genome-wide association studies actually sit, extending well beyond what Enformer or AlphaMissense address alone.
The practical result is that AlphaGenome does not replace AlphaFold or AlphaMissense; it fills in the part of the genome those models were never designed to interpret. Used together, the three models let a researcher move from a raw DNA variant, through its likely regulatory effect (AlphaGenome), through its likely effect on a protein's function if it does fall in a coding region (AlphaMissense), through to that protein's resulting 3D shape (AlphaFold) — covering most of the chain from genetic change to biological consequence within one connected suite of models, even though each was trained and released separately.
Deep Learning Architecture
Under the hood, AlphaGenome uses a layered architecture rather than a single mechanism. Convolutional layers first scan the input sequence for short, local patterns — the kind of motifs a transcription factor might bind to. Transformer layers then sit on top of that, letting the network relate information across the entire input window, which is essential because a regulatory element sitting tens or hundreds of thousands of base pairs away can still switch a distant gene on or off. Finally, modality-specific output heads translate the shared internal representation into separate predictions — gene expression, splicing, chromatin accessibility, binding — from that one underlying model, rather than requiring a patchwork of specialised tools trained separately for each task.
DeepMind also reports a notable efficiency gain in training: a single AlphaGenome model trains in around 4 hours using roughly half the computational budget that the earlier Enformer model required, with training distributed across multiple interconnected TPUs for a single sequence.
The choice of a 1-million-base-pair input window is itself a deliberate architectural trade-off. Wide enough to capture most known long-range regulatory interactions — such as an enhancer element acting on a gene hundreds of thousands of base pairs away — but still short enough to train and run without the prohibitive compute cost of modelling the full genome at once, it sits at roughly the practical ceiling for what current transformer-based approaches can handle efficiently. Output at single base-pair resolution, meanwhile, is what lets the model distinguish the effect of one specific letter change from its immediate neighbours, rather than only describing the general regulatory character of a broader region.
Benchmark Results
DeepMind evaluated AlphaGenome against a wide range of existing, often task-specific, genomics models. According to the published results:
- AlphaGenome outperformed external specialised models on 22 of 24 sequence-prediction tasks it was tested against.
- It matched or exceeded the strongest existing models on 24 of 26 variant-effect prediction evaluations.
- Reported improvements over prior best-in-class models ranged from around +3.1% (histone modification prediction) up to +25.5% (predicting the direction of RNA expression change caused by a variant).
- DeepMind describes it as the only model in the comparison that jointly predicts all of the assessed modalities from a single network, rather than requiring a different specialised model per task.
Source: Google DeepMind's AlphaGenome announcement and accompanying preprint. As with any benchmark suite, these figures describe performance on DeepMind's chosen evaluation tasks; independent replication across additional cell types and disease contexts is an ongoing part of how the genomics community will stress-test the model.
Medical Applications
The implications DeepMind highlights for medicine and biological research are significant:
- Rare and Mendelian disease: identifying the likely functional impact of mysterious non-coding mutations in patients with undiagnosed genetic conditions, where the causal variant often sits outside any protein-coding gene.
- Cancer research: understanding how driver mutations dysregulate cellular processes. In one documented example, researchers used AlphaGenome to investigate T-cell acute lymphoblastic leukaemia (T-ALL) and found it correctly predicted that a specific mutation activates the TAL1 gene by introducing a new MYB transcription-factor binding motif — replicating a disease mechanism that took researchers years to establish experimentally.
- Synthetic biology and therapeutics: designing synthetic DNA sequences with more precisely targeted, tissue-specific activity, relevant to gene-therapy vector design.
In a typical rare-disease workflow, a clinical genetics lab sequencing a patient with an undiagnosed condition might turn up dozens of non-coding variants of uncertain significance. Manually testing each one experimentally — for example with a reporter-gene assay in cultured cells — can take weeks per variant. Running each candidate through AlphaGenome first lets researchers rank them by predicted functional impact across relevant tissues in minutes, so laboratory time is spent validating the handful of variants the model flags as most likely to matter, rather than working through the full list in sequence.
Availability & the AlphaGenome Atlas
AlphaGenome launched via an API for non-commercial research use in mid-2025, alongside a community forum for researchers at alphagenomecommunity.com and a form for organisations to register commercial interest. DeepMind has been explicit that predictions are "intended only for research use and haven't been designed or validated for direct clinical purposes" — this is a scientific tool, not a diagnostic one.
In September 2026, DeepMind extended this considerably with the release of the AlphaGenome Atlas: a pre-computed, roughly 1-petabyte database scoring the predicted regulatory impact of essentially every one of the roughly 9 billion possible single-nucleotide variants across the human genome. The Atlas introduces the AlphaGenome Variant Impact (AVI) score, a single combined metric spanning both coding and non-coding regions, designed to let researchers quickly triage which variants are worth investigating further without running the model themselves. Unlike the original research API, the Atlas is accessible through a no-code web portal, lowering the barrier for labs without dedicated machine-learning infrastructure.
The distinction between the two access routes matters in practice. The original AlphaGenome API suits computational biologists who want to run the model themselves against custom or patient-specific sequences, integrate it into an existing analysis pipeline, or score variants that are not simple single-nucleotide changes. The Atlas instead suits the much larger population of biologists and clinicians who need a quick answer for a known variant of interest — a genome-wide association hit, say, or a variant flagged during clinical sequencing — without writing any code at all. DeepMind has indicated a full, broader model release is planned for the future, beyond the current non-commercial research API and the Atlas lookup tool.
Limitations
DeepMind has been reasonably candid about where AlphaGenome currently falls short:
- It struggles to capture the effects of very distant regulatory elements — those more than around 100,000 base pairs from the gene they influence.
- Cell- and tissue-specific regulatory patterns are not always fully captured, particularly for rarer cell types under-represented in training data.
- It is not designed to predict an individual's whole personal genome or phenotype; it scores the likely functional effect of a sequence or variant, not a person's outcome.
- It cannot fully predict complex trait development, which depends on environmental and developmental factors well outside what a sequence-to-function model can capture.
None of this is unusual for a first-generation model tackling a problem this hard — the same was true of early AlphaFold releases before subsequent versions closed many of the initial gaps — but it is worth stating plainly rather than implying AlphaGenome has "solved" gene regulation. It is a substantial step forward in predictive accuracy and breadth, not a complete map of how the genome works.
Verdict
AlphaGenome is a tool for scientists, not consumers, and DeepMind has been consistent about keeping its research-only framing intact even as it has expanded access through the Atlas. But its downstream impact is likely to be felt well beyond the labs using it directly. By meaningfully accelerating how quickly researchers can go from "we found a variant" to "we understand what it does", it shortens one of the slowest steps between genetic discovery and treatment. Combined with AlphaFold and AlphaMissense, it also completes a fairly clear DeepMind biology stack — sequence function, coding variant impact, and protein structure — that cements the lab's position as a genuinely serious biological research institution, not only an AI one. For more on how DeepMind's leadership and scientific priorities have been evolving, see our coverage of Demis Hassabis stepping down.
This article was last updated on 12 September 2026 to add sourced architecture, benchmark and availability detail, including the AlphaGenome Atlas released that same month.




