Thursday, October 8, 2026
English edition

Development

Google's AI genome system evaluates every possible one-base change

September 9, 2026 Development Source: Ars Technica

Google's AI genome system evaluates every possible one-base change

Share this article

But most of the non-coding DNA is junk—the remains of ancient viral infections, DNA-level parasites, genes that have been inactivated by mutation, and so on. Figuring out what’s useful and what’s not has been an ongoing challenge for biologists for many reasons. First, the proteins that interact with DNA aren’t that picky about the sequences they stick to, potentially binding at random throughout the genome and tolerating a certain degree of mutation. Many of these proteins are also cell-type specific; there’s a different population of DNA-binding proteins in liver cells, nerve cells, immune cells, and so on. In many cases, having many different protein binding sites in a compact space matters more than the presence of any one of them. We’ve developed various software tools that identify individual sites of interest in non-coding DNA. But this is exactly the sort of problem that AI is good at solving: one involving probabilities that are imprecise and rely heavily on context. So Google developed the AlphaGenome AI system, which evaluates sequences for their potential function. Google has now used this system to evaluate all possible single-base changes in the human genome. In other words, if the first base on chromosome 1 is an A, the system evaluates what changes if you swap in a G, C, or T. It then moves on to the next base and repeats the process. Let’s be clear: Nobody on Earth probably has the reference genome sequence that Google is using as its baseline. Each of us likely differs from it at millions of bases in our genome, and most of us probably carry some combinations of insertions, deletions, duplications, and flipped sequences. Many of these are of no consequence because they fall within sequences that don’t have any functional significance. At the same time, many of the sequences that are non-functional are identical in all humans, simply because there hasn’t been enough time since we’ve had a common ancestor to accumulate changes. So only a small fraction of the changes that Google evaluated are likely to both show up in an actual human genome and have a functional significance. Why would Google bother? There are a couple of reasons this could be useful. First, it essentially pre-calculates the potential mutations researchers might be interested in. So if someone finds a mutation from some genome sequencing, they can get an immediate sense of its potential consequences. The second is that it provides researchers the chance to search across the entire genome for changes that have a specific type of impact. But much of that information could have been obtained simply by looking at the ENCODE dataset that was among the data the model was trained on. The greater potential here is developing the system to the point where it can perform an analysis we can trust on things it has no training data for, such as the Neanderthal and Denisovan genomes or cell types we don’t have ENCODE data for. At the moment, it’s not clear whether we’re there yet.