Wednesday, August 19, 2026
English edition

Development

Large genome models used to design new viruses

August 7, 2026 Development Source: Ars Technica

Large genome models used to design new viruses

Share this article

But our own lack of knowledge places limits on what we can do with these models. If we prompt them with a bit of sequence from a complex cell, it will respond with a string of bases that contain what look like genes and regulatory DNA. But, since most genes can be located essentially anywhere in a eukaryotic genome, we have no idea what functions these hallucinated genes might perform (if any), so we can’t really test them in any way. Still, as a precaution while training these models, termed Evo 1 and Evo 2, the researchers did not provide them with any sequences from viruses that target complex cells. Even if we can’t understand what they output, there’s a chance that they’ll output something dangerous. But there’s a whole world of viruses that attack bacteria and can’t infect humans. And those, the Evo developers figured, are fair game. So, rather than looking at whether Evo 1 and Evo 2 could output reasonable-looking genes, they decided to test if they could output an entire genome. The model virus they chose as a test case is the catchily named ΦX174, part of a larger family of viruses that infect E. coli. In addition to the convenience of using the best-studied bacteria on the planet for tests, ΦX174 has a number of notable features. It’s fairly simple: At 11 genes spread over about 5,400 bases, all 11 of those genes have identified functions, and the virus’s infection cycle is well characterized. It also has a useful feature from the perspective of working with a large genome model: The end of the virus always has the same short sequence of bases. So, if you use those bases as a prompt, the large genome model should be able to recognize that it needs to respond by outputting a sequence that is in some way related to ΦX174. To further prepare their models, Evo 1 and Evo 2 were fed over 2 million additional bases of DNA sequences from viruses that infect bacteria (termed bacteriophages), and then fine-tuned with sequences specific to Microviridae, the group that ΦX174 belongs to. After that, they experimented with different prompts. Prompt with too much of the ΦX174 start sequence, and it would simply spit the rest of the genome back out. Too little, and it would generate lots of unrelated sequences. The researchers eventually found prompts with four to nine bases of the start sequence worked best. While the outputs were primed to return something that looked like ΦX174, they literally could be anything—from near-carbon-copies of the normal virus to sequences that are only vaguely viral. To address this, the research team put a lot of pre-conditions on the outputs designed to throw away some of the more extreme ones while still enforcing a bit of tinkering. For example, ΦX174 uses its own version of a spike protein to latch onto and infect bacteria; if the gene that encodes spike is missing or damaged, the virus simply wouldn’t work. So they discarded any outputs that had a spike protein that was less than 60 percent identical to the real one. They also threw out viruses that were too long or too short (< 4,000 bases or > 6,000), any that had strings of the same base more than 10 bases long, and any that had unusual frequencies of the two base pairings (GC and AT). Throwing out those at the computer analysis stage left them with a reasonable number of potential outputs to test: 302 proposed viral sequences. They were able to chemically synthesize the sequences of 285 of them and inserted them into bacteria to see what happened. In most cases, the answer was “nothing.” But 16 of the sequences managed to inhibit the growth of the E. coli, suggesting they were actually working as viruses. Nine were the initial sequences output by the AI, and the remaining seven had acquired additional mutations after being inserted into bacteria. This tells us the AI is far more likely to make changes that leave a viable virus behind than random mutation. This study is interesting from an intellectual perspective, but the researchers also show that it might be useful. Bacteriophages have been tested as therapies for bacterial infections that are resistant to standard antibiotics. But many bacterial strains have already evolved resistance to some families of viruses, including ΦX174. One potential solution to that is to use a cocktail of bacteriophages, since it may not be easy or even possible to evolve resistance to all of them. And here, the AI has produced a cocktail of potential viruses. So the team tested a cocktail of natural bacteriophages that normally prey on E. coli and compared it to a cocktail of the 16 viable AI products. In their tests, the natural bacteriophage cocktail failed, but the AI-generated group managed to quickly evolve the ability to infect their otherwise resistant hosts. The researchers suspect that the process of overcoming resistance involved some of the AI-designed viruses swapping segments of DNA and the appearance of additional mutations. It’s not clear whether this is a major potential benefit; phage therapies have been under consideration for years but haven’t seen widespread public use, despite the growing problems with drug-resistant bacterial strains. And the authors are very clear that large genome models come with risks. Although they excluded viruses that infect vertebrates from their training data, anyone with access to sufficient computing resources could repeat the process with those viruses included. The paper ends with a call for better governance of this and related areas, such as the ordering of custom DNA sequences. So far, however, regulation of AI has largely failed to keep up with the rapid pace of the field. Science, 2026. DOI: 10.1126/science.aec2657 (About DOIs).