Gene editing is moving from laboratory promise toward therapies, but precision remains the central safety problem. Even when a system is designed to target one gene, the human genome is large enough that similar sequences can appear elsewhere by chance.
A research team based at a variety of institutions in China used Google’s AlphaFold to examine where gene-editing proteins may allow these mistakes. Their work suggests that AI-guided structural analysis can help redesign Cas proteins so they keep useful editing activity while reducing off-target effects.
Why off-target effects matter
Gene-editing systems rely on several coordinated parts. The guide RNA is designed to base-pair with the target genome sequence. Because multiple guide RNAs can point to the same gene, one common safety step is to choose sequences that do not closely resemble other locations in the genome.
The Cas protein, named for the CRISPR system’s Cas9 protein, adds another layer of specificity. It interacts with the guide RNA and the DNA target, and it helps determine whether the complex remains bound.
The final component is the protein activity that changes DNA after Cas9 has bound to it. In the original CRISPR system, this activity cut both strands of the double helix, creating damage that can be difficult to control. Researchers have since adapted other proteins to work with Cas9 and make more focused changes, including removing a single base or making chemical modifications that alter base-pairing behavior.
The difficulty is that specificity is not absolute. Guide RNA and the genome pair over a stretch on the order of 18 bases long. A random sequence of that length should appear only about once in 70 billion bases, while the human genome is about 3 billion bases. On paper, that looks selective enough. In practice, Cas9 can still tolerate a small number of mismatches, and the acceptable number and position of those mismatches can vary.
How AlphaFold was used
The researchers began by building a broad library of off-target editing sites. They used a modified CRISPR system that converts the DNA base adenine into inosine, then isolated DNA fragments containing that change. The process was repeated with 10 different guide RNAs, giving the team many modified DNA fragments to analyze.
The next goal was to understand how the CRISPR complex behaved at those sites. Updated versions of AlphaFold can model interactions involving proteins and nucleic acids, as well as complexes made from multiple proteins. The team first gave AlphaFold the target DNA sequence, the guide RNA, the Cas9 sequence, and an enzyme that chemically modifies bases and can attach to Cas9.
That full setup did not work as intended. AlphaFold placed one of the proteins in a position that was clearly wrong.
The researchers then simplified the model. They supplied only the DNA, RNA, and Cas9 protein, because Cas9 is the main factor in sequence specificity. This version performed better, producing a structure consistent with experimentally determined structures involving real nucleic acids and proteins.
ContactSeek focused the search
By comparing AlphaFold structures for on-target and off-target sites, the team found that off-target binding often changed Cas9’s behavior in detectable ways. Many of the off-target sites, about two-thirds, caused Cas9 to adopt a slightly different overall structure. Nearly all of them, over 95 percent, changed which amino acids contacted the RNA.
That distinction mattered. Some off-target cases did not require Cas9 to take on a dramatically different shape. Instead, amino acids inside the protein could flex in ways that helped accommodate mismatched bases.
AlphaFold already included a way to estimate what the researchers called “contact probability.” In this context, it means the chance that two items, such as amino acids or nucleotides, are within a very small distance, defined here as eight Angstroms.
The team compared contact probability outputs for matched and mismatched targets. That let them identify which Cas9 amino acids changed their contacts when the guide RNA and DNA did not perfectly match. They named this analysis setup “ContactSeek.”
ContactSeek initially produced a long list of amino acids. To make that information useful, the researchers looked for clusters within Cas9 where many of those contact changes occurred. Those clusters pointed to regions that may help the protein adapt to mismatched DNA-RNA structures.
Redesigning Cas proteins
After identifying candidate regions, the researchers tested altered Cas9 versions. In all, they made 23 different swaps, replacing amino acids at one of 10 key positions found through the AlphaFold-guided analysis.
One redesigned variant kept activity similar to normal Cas9 at sites that matched the intended target. At the same time, its off-target activity dropped from 28 percent to 5 percent. The researchers reported similar results with different guide RNAs.
The approach was not limited to Cas9. The team also showed that it could work with a related system using Cas12 to recognize the DNA and guide RNA combination.
Other groups have already developed Cas9 variants with fewer off-target edits through methods such as directed evolution. When compared with those variants, the newly designed versions generally showed similar or slightly better activity and specificity.
The main difference is how targeted the new method may be. The changes identified through ContactSeek may be more specific to a particular guide RNA and mismatch combination, rather than broadly effective in every setting. That could make the approach especially useful when researchers need to understand why a particular editing setup creates off-target effects and how the Cas protein might be adjusted to reduce them.
What this means for gene editing safety
The work shows a practical use for AlphaFold beyond predicting a single protein shape. Here, the software helped compare how a gene-editing complex behaves across desired and undesired target sites. That comparison gave researchers a way to move from observing off-target edits to identifying protein regions that could be redesigned.
For gene editing, the safety challenge is not only choosing a better guide RNA. It is also understanding how Cas proteins tolerate imperfect matches. ContactSeek offers one route for doing that by turning structural differences into testable protein changes.
The result is not a claim that off-target effects disappear. It is a clearer workflow for finding the protein contacts that help create them, then testing whether changing those contacts can preserve on-target editing while reducing unwanted activity.