Why PottsMPNN pushes AI protein design past nature's examples

PottsMPNN is a machine-learning framework that helps AI design protein sequences without treating native sequences as the main target. It uses physical principles, amino acid interaction modeling, and evolutionary sequence sets to improve structural compatibility and mutation stability prediction.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 0 ►

PottsMPNN modestly advances AI’s ability to design novel biological systems beyond natural examples, though the story frames it as scientific and medical progress rather than harm.

Why PottsMPNN pushes AI protein design past nature's examples

Protein design is moving into territory where copying nature is no longer the main test. PottsMPNN, a new machine-learning framework developed in the Department of Biology, is built around a different question: can AI generate sequences that are structurally feasible even when they do not resemble any native protein?

The work, described in a paper recently published in PNAS, focuses on a central challenge in AI protein design. A protein’s function depends on its structure, and that structure depends on how its amino acid sequence folds. But nature does not offer just one valid answer for every fold.

Why native sequences are not enough

Many protein design methods begin with a target structure. A machine-learning system then proposes amino acid sequences that might fold into that structure. This process is important because designed proteins could include examples that bind to a disease-causing molecule in our cells.

The difficulty is that evolution selected only some sequences, not every useful one. Many different amino acid sequences can fold into the same structure. At the same time, one sequence can potentially adopt different structures, depending on a protein’s flexibility or a functional trigger.

That means a model can be too focused on reproducing what already exists. Amy E. Keating, Department of Biology head, Jay A. Stein (1968) Professor of Biology, professor of biological engineering, and senior author of the paper, frames the problem directly: “For years, the field has measured success by asking whether a model can reproduce the protein sequence that evolution happened to select — our work shows that this isn’t the best metric for protein design,”

For a completely novel protein structure, a native sequence may not exist as a useful reference. In that case, success has to be judged by whether the generated sequence is likely to fold into the intended structure and whether the model understands the relationship between sequence and stability.

What PottsMPNN changes

PottsMPNN incorporates physical principles that govern protein structure and stability. The goal is to improve sequence generation and strengthen predictions about how mutations will affect a protein’s stability.

The framework is described as having a better understanding of the sequence-energy landscape. In plain language, that means it models how the identity of each amino acid relates to the stability of the protein.

Graduate student and lead author Foster Birnbaum explains the shift in evaluation this way: “What we actually care about is how likely the generated sequences are to fold into the desired structures, how well the model understands the sequence-energy landscape, and how well it can predict the effect of mutations on the stability of the protein.”

That emphasis matters because protein design is not just about generating something that looks familiar. It is about finding sequences that can physically work for a target structure.

Learning from noise and amino acid interactions

Machine learning has recently changed the pace and breadth of fundamental biological research. The source article notes that only recently has it become possible to reliably use a computational model to generate a protein structure or sequence. It also notes that perhaps the most widely used model today was released in 2022.

Birnbaum studied strategic uses of what researchers call “noise,” which means adding variations to a protein structure during training. This reduces a model’s tendency to overcopy native sequences and increases the diversity of structures for which it can generate sequences.

PottsMPNN also uses a pairwise distribution to capture interactions between amino acids. This lets the framework account for physical interactions between all 20 possible sequence options at a pair of positions in the protein.

That pairwise modeling is presented as a key reason PottsMPNN more accurately models the sequence-energy landscape than other methods. Instead of treating amino acid choices as isolated decisions, the framework considers how pairs of positions can affect one another.

Evolution still helps, but it is not the endpoint

The researchers also introduced sets of evolutionarily related sequences into training PottsMPNN. The purpose was to teach the model that different sequences can adopt the same folded structure.

Birnbaum acknowledges a tension in that choice. A framework designed to move away from strict dependence on native sequences still uses evolutionary information during training. But the reported result is that as the model depends less and less on native sequences, structural compatibility and energy prediction improve, including for novel proteins.

This is the main implication of the work. Evolutionary examples remain useful, but they should not define the full search space for protein design. PottsMPNN points toward AI systems that can explore more possible answers while staying grounded in the physical requirements of folding and stability.

What it could mean for biological engineering

Birnbaum describes the long-term stakes plainly: “Once we can design any protein we want, that enables us to do a potentially scary amount of biological engineering,” He adds, “It’s a difficult task, but I’m really optimistic about this century’s progress in biology.”

The model could also be further improved and fine-tuned for a specific task. The source notes that such fine-tuning has in the past led to better predictions, including predictions about the outcome or consequence of a particular mutation.

Keating’s view is that the methods move the field toward useful new-to-nature proteins for diverse applications while providing a stronger foundation for future advances. The significance is not that AI can ignore biology. It is that AI protein design may become more useful when it stops treating nature’s selected sequences as the only correct answers.