Protein researchers can now draw on a large open collection of sequence alignments and related information. OpenProteinSet brings together data that can help scientists study how protein sequences relate to structure and function, and gives machine learning researchers a resource for developing and training models.
Why sequence alignments matter
Proteins carry out many of life’s functions, and their amino acid sequences encode information about what they do. Over generations, related proteins can accumulate small sequence changes while retaining much of their structure and function.
Multiple sequence alignments, or MSAs, arrange evolutionarily related sequences so corresponding amino acids appear in the same columns. Gaps are inserted where needed to make the comparison possible. Patterns across an alignment can offer clues about a protein’s structure and function.
These alignments have long supported protein research. Their importance grew with AlphaFold 2, which uses large amounts of MSA data to predict protein structures with near-experimental accuracy. Although AlphaFold 2 is open source, its training data was not made available to the research community, leaving a gap for researchers seeking to reproduce or build on that work.
What OpenProteinSet includes
OpenProteinSet provides 16 million MSAs and associated data. The collection includes alignments for all 140,000 proteins in the Protein Data Bank (PDB), which the source describes as the definitive database of experimentally determined protein structures.
It also draws on sequences from the UniProt knowledge base, grouped into clusters by similarity. For PDB proteins, the dataset supplies raw alignments from multiple sequence databases and includes structurally similar proteins found through searches of the PDB.
Another part of the collection is predicted structures from AlphaFold2 for 270,000 different UniProt clusters. Together, the alignments, sequence groupings, related proteins, and predicted structures give researchers several kinds of material to use when studying proteins or developing computational methods.
Testing the data with OpenFold
The developers used OpenProteinSet to train OpenFold, an open recreation of AlphaFold 2. They report that OpenFold performs on par with the original, which they present as evidence that the open dataset is sufficient for training a comparable system.
That result matters because releasing model code alone does not give other researchers everything needed to recreate a system. Training data can shape what a model learns, and access to a substantial open collection gives research teams a way to conduct their own work with the underlying sequence information.
The developers say OpenProteinSet increases both the quantity and quality of precomputed MSAs available to molecular machine learning communities. They also point to immediate uses across diverse tasks in structural biology. The dataset is hosted and available on AWS.
A resource for further research
OpenProteinSet addresses a practical constraint in structural biology: researchers need extensive sequence comparisons to investigate protein structure, while a major model’s training data remained private. By making a large collection of alignments and associated materials available, the project gives scientists an open starting point for new analyses and model development.
The reported OpenFold result offers one demonstration of that potential. The source does not claim that every application will produce the same outcome, but the dataset expands what researchers can access as they pursue questions about protein structure, function, and machine learning.