The Allen Institute for AI (AI2) has released Dolma, a text dataset intended to support research on its planned open language model, OLMo. The release gives researchers a way to inspect information about the material and processing behind a language model training dataset—details that AI companies often keep private.
Opening up the material behind a model
AI2’s stated aim is to make both OLMo and the data used to create it available to the AI research community. Dolma’s name refers to “Data to feed OLMo’s Appetite,” and the dataset is the first data artifact AI2 has published for the project.
Many organizations share some basic information about the datasets behind their language models, but the source article says much of the underlying detail remains proprietary. That can make it difficult for outside researchers to understand what was included, what was removed, or how the material was judged and prepared.
Those questions matter for studying how a dataset shapes a model. For example, researchers may want to know what counted as high-quality text, whether personal information was removed, and why particular material was excluded. Without documentation, comparing or replicating work becomes harder.
What AI2 says it is documenting
AI2 presents Dolma as a more inspectable alternative. The organization says it is publicly documenting the dataset’s sources and the steps used to prepare it, including how and why it was limited to original English-language texts.
The publication describes Dolma as the largest open dataset of its kind at the time, containing 3 billion tokens. Tokens are a measure of content volume used in AI. The source also characterizes the dataset’s permissions as straightforward, and says it is offered under the “ImpACT license for medium-risk artifacts.”
More information was still to come: AI2 said a more comprehensive paper was in the works. The release therefore offers researchers an initial account of the dataset and its preparation, with a fuller treatment planned separately.
Why openness matters for research
Publishing sources and methods can help people outside a model-building organization examine how training data was assembled. Researchers can use that information to better understand what a model may have learned from its data and to study the choices involved in preparing a large text collection.
The source article also raises concerns about whether some closed datasets may include material obtained without permission, such as pirated copies of authors’ books. It presents that as speculation, not as an established claim about any particular dataset. More generally, limited disclosure makes it difficult for outsiders to assess the contents and handling of training data.
Making a dataset open does not answer every question about its contents or use. But documenting sources, filtering decisions, and permissions gives researchers more to examine than a release that shares only a few summary statistics.
Access and personal data requests
Dolma is available through Hugging Face, according to the source. AI2 also provides a form for people who believe their personal information may have been included and want to request its removal.
The article describes that form as intended for specific cases, rather than a general request to opt out of data use. Together, the access point and removal process give prospective users and affected individuals distinct ways to engage with the dataset.
For AI2, Dolma is part of a broader effort to make the foundations of OLMo available for scrutiny alongside the model itself. Researchers can consult the dataset and its documentation to investigate how its sources and preparation support that open research goal.