MosaicML has released MPT-7B, a language model with nearly 7 billion parameters that the company says can match Meta’s 7-billion-parameter LLaMA model. Its commercial-use license and a family of specialized variants give the release several distinct uses, including processing much longer passages than the base model’s headline size might suggest.
A model built for commercial use
MosaicML trained MPT-7B on its own dataset of nearly a trillion tokens. The training followed the regimen of Meta’s LLaMA model and took 9.5 days on the MosaicML platform, at a cost of nearly $200,000.
According to MosaicML, MPT-7B reaches the performance of LLaMA, making it the first open-source model to do so and putting it ahead of OpenLLaMA. The article highlights a practical difference: unlike Meta’s models, MPT-7B is licensed for commercial use.
That licensing distinction matters to organizations evaluating whether a model can be used in commercial work. The release offers a starting point for those uses, while its different versions are designed for different kinds of interactions and workloads.
Four versions for different tasks
MosaicML released the “MPT-7B Base” model along with three variants: MPT-7B-StoryWriter-65k+, MPT-7B-Instruct and MPT-7B-Chat. Together, they extend the release beyond a single general-purpose model.
- MPT-7B Base is the core model in the family.
- MPT-7B-Instruct is designed to follow instructions.
- MPT-7B-Chat is a chatbot variant in the style of Alpaca or Vicuna.
- MPT-7B-StoryWriter-65k+ is built to read and write stories with very long context lengths.
The variants reflect different ways people may want to use an open-source language model. A chat interface, instruction-following behavior and long-form story work call for different capabilities, so the model family gives users options beyond the base release.
StoryWriter can work with entire novels
The StoryWriter variant was fine-tuned with a context length of 65,000 tokens using a subset of the books3 dataset. Context length describes how much text a model can handle at once. In this case, MosaicML says the model could read entire novels and write an epilogue.
The article compares that capacity with the largest GPT-4 variant of OpenAI, which is able to handle 32,000 tokens. The comparison points to the scale of the StoryWriter context window: it is designed to accommodate a much longer input than a typical short prompt, making whole-book reading and continuation possible.
MosaicML says the model can scale beyond 65,000 tokens with some optimizations. The team demonstrated up to 84,000 tokens on a single node using Nvidia A100-80GB GPUs. Those figures describe what the team reported for this setup, while the released StoryWriter model was fine-tuned for a 65,000-token context.
What the release offers
MPT-7B combines a nearly 7-billion-parameter base model with commercial-use licensing and specialized versions for instructions, chat and long-form writing. The StoryWriter model’s extended context is its clearest distinction: it is intended to work across very long texts, including novels.
All MPT-7B models are available on GitHub. For developers and organizations, the release presents a commercially licensed open-source option with multiple variants to explore, from conversational use to reading and continuing long stories.