MPT-7B Brings Commercial Use and Long Context to Open Models

MosaicML released MPT-7B, an open-source language model licensed for commercial use, alongside instruction-following, chat and long-context variants. Its StoryWriter version was fine-tuned for 65,000 tokens, enabling it to work with entire novels and produce an epilogue.

WTF Index IDIOCRACY
◄ Terminator 0 Idiocracy 1 ►

The release is a routine model launch, with only a mild lean toward dependence through instruction-following and chatbot variants.

MPT-7B Brings Commercial Use and Long Context to Open Models

MosaicML has released MPT-7B, a language model with nearly 7 billion parameters that the company says can match Meta’s 7-billion-parameter LLaMA model. Its commercial-use license and a family of specialized variants give the release several distinct uses, including processing much longer passages than the base model’s headline size might suggest.

A model built for commercial use

MosaicML trained MPT-7B on its own dataset of nearly a trillion tokens. The training followed the regimen of Meta’s LLaMA model and took 9.5 days on the MosaicML platform, at a cost of nearly $200,000.

According to MosaicML, MPT-7B reaches the performance of LLaMA, making it the first open-source model to do so and putting it ahead of OpenLLaMA. The article highlights a practical difference: unlike Meta’s models, MPT-7B is licensed for commercial use.

That licensing distinction matters to organizations evaluating whether a model can be used in commercial work. The release offers a starting point for those uses, while its different versions are designed for different kinds of interactions and workloads.

Four versions for different tasks

MosaicML released the “MPT-7B Base” model along with three variants: MPT-7B-StoryWriter-65k+, MPT-7B-Instruct and MPT-7B-Chat. Together, they extend the release beyond a single general-purpose model.

  • MPT-7B Base is the core model in the family.
  • MPT-7B-Instruct is designed to follow instructions.
  • MPT-7B-Chat is a chatbot variant in the style of Alpaca or Vicuna.
  • MPT-7B-StoryWriter-65k+ is built to read and write stories with very long context lengths.

The variants reflect different ways people may want to use an open-source language model. A chat interface, instruction-following behavior and long-form story work call for different capabilities, so the model family gives users options beyond the base release.

StoryWriter can work with entire novels

The StoryWriter variant was fine-tuned with a context length of 65,000 tokens using a subset of the books3 dataset. Context length describes how much text a model can handle at once. In this case, MosaicML says the model could read entire novels and write an epilogue.

The article compares that capacity with the largest GPT-4 variant of OpenAI, which is able to handle 32,000 tokens. The comparison points to the scale of the StoryWriter context window: it is designed to accommodate a much longer input than a typical short prompt, making whole-book reading and continuation possible.

MosaicML says the model can scale beyond 65,000 tokens with some optimizations. The team demonstrated up to 84,000 tokens on a single node using Nvidia A100-80GB GPUs. Those figures describe what the team reported for this setup, while the released StoryWriter model was fine-tuned for a 65,000-token context.

What the release offers

MPT-7B combines a nearly 7-billion-parameter base model with commercial-use licensing and specialized versions for instructions, chat and long-form writing. The StoryWriter model’s extended context is its clearest distinction: it is intended to work across very long texts, including novels.

All MPT-7B models are available on GitHub. For developers and organizations, the release presents a commercially licensed open-source option with multiple variants to explore, from conversational use to reading and continuing long stories.