Mistral AI has released Mistral 7B, a language model with 7.3 billion parameters. The company says it performs better than some larger Meta Llama models on measured benchmarks, while requiring less memory and supporting higher data throughput.
What Mistral says the benchmarks show
Mistral AI reports that Mistral 7B outperformed Llama 2 13B across all the benchmarks it measured. It also says the new model beat Llama 1 34B on many benchmarks. These comparisons cover reasoning, world knowledge, reading comprehension, math and code.
The results are the company’s claims, and the source does not provide detailed benchmark scores. Mistral says the model approaches CodeLlama 7B’s programming performance while remaining capable on English-language tasks. It also describes Mistral 7B as comparable to a theoretical Llama 2 model more than three times its size.
There are limits to that picture. Mistral attributes weaker results than Llama 1 34B on knowledge questions to Mistral 7B’s lower parameter count. That distinction matters: performance varies by task, so a strong overall comparison does not mean the smaller model leads in every area.
Efficiency comes from attention design
Mistral points to two Transformer optimizations as reasons the model can work efficiently. Grouped Query Attention, or GQA, processes multiple queries simultaneously. The goal is to raise computational efficiency while maintaining model performance.
Sliding Window Attention, or SWA, limits attention to a defined context window within a sequence. This approach is meant to balance computational cost and output quality. Mistral says SWA doubles speed for sequence lengths of 16k with a context window of 4k.
These design choices support the company’s broader claim that Mistral 7B can deliver useful results with a smaller model footprint. Lower memory needs and increased data throughput could matter for teams deploying models themselves, though the source does not quantify those savings.
Available to download and adapt
Mistral 7B is offered under the Apache 2.0 license. The source says it can be downloaded for free and deployed using the reference implementation, through vLLM Inference Server and Skypilot in AWS, GCP or Azure, or via HuggingFace.
Mistral AI says the model can be adapted to tasks such as chat and instruction following through fine-tuning. The company also describes an instruction-tuned version, Mistral 7B Instruct, created using HuggingFace instruction datasets. It reportedly outperforms all 7B models on MT-Bench and competes with 13B chat models.
That combination of an open license and adaptation options gives developers a route to experiment with the model in different settings. The reported benchmark results offer one guide to its capabilities, while task-specific fine-tuning may shape how useful it is for a particular application.
A first release in a broader plan
Mistral AI is a French startup whose team includes former Meta and Google DeepMind employees. The company made headlines in June after announcing a $105 million European seed round before it had a product. Former Google CEO Eric Schmidt is among its high-profile investors.
The company’s stated business model is to distribute powerful open-source models and offer specific paid features to customers. A leaked pitch letter reportedly said that top-of-the-line models could be paid for, and that Mistral planned a family of text-generation models by the end of 2023. The letter said those models would “significantly outperform” ChatGPT with GPT-3.5 and Google Bard, with some of the family open source.
Mistral 7B is therefore presented as both a usable release and an early sign of the company’s plans. Its reported performance, deployment options and fine-tuning path give developers several ways to evaluate it, while Mistral’s later ambitions remain claims described in the leaked letter.