Why Beam puts open-weight AI competition on efficiency

Reflection has announced Beam, an open-weight model aimed at coding, reasoning, and agentic tasks. The company says Beam matches GLM 5.2 on key reasoning benchmarks while using three to four times less compute, with weights planned under Apache 2.0 "later this month."

Why Beam puts open-weight AI competition on efficiency

Reflection is preparing to release Beam, its first freely available AI model, with a clear message: capability matters, but compute efficiency may matter just as much for teams running coding and reasoning systems at scale.

The model is positioned against Chinese open-weight systems from Deepseek and Qwen, while Reflection says its focus is not simply topping every benchmark. Instead, Beam is built to offer strong performance with lower operating demands.

A model built around efficient reasoning

Beam is designed for coding, logical reasoning, and agentic tasks. Reflection describes it as a mixture-of-experts model with 501 billion total parameters, but only 23 billion parameters active per token.

That architecture is central to the company’s claim. By activating only part of the model for each token, Beam is meant to keep compute costs relatively low while still handling difficult work.

According to Reflection, Beam matches GLM 5.2 on demanding reasoning tasks while using three to four times less compute. For businesses using AI for software development and automated workflows, that tradeoff is the heart of the pitch: useful reasoning without letting inference costs dominate deployment decisions.

Reflection also says Beam comes close to Qwen3.8-Max on coding and agent benchmarks. The company acknowledges that stronger open models such as Kimi K3 still lead on raw performance.

What the training run says about Reflection’s strategy

Beam’s abilities come from large-scale pretraining combined with a major reinforcement learning phase. Reflection says that phase used 10,500 Nvidia GB300 GPUs for over four weeks, and describes it as one of the largest training runs carried out by an open lab.

The company says performance continued improving through the end of that run without hitting a ceiling. That detail matters because it frames Beam not as a finished endpoint, but as evidence that more reinforcement learning may still improve the model family.

Reflection says it is already training a successor intended to close the remaining gap with the strongest open models. Beam, then, is both a product release and a signal about the company’s direction: more coding, more reasoning, and more work on models that can act through computers.

Users will also be able to control how deeply Beam reasons through a task. A tunable parameter lets the model answer quickly or spend more time on harder problems, trading compute cost against output quality.

Agentic behavior beyond the training mix

Reflection says it observed what it calls "emergent capabilities" during Beam’s reinforcement learning work. While the training mix included reasoning, software engineering, and terminal tasks, the company noticed the model improving at web browsing even though browsing tasks were not included in that mix.

With web access, Beam independently learned to query other language models and retrieve documents from external services, according to Reflection. That is especially relevant to agentic AI, where the model is expected to do more than generate text in isolation.

The demos shared by Reflection also emphasize this direction. They include a live-updating New York City subway map, a small 3D game, and a notebook for fine-tuning another AI model.

Beam itself is text-only. However, Reflection says it can handle content from other media formats when that content is represented as text.

Safety, release plans, and company backing

Reflection trained a second model for safety and alignment, then merged it with Beam. The safety guidelines include hard rules the model must never break, as well as quality standards such as factual accuracy, admitting uncertainty, and using a direct, thorough, and proactive response style.

The company says it plans to publish safety test results in a technical report and open-source the evaluation methods it developed. Beam is still undergoing final safety testing, and an early version is currently available to select users.

Reflection says the technical report, developer documentation, and model weights will ship under the Apache 2.0 license "later this month." The company releases model weights, but keeps its training data and pipelines proprietary.

The startup was founded in 2024 by former Google Deepmind researchers Misha Laskin and Ioannis Antonoglou. Laskin led reward modeling for Gemini, while Antonoglou helped build AlphaGo.

Reflection launched in March 2025 with $130 million in seed funding and a goal of building superintelligence through autonomous coding. Its idea was that language models could learn to act independently on computers through reinforcement learning, in a way connected to how AlphaGo plays Go.

In summer 2025, Reflection released Asimov, an agent for analyzing large codebases. In October 2025, the company raised $2 billion at an $8 billion valuation, with Nvidia among the investors.

Since then, Reflection has positioned itself as a Western counterpart to Deepseek and Qwen. More recently, it signed billion-dollar compute deals with SpaceX and cloud provider Nebius.

Why Beam matters for open-weight AI

Beam’s importance is not only that it is another large model. Its larger claim is that open-weight AI competition is becoming a contest over efficiency, reasoning depth, and practical deployment cost.

If Reflection’s claims hold up in broader use, Beam could appeal to teams that need coding assistance, reasoning, and automated workflows but cannot treat compute as unlimited. That is the market Reflection is clearly targeting: AI systems that are strong enough for hard tasks, but efficient enough to run in real products.