Kolibri puts European AI sovereignty into practical terms

Aleph Alpha has released Kolibri, a German-English open-weight language model with 78 billion parameters. The company presents it as a cost-conscious model for public administration, aviation and industry, developed under European law with the EU AI Act in mind.

Kolibri puts European AI sovereignty into practical terms

Aleph Alpha has introduced Kolibri, a German-English language model that puts several current AI priorities into one release: open weights, European development, long context, and a focus on practical use in regulated environments.

The model has 78 billion parameters, with about three billion active per token through a mixture-of-experts architecture. Aleph Alpha says Kolibri is aimed at public administration, aviation, and industry, which places the release in areas where cost, language coverage, and legal context matter as much as raw model capability.

What Aleph Alpha released

Kolibri is a German-English language model. Its weights are available under an Apache 2.0 license on Hugging Face, making the release an open-weight model rather than a closed system available only through a managed interface.

The model uses a mixture-of-experts architecture. In the figures provided by Aleph Alpha, Kolibri has 78 billion parameters in total, while about three billion are active per token. That distinction matters because the company is presenting the model not only around size, but also around operating cost.

Aleph Alpha claims Kolibri sits on the Pareto front of quality and operating cost in both German and English. In plain terms, the company is arguing that the model is competitive on performance while also being efficient to run, compared with models using similar architectures.

The source notes that some of the compared models are significantly older, specifically from March and April 2026. That context is important because model comparisons can depend heavily on what is being compared, how recent those models are, and which metrics are used.

Why German-English training matters

Kolibri is positioned around German and English rather than as a general language release with no clear regional emphasis. According to Aleph Alpha, German accounts for 21.3 percent of the training data.

The company also built a dedicated German data pipeline for the project. That detail is central to the sovereignty argument around Kolibri: the model is not only being described as European because of where it was developed, but also because of how its German-language training data was handled.

The source also states that Chinese models were used to generate synthetic training data. That means Kolibri's training process included generated material from outside the German-English focus of the final model, even while the release itself is framed around European AI sovereignty.

For organizations that operate in both German and English, the language mix is one of the practical points of interest. A model serving public administration, aviation, or industry needs to handle domain work in the languages those organizations use, and Aleph Alpha is presenting Kolibri as a model built with that requirement in mind.

Built with European law in view

Aleph Alpha says Kolibri was developed under European law with the EU AI Act in mind. The source does not provide a detailed compliance analysis, so the strongest supported point is the company's positioning: Kolibri is meant to fit a European legal and governance context.

That framing is especially relevant because the target sectors named by Aleph Alpha are not casual consumer use cases. Public administration, aviation, and industry often involve formal processes, institutional requirements, and systems where deployment choices can carry long-term operational consequences.

The model was trained on 768 B200 GPUs in Germany and Finland, according to the tech report cited in the source article. Those locations reinforce the European framing of the release, alongside the German-English language focus and the stated attention to European law.

Kolibri also supports context windows of up to one million tokens. That gives the model room to work with very large inputs, although the source does not specify benchmark results for particular long-context tasks.

The practical case Aleph Alpha is making

The central argument around Kolibri is not simply that another large language model has arrived. Aleph Alpha is presenting the release as a combination of capability, cost awareness, openness, and European alignment.

Several facts support that positioning:

  • Kolibri is a German-English model with 78 billion parameters.
  • About three billion parameters are active per token through a mixture-of-experts architecture.
  • German accounts for 21.3 percent of the training data.
  • The model supports context windows of up to one million tokens.
  • The weights are available under an Apache 2.0 license on Hugging Face.
  • Aleph Alpha says the model targets public administration, aviation, and industry.

For the AI market, the release highlights a familiar tension: organizations want strong models, but they also care about operating cost, language quality, licensing, and legal environment. Aleph Alpha's claim that Kolibri sits on the Pareto front of quality and operating cost is designed to speak directly to that tradeoff.

The open-weight release also changes how the model can be evaluated. With weights available on Hugging Face under Apache 2.0, Kolibri is being offered in a form that can be inspected and used beyond a single hosted product path.

The broader message is clear from the facts Aleph Alpha chose to emphasize. Kolibri is not being framed as a generic model release alone. It is being used to argue that European AI sovereignty can be expressed through model architecture, training choices, legal context, infrastructure location, and availability of weights.