Europe has spent years debating whether it can build frontier artificial intelligence on its own soil, and on October 3, 2026, German Unity Day, one of the continent's best known contenders answered with downloadable weights. Heidelberg based Aleph Alpha published Kolibri-1, an English and German language model the company describes as sovereign, open weight and aimed at regulated, mission critical work.
The model arrives in a field dominated by American and Chinese labs. European public institutions have repeatedly asked for alternatives that do not route sensitive workloads through foreign infrastructure or foreign legal regimes, and Aleph Alpha has positioned itself as that alternative since its earliest funding rounds. Kolibri-1 is the company's largest open artifact to date: a mixture of experts transformer with 78.1 billion total parameters, about 3.46 billion active per token, a context window stretching to 1,048,576 tokens, and an Apache 2.0 license that permits commercial reuse.
It did not appear out of nowhere. Aleph Alpha released a smaller 30 billion parameter predecessor, Kolibri Origin, on June 11, 2026, and finished pre training of the larger model on September 11, 2026. Roughly three months separate those milestones, a cadence that says as much about the company's engineering pipeline as about the model itself.
The target buyers are public administration, industrial firms, aerospace and defense, the customers that ask where data is processed, who controls the stack and which regulator holds jurisdiction. Aleph Alpha says Kolibri was built by teams in Germany and trained on infrastructure in Germany and Finland under European and German law, with no foreign control, and that the company owns the pipeline from data curation through pre training and post training to optimization.
Key Facts
Scale first. The Hugging Face model card for Aleph-Alpha/Kolibri-1 lists 78,103,074,560 total parameters and 3,457,573,120 active parameters per token, with a maximum context length of 1,048,576 tokens. Hugging Face reported on October 3 that the card recommends staying at or below 262,144 tokens when serving efficiency and complex task quality matter. The weights ship in float8_e4m3fn (FP8) precision and occupy roughly 78 GB, against about 156 GB in BF16. Minimum hardware is two A100 80GB GPUs, two H100 SXM5 GPUs, one H200, one B200 or one B300.
Training was done on 768 Nvidia B200 GPUs spread across 96 HGX 8xB200 nodes, in three stages. Aleph Alpha reported on October 3 that pre training covered 20 trillion tokens at a 16,384 token sequence length over 21 days, which works out to 511 hours, 392,000 GPU hours and 6.4e23 FLOPS. Mid training added 3.44 trillion tokens at 65,536 tokens over five days and 90,000 GPU hours, and long context adaptation added 201 billion tokens at 262,144 tokens over 13 hours and 10,000 GPU hours. The total sits near 24 trillion tokens; ForkLog reported on October 4 that the figure comes to about 23.64 trillion tokens.
The data mix is unusually German heavy for a bilingual model. German accounts for 21.3 percent of pre training tokens, roughly 4.3 trillion of them, against about 62 percent English and about 14 percent code. The model card breaks that down as 62.5 percent English, 23.9 percent German and 13.6 percent code. Only about 6 percent of the German material was machine translated, and the unique German corpus held about 2.4 trillion tokens. The stated knowledge cutoff is June 18, 2026, for both English and German.
Architecturally, Kolibri has 50 layers, all of them mixture of experts with one shared expert, 384 experts with six routed active per token, and a 4 to 1 sliding window to full attention pattern. Training used the Muon optimizer and a technique Aleph Alpha calls Exact Quantile Balancing. Users can set reasoning effort to none, low, medium or high, and the model supports tool calling. The card says Kolibri abstains when the supplied context does not support an answer, following what it calls a Merlin-Arthur protocol, and that it is built for human reviewed workflows rather than unsupervised operation.
Energy and verification round out the picture. AlexTech.ai reported on October 3 that Aleph Alpha estimates 950 MWh for pre training, mid training and the long context phase, including node power and data center overhead but excluding supervised fine tuning and reinforcement learning. Crypto Briefing reported on October 3 that Kolibri scored 96.9 percent on AIME 2025, a math competition benchmark, based on Aleph Alpha's own reporting. ForkLog noted that all benchmark results come from the developer and were not independently verified, and Aleph Alpha published a 189 page technical report covering architecture, training methods and evaluation.
Analysis
The bigger picture here is that European sovereign AI has stopped being a policy slogan and started being a shipping artifact. The interesting question is no longer whether a German lab can describe a frontier class model, but whether it can put weights on a public repository that customers in regulated sectors will actually deploy. Kolibri-1 clears that bar in the narrow sense: Apache 2.0 weights, a documented model card, a technical report and a hardware floor that a mid sized enterprise data center can meet with two accelerators.
What this really means is that Aleph Alpha is competing on governance and locality rather than on raw benchmark leadership. The company frames the model around the European AI Act, the General Purpose AI Code of Practice and GDPR, and it trains the model to refuse when context is insufficient, a behavior that matters to auditors even if it frustrates users who want a confident answer. That is a deliberate trade: slightly less fluent improvisation in exchange for an audit trail.
The efficiency story cuts both ways. With 3.46 billion active parameters out of 78.1 billion, computation per generated token is modest, which keeps inference economically plausible at scale. But the entire model must be resident in memory, so nobody is running Kolibri on a laptop. The roughly 78 GB FP8 footprint and the two accelerator minimum make this a data center product, and Aleph Alpha's recommendation to serve at 262,144 tokens rather than the full 1,048,576 suggests the headline number is a capability ceiling, not a default operating point.
Competitive claims deserve caution. The 96.9 percent AIME 2025 result is self reported, and ForkLog pointed out that no independent evaluation has confirmed it. German heavy pre training may help on administrative and legal text, but it does not automatically translate into frontier reasoning performance. The strongest honest claim available today is that Kolibri-1 is a credible, inspectable, legally clean option for European buyers who value jurisdiction over leaderboard position.
Why It Matters
For public sector buyers, the release expands the shortlist. Agencies that previously chose between American APIs and building in house now have an open weight model with a European training footprint, German language depth and a license that allows internal modification. That shortlist effect matters more than any single benchmark row, because procurement cycles in government and defense move slowly and are hard to reverse once a vendor is embedded.
For the wider open weight ecosystem, Kolibri-1 adds a genuinely multilingual entry at a scale where most releases default to English. A 21.3 percent German share in pre training is a real design decision, not a marketing line, and it produces a model that may handle German administrative, legal and industrial text better than larger English first alternatives. That is a niche, but it is a niche with money in it.
For Europe's regulatory agenda, the timing is symbolic. Announcing on German Unity Day, under an Apache 2.0 license, with a technical report and an abstention protocol, is a statement that compliance can be designed in rather than bolted on. Whether regulators and customers agree will be decided in procurement, not in a blog post.
Next Up
The immediate tests are practical. Independent evaluators will need to reproduce the self reported results, and third parties will want to measure the model at its recommended 262,144 token serving window rather than the full million. Enterprises will probe tool calling, abstention behavior and German language quality on their own documents, and hardware vendors will publish throughput numbers for the 78 GB FP8 checkpoint on H200, B200 and B300 parts.
Beyond that, attention turns to what Aleph Alpha builds next on the same pipeline, and to whether European public sector contracts follow. The company has shown it can train a 78.1 billion parameter model on 768 B200 GPUs in roughly 21 days of pre training. The harder question is whether sovereign AI can convert technical output into deployments at scale, and that answer will take longer than a single release.
Comments (0)
Log in or sign up to leave a comment.
No comments yet. Be the first to share your thoughts.