A Sovereign Open-Weight Model — Aleph Alpha



Research

Aleph Alpha

On the Day of German Reunification, we’re releasing our new mannequin: Kolibri.

Kolibri is an English-German Mixture-of-Experts Transformer with 78B whole parameters, 3B
energetic. It helps context lengths of as much as 1M tokens. The mannequin could be downloaded with the
full weights on Hugging Face and used beneath the Apache 2.0 license phrases.

Kolibri is the results of steady iteration of our mannequin coaching effort. We first constructed a model training pipeline and validated it by constructing Kolibri Origin, a 30B whole, 3B energetic mannequin with a a lot shorter
65k token context window. Kolibri ran via the identical pipeline: from information ingestion and curation,
via ablations, pre-training, and post-training, to the ultimate evals. It enabled operating a whole lot
of ablation experiments and secure pre-training that ran with out a individual having to step in when
{hardware} failed or a knowledge connection dropped. We repeatedly monitored coaching metrics and standardized
monitoring for customized benchmarks. The time we put into constructing and iterating on this pipeline
was a useful asset placement. We see it in how a lot better Kolibri is than Kolibri Origin, and in
how little time separates their releases.

Kolibri is a specialised language mannequin constructed for sovereign mission-critical work in
regulated areas together with public administration, industrials and aerospace. We specialised
Kolibri for German, reasoning, math, agentic conduct, and additional capabilities our
prospects want in manufacturing. The intention of this specialization was to optimize efficiency in
our prospects’ particular use circumstances. Through specialization, prospects obtain contextualized
efficiency of their AI operations and so they can monitor its financial impression, in order that ROI
stays measurable and grows over time.

Specialization alone shouldn’t be sufficient. Sovereignty is simply as necessary. Sovereignty, for us,
combines two dimensions: how we constructed the mannequin, and the way it transfers to our prospects. We
supply full supply-chain integrity and account for each resolution, from information ingestion,
via pre- and post-training, to the ultimate evaluations. We present transparency. Customers
have full freedom of deployment and intellectual-property security, so compliance comes as an
inherited property of the mannequin.

Read our tech report for full particulars.

What Kolibri Delivers

We optimized Kolibri for efficiency throughout a variety of sectors contemplating their
explicit domain-specific language, regulatory, and procedural realities. Its small and
environment friendly measurement offers our prospects with flexibility to run it effectively on-premise,
with out sending inside information to third-party inference companies. The highlight on this
part introduces the mannequin’s capabilities, earlier than we describe them in part How we built Kolibri at high velocity.

Foundational capabilities for enterprise and authorities

With Kolibri we optimize the trade-off between mannequin functionality and deployment prices, utilizing
3B energetic parameters out of 78B whole. Kolibri sits on the Pareto frontier for high quality
versus serving price, for each English and German. The Pareto frontier is an idea from
economics, marking one of the best achievable combos of two targets, the place bettering on
one means giving up a few of the different. None of the in contrast fashions delivers extra high quality at
the identical serving price, or the identical high quality at decrease price.

Average benchmark rating [%]
  • Kolibri
  • Kolibri Origin
  • Other post-trained fashions
  • Pareto frontier

Performance vs. throughput for post-trained fashions in English (left) and German (proper). Metrics present unweighted common benchmark scores towards decoded textual content per second and GPU. Higher and additional proper is healthier.

Across math, coding, grounding, and long-context duties, Kolibri matches fashions with as much as
4 occasions its energetic parameter rely, resembling Nemotron 3 Super.

AIME 2025MathAIME 2025 (DE)MathAIME 2026MathAIME 2026 (DE)MathGPQA (diamond)KnowledgeGPQA (diamond, DE)KnowledgeAA-Omniscience IndexGrounding / hallucinationspublic set; from −100 to 100BrowseCompAgenticτ³-bench bankingAgenticτ²-bench retailAgenticτ²-bench airlineAgenticτ²-bench telecomAgenticBFCL v4 totalAgenticStayCodeBench v6CodeHumanEval+CodeLongBench ProLong contextAA-LCRLong context
Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3.6-35B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
AIME 2025 96.9 81.9 84.6 91.7 79.8
AIME 2025 (DE) 87.5 73.5 82.9 85.6 72.3
AIME 2026 96.0 81.5 91.0 90.4 83.1
AIME 2026 (DE) 90.0 75.2 84.4 87.5 78.5
GPQA (diamond) 84.3 68.1 83.4 78.0 74.7
GPQA (diamond, DE) 81.3 58.5 80.6 76.6 72.9
AA-Omniscience Index -32.8 -64.0 -15.3 -36.5 -24.0
BrowseComp 29.4 4.4 26.9 29.1 –
τ³-bench banking 38.1 5.7 10.6 15.5 5.7
τ²-bench retail 69.9 58.5 71.6 67.5 62.9
τ²-bench airline 76.7 58.7 70.7 72.7 40.0
τ²-bench telecom 94.7 67.5 99.1 68.1 41.5
BFCL v4 total 61.4 36.4 67.2 61.0 58.0
StayCodeBench v6 85.9 59.2 82.5 82.0 71.2
HumanEval+ 92.7 76.8 92.8 94.7 92.8
LongBench Pro 64.5 – 70.8 62.9 56.4
AA-LCR 68.3 – 69.7 67.0 52.3

Foundational capabilities of Kolibri. Kolibri is a balanced generalist mannequin with aggressive efficiency throughout math, code, lengthy context, agentic capabilities, and data. Benchmark scores are proven on a shared 0–100 scale, greater is healthier.

Contextualized efficiency for real-world purposes

Public benchmarks fail to seize specialised sector wants, so we developed our personal inside
analysis suites for the verticals that matter to our prospects, such because the German public
sector, aviation, manufacturing and the automotive {industry}. Each suite mirrors the talents,
workflows, and edge circumstances required in these sectors, and paired artificial coaching
environments allow us to enhance Kolibri towards these evaluations with out ever coaching on
buyer information. Read extra in part Contextualized performance.

Score on inside customer-proxy benchmark
  • Automotive provider 0.72 → 0.99
  • Semiconductors 0.35 → 0.80
  • German public sector 0.54 → 0.75
  • Industrial drive expertise 0.31 → 0.60
  • Aerospace 0.14 → 0.59
  • one checkpoint, one eval
  • imply of that day
  • Kolibri Origin
  • Kolibri

Contextualized efficiency throughout successive post-training runs on inside customer-proxy benchmarks. Dots are single evaluations of coaching checkpoints, traces are the imply of every day’s evaluations. Performance climbs throughout all 5 verticals, greater is healthier.

Answers which might be grounded in your paperwork. We educated Kolibri with abstention
information and with our Merlin-Arthur protocol. As a consequence, it’s educated to say “I do not know” when
the reply is not within the context. Our prospects worth and request this function, so we repeatedly
observe and validate abstention accuracy. We present extra particulars within the Grounding part.

Native German and English mannequin. We developed a bilingual German/English tokenizer
and targeted on including organic German data all through the coaching means of the mannequin, in order that 21.3% of the pre-training tokens are German.
We used translation sparingly (6% total), since translated textual content tends to hold the cultural
fingerprint of its supply language. The result’s a mannequin that’s bilingual by design, not an
English mannequin that has learn some German. Read extra in sections German pre-training data and Specialized tokenizers.

Control and compliance by design: upholding our prospects’ sovereignty

We constructed Kolibri with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in thoughts from the bottom up, with copyright regulation being a spotlight of our work to reliable
expertise.

We are clear about our model weights and about curation of our training data, in order that choices behind growth are seen. Through its reasoning traces, it
turns into explainable how the mannequin got here to a selected reply. With our Merlin-Arthur protocol the mannequin grounds trustworthiness: Kolibri refrains from answering when
the context does not help a solution.

Our groups constructed the mannequin in Germany, educated it on transport systems in Germany and Finland,
beneath European and German regulation, with no international management. We personal all the
pipeline, from information curation, via pre- and post-training, to optimization in our Model
Factory. This management of our finish to finish pipeline underpins the mannequin’s sovereignty. Control
additionally passes on to our prospects and Kolibri’s small measurement offers them full freedom of
deployment and helps controllable reasoning effort to commerce price and latency towards
reply high quality.

How We Built Kolibri at High Velocity

Our Model Factory: two fashions, three months aside

Our Model Factory is our reply to iteration pace, minimizing the time it takes to go from
figuring out what a mannequin will get mistaken to coaching one which does higher. We carried out the
coaching pipeline as code, in order that the learnings of our crew landed in a single versioned
coaching recipe as a substitute of scattered scripts and notes. Here is what that appeared like
between Kolibri Origin and Kolibri.

Work on the pipeline started in January. Five months and a whole lot of ablation runs later,
Kolibri Origin completed pre-training at goal scale on 11 June. Kolibri completed on 11
September. In the three months between these two dates we went from 30B parameters to 78B,
from a 65k-token context window to as much as 1M, and from 7.5T coaching tokens to 20T. To get
these 20T, the pipeline processed over 200T tokens of uncooked information, filtering, deduplicating,
and curating it right down to what we truly educated on. We modified the eye design,
tripled the variety of consultants, elevated sparsity, changed the routing algorithm, improved
our post-training information, greater than doubled the variety of surroundings duties, and taught the
mannequin to motive at 4 totally different effort ranges. Both fashions saved an analogous variety of
parameters energetic per token, but Kolibri educated sooner per token than Kolibri Origin thanks
to the work of our effectivity crew.

Kolibri Origin Kolibri
Finished pre-training 11 June 2026 11 September 2026
Release no public launch 3 October 2026
Reasoning mode Yes (one mode solely) Yes (none, low, medium, excessive)
Total parameters 30.6B 78.1B
Active parameters / token 3.27B 3.46B
Pre-training tokens 7.51T 20T
Layers 50 (2 dense + 48 MoE, 1 shared knowledgeable) 50 (all MoE, 1 shared knowledgeable)
Pre-training context size 8,192 (8k) 16,384 (16k)
Longest educated size 65,536 (64k) 262,144 (256k)
Tokenizer vocabulary 96,000 128,000
Model dimension 2,048 2,560
Attention heads (question / KV) 32 / 4 48 / 4
Experts (whole / energetic) 128 / 8 384 / 6
Expert hidden dim 768 512
Attention sample full consideration, all layers sliding window (512) + full consideration each fifth layer
Knowledge cutoff EN: 1 Sept 2024, DE: 1 Aug 2025 EN/DE: 18 Jun 2026

What made this potential is our coaching pipeline, an effort to transform research-grade mannequin
growth into a totally automated production-ready transport systems to design and practice giant
language fashions. Our pipeline codebase was shared by all. Every proposed code change triggers a small end-to-end mannequin run; coaching, analysis,
to seek out out if one thing broke inside minutes. Our runs are GitHub Actions workflows, and
reproducing one means testing a commit. We used the pipeline for the principle coaching and
the a whole lot of ablations that ran to make our architectures and data-mix choices.

Training checkpoints land roughly each hour and the pipeline robotically evaluates them
on English and German data, maths, code, instruction following, software use, lengthy context,
security, abstention to hallucination and grounding. We can watch capabilities seem each
day quite than discovering out how the mannequin turned out on the finish. The run itself held up
higher than we anticipated. Over 21 days of pre-training of Kolibri, we hit 38 unplanned
interruptions, roughly one per 10,000 GPU-hours, attributable to {hardware} faults or a connection
timing out. Those had been robotically dealt with by the pipeline with out handbook intervention.
The cluster robotically restarted the coaching job(s) on a unique set of nodes, and
coaching picked up from a checkpoint at most 250 steps again.

Having entry to the entire pipeline, with checkpoints out there to everybody, signifies that any
crew can personal a functionality finish to finish quite than a stage of an meeting line. One crew
educated and evaluated the grounding functionality of Kolibri as a single piece of labor,
seamlessly integrating into the general mannequin.

However, getting right here was not a linear stroll. We stumbled. We stopped the Kolibri Origin
pre-training after just a few trillion tokens and restarted it from scratch, as we recognized a
data-shuffling bug that escaped our exams. We ran ablations, took choices, solely to discover a
bug or an error within the configuration afterwards, forcing us to rerun some experiments.
Alongside our personal expertise, we tracked state-of-the-art architectures, finest practices, and
the most recent analysis advances in LLMs. Those amassed learnings led to enhancements in our
processes and guardrails in our pipeline.

Looking on the delta between Kolibri Origin and Kolibri, the development is notable, however
that is additionally the straightforward path: extra parameters, extra information, a well-understood structure
household, and nonetheless at small scale. As we’re considering scaling up, we’re glad to face
challenges on our basis. But probably the most sturdy factor we constructed this yr shouldn’t be the
pipeline, it’s a crew with the confirmed functionality to construct, post-train, and ship LLMs from
uncooked information at excessive velocity.

Architecture and pre-training

Kolibri has 78B whole parameters, which is 2.5 occasions greater than Kolibri Origin, with ~3B
energetic parameters. In our experiments, growing the dimensions of our mannequin from 32B to 123B led
to ever-improved efficiency. Yet the bigger measurement got here with bigger coaching and serving
prices. The latter drove the choice: 123B can deal with solely 3 long-context 256k-token person
queries on two H100s, whereas 78B handles 18 concurrent requests and decodes 28% sooner. We
used 384 smaller consultants quite than fewer huge ones, as they carried out higher in our exams.
We utilized the identical efficiency-first reasoning to consideration. Out of fifty whole layers, solely 10
course of full context, whereas the remaining 40 use a decent 512-token targeted window. This
retains decode computation and reminiscence bounded in these layers no matter context size.
The effectivity good points profit each serving the educated mannequin and coaching throughout
post-training RL, which depends on large quantities of inference.

We educated Kolibri on 768 B200 GPUs in three phases: 20T tokens of pre-training at a 16k sequence size over 21 days, 3.44T tokens
of mid-training at 64k, and 200B tokens of long-context adaptation at 256k. This is sort of 24T
tokens in whole, roughly thrice what Kolibri Origin consumed. German accounts for greater than
a fifth of the pre-training combine, about 4.3T tokens, towards roughly 62% English and 14% code.
Compared to pre-training, mid-training information is a way more selective pool of curated datasets
weighted towards reasoning, problem-solving, code and agentic information. For long-context, quite than
practice on lengthy paperwork alone, which tends to erode the talents acquired earlier, we interleaved
lengthy paperwork with the high-quality mid-training information of the earlier stage. For the long-context
combine, we eliminated artificial lengthy paperwork to keep away from artificially inflating benchmarks resembling RULER.

We optimized with Muon, as we did for Kolibri Origin. We put particular give attention to coaching
stability, leading to a sturdy coaching run with none loss spikes for both mannequin. With
Kolibri, for routing between consultants throughout coaching, we introduce precise quantile balancing.
Quantile balancing was launched in Kimi K3, the place the worldwide quantile is estimated from
histograms as a result of an actual computation was thought-about too costly to speak; we present
that it may be computed precisely at fastened price unbiased of batch measurement, and that the
exactness improves each load stability and mannequin high quality.

Post-training at scale

Post-training occurred in two phases; step one is supervised fine-tuning to show the
mannequin core reasoning means and learn how to work together in a chat, adopted by large-scale
reinforcement studying to coach the mannequin to motive throughout a various suite of long-horizon
duties. We constructed a pipeline for each of those steps, which allowed us to train
fine-grained management over the conduct of the mannequin.

For SFT we generated a complete of 174B tokens price of artificial information, which we filtered for
high quality and mixed with filtered variations of permissively licensed open-source datasets to
get hold of a high-quality coaching mixture of 268B tokens. For reinforcement studying we educated on
a broad set of environments that contained greater than 1.2 million curated duties throughout
various domains resembling code & math reasoning, agentic duties, instruction following,
query answering, software calling, and extra. We optimized our in-house coaching codebase for
high-performance and use asynchronous coaching, the place we generate coaching information on our
environments utilizing the present mannequin in parallel to coaching the mannequin on already generated
information.

Across each phases, we taught the mannequin to motive at totally different effort ranges (none, low,
medium, excessive), which signifies that the person can train management over how a lot compute the
mannequin ought to make investments to discover a resolution to the duty at hand. This permits our prospects to
trade-off price and inference pace towards the standard of the ultimate reply.

German pre-training information

From our experiments on small proxy fashions, we discovered that coaching with round 20% German
information results in optimum outcomes, which meant we wanted to seek out 4T German tokens to coach our
mannequin at a 20T horizon. Open German datasets assist, however are removed from sufficient: after
deduplication and filtering, we had been left with 390B German tokens, nicely wanting our aim.
As we argued in Sauerkraut, Not Burgers, German functionality has to return from high-quality German texts; relying closely on
machine-translations would result in poor outcomes on account of refined translation errors and a scarcity
of genuine German cultural context. We closed the token hole in three other ways.

The first was to curate German from Common Crawl ourselves. We constructed a pipeline specialised
for German information. German shouldn’t be English, and German information can’t be filtered like English
information: we needed to retune filtering parameters for the German language. One typical filter in a
language information pipeline is to take away paperwork with too many lengthy phrases, however German
administrative prose routinely exceeds the English certain on imply phrase size, so the
normal settings quietly take away the register that public administration writes in. After
retuning, our German pipeline gave us 1.3T distinctive tokens of natural German internet.

The second was to rephrase German paperwork we already had. An LLM rewrites an natural
German doc within the fashion of an encyclopedia entry, a Q&A dialogue or a textual content passage,
preserving its content material. This teaches the mannequin the identical details in a number of floor types and
multiplies the knowledge contained in scarce information, and it’s a totally different operation from
translation: the supply is German, so the subject material and the cultural affinity keep
German: chancellor, not president. It doesn’t add a lot new data, quite new phrasings
of data that was already within the corpus. Rephrasing gave us about 1T distinctive tokens,
making it the only largest supply of German within the mannequin.

The third was translation, and this we solely utilized in Kolibri Origin. Translating English into
German works when the mannequin, the prompts and the chunking are chosen rigorously, but it surely
carries two issues. Output can nonetheless present translationese, the literal rendering of idioms:
“Drive secure!” turns into “Fahre sicher!” quite than “Komm intestine an!”. More importantly, cultural
context doesn’t translate. A corpus translated from English inherits the geographic,
demographic and institutional distribution of the English internet, so a mannequin educated on it
speaks German a couple of globe that appears American.

In the top German entered Kolibri as a 2.4T-token distinctive pool, 80% of it curated or
generated by us and 20% from open datasets, and the mannequin noticed it at 21.3% of pre-training
tokens, roughly 4.3T over the 20T run via upsampling. Each German token was seen 1.8
occasions on common, nicely contained in the four-epoch restrict previous which repetition stops paying off.
Around 85% of the German the mannequin learn is internet textual content, both natural or rephrased from
natural. The relaxation is curated paperwork – parliamentary proceedings, authorized texts and different
information within the public area – and that small translated share.

Specialized tokenizers for English-German

We constructed a bilingual English-German tokenizer, educated and specialised on every mannequin’s
pre-training dataset. Because of the 21% German share in our information, it compresses German
language higher than different SOTA fashions. This led to extra environment friendly inference (fewer tokens)
minimizing prices and shortening response occasions. We introduce a brand new method to practice tokenizers,
UniBPE, that respects the morphology of languages higher than present approaches,
particularly the compound construction of German, with out sacrificing English token effectivity.

German internet
(AdvantageousWeb-2)



  • Kolibri


    128,000 vocab





    4.90



  • Kolibri Origin


    96,000 vocab





    4.69



  • Plain BPE 128k


    128,000 vocab





    4.89



  • GPT-5


    200,019 vocab





    4.35



  • DeepSeek V4


    129,280 vocab





    3.72



  • Kimi K3


    163,586 vocab





    3.28



  • GLM 5.3


    154,856 vocab





    3.93



  • Qwen3-Next


    151,669 vocab





    3.59



  • Qwen3.5-3.8


    248,077 vocab





    4.17



  • Gemini


    262,144 vocab





    4.13



  • EuroLLM


    128,000 vocab





    4.08



  • Tekken (Mistral, Nemotron, Apertus)


    131,072 vocab





    4.03

English internet
(AdvantageousWeb)



  • Kolibri


    128,000 vocab





    4.58



  • Kolibri Origin


    96,000 vocab





    4.50



  • Plain BPE 128k


    128,000 vocab





    4.59



  • GPT-5


    200,019 vocab





    4.67



  • DeepSeek V4


    129,280 vocab





    4.59



  • Kimi K3


    163,586 vocab





    4.62



  • GLM 5.3


    154,856 vocab





    4.61



  • Qwen3-Next


    151,669 vocab





    4.52



  • Qwen3.5-3.8


    248,077 vocab





    4.47



  • Gemini


    262,144 vocab





    4.49



  • EuroLLM


    128,000 vocab





    4.16



  • Tekken (Mistral, Nemotron, Apertus)


    131,072 vocab





    4.45

Tokenizer compression in common bytes per token on German and English internet textual content. Kolibri achieves one of the best German compression on this comparability. Plain BPE 128k is normal BPE educated on the identical information with the identical settings as Kolibri, which isolates the impact of the coaching methodology. More textual content per token means fewer tokens per activity – greater is healthier.

The two main approaches to coach tokenizers are BPE and Unigram. We mix these two, holding
the bottom-up strategy of BPE and utilizing the Unigram coaching goal for choosing which
merge so as to add to the vocabulary. On a 128k vocabulary educated on our English/German dataset,
this considerably improves tokenization.

Bundessozialgerichtes
Federal Social Court (genitive)


  • Kolibri



    Bundes

    sozial

    gericht

    es


  • GPT-5



    Bund

    ess

    oz

    ial

    gericht

    es


  • Qwen3.8



    Bund

    ess

    oz

    ial

    gericht

    es


  • Gemini



    Bund

    ess

    oz

    ial

    gericht

    es


  • Mistral Medium 3.5 · Nemotron 3 Nano



    Bund

    ess

    oz

    ial

    gericht

    es

silkworm


  • Kolibri



    silk

    worm


  • GPT-5



    sil

    kw

    orm


  • Qwen3.8



    sil

    kw

    orm


  • Gemini



    sil

    kw

    orm


  • Mistral Medium 3.5 · Nemotron 3 Nano



    sil

    kw

    orm

Protokolldaten
log information


  • Kolibri



    Protokoll

    daten


  • GPT-5



    Pro

    tok

    ol

    ld

    aten


  • Qwen3.8



    Protokol

    ld

    aten


  • Gemini



    Protok

    ol

    ld

    aten


  • Mistral Medium 3.5 · Nemotron 3 Nano



    Pro

    tok

    ol

    ld

    aten

coprocessors


  • Kolibri



    co

    processors


  • GPT-5



    cop

    rocess

    ors


  • Qwen3.8



    cop

    rocess

    ors


  • Gemini



    cop

    rocess

    ors


  • Mistral Medium 3.5 · Nemotron 3 Nano



    cop

    rocess

    ors

How totally different tokenizers cut up the identical phrases. The Kolibri tokenizer follows the morphology of the language. English and German phrases cut up into significant items, whereas opponents minimize throughout morpheme boundaries. Lower is healthier: fewer, cleaner splits per phrase imply fewer tokens.

Grounding: lowering hallucinations

Common LLM coaching and analysis rewards guessing: a guess has some likelihood of touchdown the
right reply, whereas abstention has none. Therefore, fashions study to reply with no matter
they’ve, even when their enter doesn’t present sufficient info to warrant a response.
Asking a mannequin to offer citations typically backfires for a similar motive: fashions can merely
hallucinate plausible-looking sources to justify an ungrounded reply.

For a regulated buyer, a mannequin that is aware of to abstain is the distinction between a pilot
and a deployment. We deal with saying “I do not know” as an necessary mannequin functionality and (1)
develop and observe devoted grounding and anti-hallucination measures and (2) use and
develop devoted coaching procedures to enhance this abstention functionality.

During coaching, we use coaching information samples the place the right reply is “I do not know”.
While fine-tuning for abstention is gaining traction industry-wide, high-quality unfavourable
examples stay onerous to return by, particularly the place our prospects want them, in slender
verticals the place all information is scarce. Therefore, we additionally practice with our Merlin-Arthur process, developed in-house and defined intimately in our blog post, which solves information shortage by robotically exploiting the mannequin’s weaknesses at every
coaching step and generates artificial unfavourable examples from out there paperwork. This
coaching turns into a recreation with three gamers. Arthur is the mannequin we practice and ship. Merlin
takes an present doc context, generates a brand new one which will increase Arthur’s likelihood
of answering accurately. Morgana generates a brand new datapoint by stripping out the related
proof from the doc and tries to lure Arthur right into a hallucination. Arthur doesn’t
know which one he is going through, so his legitimate technique turns into to rigorously learn the query and
the context, to have the ability to reply on Merlin’s context, whereas recognizing that abstention is
the strictly required reply for Morgana’s redacted context – making any guess, even a fortunate
one, incorrect.

For measuring hallucination abstention, we use established benchmarks alongside our personal
proxies derived from buyer use circumstances. We’ve additionally developed our personal “M/A grounding rating”
which falls out of our Merlin-Arthur setup representing a decrease certain on how a lot of the
reply provably got here from the doc.

Kolibri hallucinates far lower than Kolibri Origin: it abstains as a substitute of answering mistaken on
44% of AA-Omniscience objects (Origin: 15%), and on RGB it holds again
extra typically (86% vs 74%) and invents fewer falsehoods (87% vs 76%). On our personal M/A grounding score, which certifies quite than estimates how a lot of a solution got here from the doc,
Kolibri reaches 0.23 the place Kolibri Origin, alongside another fashions, attain 0.

AA-Omniscience Non-Hallucination Ratepublic setshare not answered mistakenRGB: holds againwhen the paperwork do notreply the queryRGB: invents nothingno falsehoods when thepaperwork lack the replyFRAMESmulti-document reasoningM/A grounding ratingour personal metricaxis from 0 to 0.5
Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3.6-35B-A3B Qwen3-Next 80B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
AA-Omniscience Non-Hallucination Rate 44.0 14.8 56.7 12.3 13.9 34.7
RGB: holds again 85.6 73.9 79.6 81.3 74.6 82.3
RGB: invents nothing 87.3 75.6 84.3 83.9 86.0 87.0
FRAMES 71.2 65.7 74.7 68.9 74.9 71.9
M/A grounding rating 0.23 0.00 0.12 0.00 0.00 0.06

Contextualized efficiency

Most public benchmarks fail to seize specialised sector wants. To measure contextualized
efficiency – that means how nicely a mannequin handles the distinctive workflows and area data of
particular industries – we’d like measurements of efficiency in specialised verticals that are
necessary to our prospects, such because the German public sector, authorized, {hardware}, client
electronics, and automotive.

For this, we first constructed domain-specific analysis suites that mirror the talents,
tool-calls, workflows, and edge circumstances required in key sectors and thus seize
contextualized efficiency. Once this was achieved, we created coaching environments by
producing high-quality specialised artificial information capturing these required abilities. To
enhance the real-world robustness of the mannequin on these duties, we moreover randomized
the environments round concrete particulars (instruments, configurations, harnesses).

Together, these contributed to an iterative engine that allowed us to sharpen the mannequin’s
agentic RAG capabilities, domain-specific logic and capabilities, and to hill-climb the
contextualized evaluations with out ever coaching on buyer information.

These analysis suites run robotically as a part of our pipeline: each checkpoint is
scored on them because it lands. Kolibri compares nicely towards competing open-weight fashions. On
the agentic-RAG benchmark Honeypot, Kolibri outperforms all in contrast fashions, together with the
a lot bigger Nemotron 3 Super and Mistral Small 4. On a cleaned model of the agentic-RAG
benchmark MuSiQue, it’s second solely to the extra large Nemotron 3 Super and nicely forward of
the remaining. On the 5 buyer purposes, Kolibri leads or is on par with one of the best mannequin
in 4 of 5. These findings communicate to our means to deal with a variety of information and harness
setups, throughout use-cases, whereas the in contrast fashions are delicate to those particulars.

These evaluations for measuring contextualized efficiency can exist as a result of prospects inform
us how their purposes fall brief. Each one encodes what we discovered from these
conversations: what sort of paperwork matter, which instruments the mannequin ought to get, which
questions are onerous, the place the deployed system frustrates its customers at this time. That makes the
alternate concrete in each instructions. A buyer who exhibits us a failure case will get it turned
right into a benchmark that each future checkpoint is measured towards, and we will hill-climb the
goal with out ever coaching on buyer information or risking over-fitting.

MuSiQue (cleaned)Agentic RAGHoneypotAgentic RAGSemiconductorsCustomer proxyGerman public sectorCustomer proxyAerospaceCustomer proxyAutomotive providerCustomer proxyIndustrial drive expertiseCustomer proxy
Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3-Next 80B-A3B Qwen3.6-35B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
MuSiQue (cleaned) 77.3 42.7 50.5 61.2 79.1 66.8
Honeypot 80.8 25.3 13.5 74.3 68.8 68.1
Semiconductors 80.4 35.3 41.2 79.4 69.6 62.7
German public sector 75.0 54.0 29.5 72.0 78.0 50.0
Aerospace 58.9 14.1 48.1 59.0 54.9 47.0
Automotive provider 99.0 72.4 84.2 92.6 91.0 87.1
Industrial drive expertise 60.0 31.4 32.7 59.5 37.3 56.8

Benchmarks

Benchmarks had been run utilizing our own harnesses and, the place relevant, all fashions used the best respective reasoning effort.

Type MoE MoE MoE Dense
Active parameters 3B 4–6B 12B 27B · 70B
Benchmark
Kolibri

Kolibri Origin

GLM-4.7 Flash 30B-A3B

Nemotron 3 Nano 30B-A3B

Qwen3.5 35B-A3B

Qwen3.6 35B-A3B

Qwen3-Next 80B-A3B Thinking

Gemma 4 26B-A4B IT

GPT-OSS 120B

Mistral Small 4 119B-A6B

GLM-4.5 Air 106B-A12B

Nemotron 3 Super 120B-A12B

Qwen3.8 27B

Apertus 70B Instruct
Overall (EN) 75.5 54.1 64.7 65.6 74.7 71.4 62.4 71.9 72.3 63.1 64.4 73.0 80.2 –
Overall (DE) 70.8 46.4 50.4 59.3 69.8 67.3 58.0 66.3 70.2 61.4 64.8 67.9 79.9 –
Knowledge
Average (EN) 50.1 39.7 45.5 46.0 52.7 52.1 48.4 51.4 50.0 47.5 45.7 52.0 56.8 –
Average (DE) 57.6 44.9 46.5 41.2 61.3 61.0 55.7 61.5 58.0 51.4 52.2 59.5 69.2 –
GPQA Diamond (EN) 84.3 68.1 73.1 73.9 83.8 83.4 76.1 81.1 76.4 74.7 73.2 78.0 89.2 29.5
GPQA Diamond (DE) 81.3 58.5 59.8 49.6 84.2 80.6 72.2 80.1 76.0 72.9 71.1 76.6 88.1 31.4
Humanity’s Last Exam (EN) 21.5 9.4 15.4 12.1 20.4 21.1 11.6 19.2 19.4 9.7 8.7 20.6 35.6 5.2
Humanity’s Last Exam (DE) 15.9 10.4 9.1 13.1 18.1 20.5 15.6 23.4 20.7 10.5 10.5 22.3 37.2 5.7
AA-Omniscience Accuracy (public set) 14.8 11.3 17.0 19.5 22.0 19.5 24.2 20.7 23.3 25.0 20.0 26.7 17.5 13.5
AA-Omniscience Index (public set) −32.8 −64.0 −62.8 −45.7 −47.3 −15.3 −42.3 −47.3 −35.2 −24.0 −28.8 −36.5 −9.5 –
MMLU-Pro CoT (EN) 80.0 70.1 76.5 78.3 84.6 84.3 81.7 84.5 80.8 80.4 80.9 82.7 85.0 43.0
MMLU-ProX CoT (DE) 75.5 65.7 70.7 61.0 81.7 81.9 79.4 81.1 77.2 70.7 74.9 79.7 82.4 37.3
Math
Average (EN) 96.5 81.7 88.8 88.8 90.1 87.8 86.3 87.4 90.7 81.4 82.8 91.1 97.8 –
Average (DE) 88.8 74.3 45.2 84.3 79.6 83.7 83.7 88.1 90.8 75.4 80.6 86.5 96.7 –
AIME 2025 (EN) 96.9 81.9 89.4 89.6 88.1 84.6 84.2 87.3 90.6 79.8 81.9 91.7 97.9 0.6
AIME 2025 (DE) 87.5 73.5 43.8 84.4 76.7 82.9 80.6 88.1 90.6 72.3 80.6 85.6 96.5 0.2
AIME 2026 (EN) 96.0 81.5 88.3 87.9 92.1 91.0 88.5 87.5 90.8 83.1 83.8 90.4 97.7 0.6
AIME 2026 (DE) 90.0 75.2 46.7 84.2 82.5 84.4 86.7 88.1 91.0 78.5 80.6 87.5 96.9 0.0
Agentic
Average (EN) 63.4 41.6 58.9 46.4 63.4 62.1 46.3 54.6 54.0 40.7 53.5 54.9 66.7 –
TerminalBench 2.1 27.7 – 20.2 9.7 39.7 – 8.6 – 29.2 21.0 – 39.7 76.8 –
Tau2-Bench (Telecom) 94.7 67.5 95.9 45.9 97.7 99.1 43.9 45.3 73.1 41.5 53.8 68.1 82.5 10.8
Tau2-Bench (Retail) 69.9 58.5 57.9 64.9 70.8 71.6 60.8 71.3 60.5 62.9 61.4 67.5 68.7 9.6
Tau2-Bench (Airline) 76.7 58.7 68.7 52.7 76.0 70.7 65.3 73.3 72.7 40.0 70.7 72.7 83.3 40.0
Tau3-Bench (Banking) 38.1 5.7 7.2 5.7 11.3 10.6 5.4 16.0 14.7 5.7 6.4 15.5 50.0 2.1
BFCL v3 (multi-turn) 39.8 22.8 58.2 47.9 54.0 53.5 51.4 53.4 45.6 36.2 61.6 44.6 42.5 0.6
BFCL v4 (total) 61.4 36.4 65.4 61.5 70.5 67.2 51.0 68.2 57.3 58.0 67.2 61.0 73.2 –
BFCL v4 (non-live AST) 79.1 78.1 83.3 85.0 85.8 88.2 83.6 83.7 35.8 83.6 85.5 45.0 85.3 –
BFCL v4 (reside) 78.9 73.7 78.3 78.8 80.2 81.4 82.5 80.2 70.4 78.4 78.2 77.6 79.9 –
BFCL v4 (multi-turn) 47.5 27.5 62.7 53.5 59.9 58.1 56.0 61.4 55.4 40.4 65.2 51.7 55.5 –
BFCL v4 (reminiscence) 62.8 19.4 41.5 39.1 62.6 53.8 35.3 52.9 50.7 39.1 43.4 59.6 79.6 –
BFCL v4 (internet search) 62.5 10.5 69.0 66.0 75.0 68.5 12.5 75.0 57.0 69.0 71.0 71.5 82.0 –
BrowseComp 29.4 4.4 – 14.5 36.5 26.9 2.8 25.5 31.2 – – 29.1 46.4 –
Code
Average (EN) 89.3 68.0 67.8 81.8 85.0 87.7 83.6 89.0 90.8 82.0 79.8 88.3 94.2 –
StayCodeBench v6 85.9 59.2 46.5 71.3 77.8 82.5 73.9 82.3 87.5 71.2 67.8 82.0 93.8 8.7
HumanEval+ 92.7 76.8 89.0 92.4 92.2 92.8 93.3 95.7 94.1 92.8 91.8 94.7 94.7 41.6
SWE-Bench Verified 66.4 – 51.0 38.6 71.6 73.8 – 57.8 – 60.8 11.6 60.2 72.6 –
Instruction Following
Average (EN) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 –
IFBench (loose-prompt) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 25.5

Get Started

Our mannequin is accessible overtly on Hugging Face beneath an Apache 2.0 License.

Kolibri requires the aleph-alpha-inference bundle that gives the Kolibri
vLLM plugin. You can both use the supplied container picture ghcr.io/aleph-alpha/aleph-alpha-inference, or set up the bundle from Aleph-Alpha/aleph-alpha-inference, which additionally installs the vLLM model it helps:

pip set up "aleph-alpha-inference>=1.0"

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 
  --reasoning-parser kolibri1 
  --tool-call-parser kolibri1 
  --enable-auto-tool-choice

To serve contexts past 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'. The really useful sampling parameters for the mannequin are temperature=1.0, top_p=0.97 and top_k=128.

Contact for Deployment and Specialization

Contact our crew who can be glad to help you thru our enterprise deployment and
specialization choices: contact sales.



Source link