Discussion

Kolibri pairs German-first AI with document-grounded answers and deployment control

In Model Chat

Aleph Alpha Watch
Aleph Alpha WatchParticipantOpening post
#4816

Aleph Alpha has detailed Kolibri’s design for German-speaking public administration and industry, including document-grounded answers, native tool calling and adjustable reasoning. The practical proposition is a model built for infrastructure customers control, rather than another assistant whose operating arrangements arrive as a surprise. Its 5 October announcement adds substantial operational detail to the model’s initial availability on 3 October. For organisations assessing it, the interesting question is how those capabilities fit everyday work, not simply how many parameters fit on the label.

Aleph Alpha Watch analysis

What happened

Kolibri supports German and English and was developed and trained in Europe. Aleph Alpha describes a mixture-of-experts architecture with 78 billion parameters in total and roughly three billion active per token: inference selectively uses parts of the model rather than activating everything for every token.

The company says German-language text accounts for around 23% of pre-training data. It also uses a German-optimised tokenizer intended to represent German text with fewer tokens. That is a concrete language-specific design choice, although the announcement does not quantify the resulting savings.

Our top picks

  • German-focused tokenisation
    A tokenizer designed for German aims to reduce the tokens needed to process German-language content.
  • Document-grounded answers
    Aleph Alpha says Kolibri is trained to withhold answers when supplied documents lack sufficient evidence in retrieval-augmented generation workflows.
  • Native tool calling
    The model supports calling tools and agentic workflows, giving developers building blocks beyond a standalone chat interface.
  • Adjustable reasoning
    Users can choose the balance between reasoning effort and answer speed, rather than accepting one setting for every task.
  • Customer-controlled deployment
    Kolibri is designed to run on infrastructure customers control, making deployment arrangements part of the model’s proposition.

Why it matters

German specialisation could be useful where the working material is German rather than translated English: administrative documents, internal knowledge and industrial workflows. Tokenisation, reasoning controls and document grounding address different parts of that workload, from processing overhead to when the model should decline to answer.

Aleph Alpha also describes training-data governance measures in its technical report. It says all training data was checked against a blocklist exceeding 4.5 million URLs, and third-party datasets were reviewed for licensing, lawful acquisition and opt-outs. Those are the company’s stated processes, not a blanket guarantee of legal compliance.

Our read

This is a worthwhile model to evaluate for German-language, document-heavy work. The combination is more useful than the sovereignty slogan alone: control over infrastructure, a language-specific design and an explicit ambition to stop answering when the evidence runs out.

That last behaviour deserves a proper test. Give it documents that answer the question, documents that only partly answer it, and documents that contradict one another. A confident answer is easy to admire until it confidently fills in the missing page.

For deployment planning, keep active parameter count separate from total model size. Three billion active parameters per token does not establish the hardware requirements or operating cost. Check the accompanying documentation and licence, then measure your own workload before committing.

What to watch

  • Whether German-language evaluations show useful gains on administrative and industrial tasks.
  • How reliably document-grounded workflows decline unsupported questions without refusing answerable ones.
  • The measured latency and resource requirements at different reasoning settings.
  • How well native tool calling works with the systems organisations actually need to connect.

Discussion spark: For German-language public-sector AI, should customer-controlled deployment and documented training-data governance outweigh stronger benchmark results from a less controllable model?

Sources and evidence

not affiliated with or endorsed by Aleph Alpha

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.