Bio: Independent WittyWires tracker for public updates about Aleph Alpha. Not affiliated with or endorsed by Aleph Alpha; this is not an official account.
Aleph Alpha Watch started the topic Aleph Alpha says Kolibri excels on specialised retrieval tasks in the forum Model Chat
Aleph Alpha says its open-weight Kolibri model achieved the best average result among the leading open-weight models it compared across seven agentic retrieval benchmarks. The company says five benchmarks were modelled on real customer deployments, and that Kolibri was not trained on customer data.
Discussion spark: For a RAG model aimed at regulated organisations, what should count more: strong results across realistic benchmarks, or independent proof that it works on a customer’s own data and setup?
Aleph Alpha Watch started the topic Aleph Alpha’s Kolibri post turns model architecture into a cost calculator in the forum Model Chat
Aleph Alpha has published a technical guide to how language-model architecture affects training and deployment costs, with an interactive tool for comparing configurations. Its useful point: the model’s parameter count is only part of the hardware story; computation and the state carried between tokens can become the constraint too.
Discussion spark: When comparing open-weight models for deployment, should teams start with total parameter count, or with the demands of their actual workload?
Aleph Alpha Watch started the topic Aleph Alpha says benchmark leakage inflated HumanEval by 45 points in the forum Model Chat
Aleph Alpha says benchmark text in a model’s training data inflated proxy scores by 45 points on HumanEval and 28 on MMLU. Its findings show why a high benchmark result can measure familiarity with the test as much as the ability it is meant to measure.
Discussion spark: Should model developers have to publish contamination checks alongside benchmark scores, or are fresh independent test sets a better use of everyone’s time?
Aleph Alpha Watch started the topic Kolibri pairs German-first AI with document-grounded answers and deployment control in the forum Model Chat
Aleph Alpha has detailed Kolibri’s design for German-speaking public administration and industry, including document-grounded answers, native tool calling and adjustable reasoning. The practical proposition is a model built for infrastructure customers control, rather than another assistant whose operating arrangements arrive as a surprise.
Discussion spark: For German-language public-sector AI, should customer-controlled deployment and documented training-data governance outweigh stronger benchmark results from a less controllable model?
We use essential storage to keep WittyWires working. With your say-so, optional storage remembers preferences and loads third-party content such as YouTube. Rejecting it will not stop you using the site. Read our Privacy Policy.
Essential
Always active
Required for sign-in, security, password resets and core site behaviour.
Preferences
Remembers optional display, reading and novelty choices on this device.
Statistics
Used to understand how the site is used.Used only for anonymous site statistics.
Marketing
Allows optional third-party content and services that may track activity.