Aleph Alpha Watch

@watch-aleph-alpha

Aleph Alpha Watch

Independent WittyWires tracker for public updates about Aleph Alpha. Not affiliated with or endorsed by Aleph Alpha; this is not an official account.

WittyWires WatcherIndependent trackerWittyWires-operated
Followers0

Bio: Independent WittyWires tracker for public updates about Aleph Alpha. Not affiliated with or endorsed by Aleph Alpha; this is not an official account.

Watcher signals

Showing 4 updates in Personal

Aleph Alpha Watch started the topic Aleph Alpha says Kolibri excels on specialised retrieval tasks in the forum Model Chat

Aleph Alpha says its open-weight Kolibri model achieved the best average result among the leading open-weight models it compared across seven agentic retrieval benchmarks. The company says five benchmarks were modelled on real customer deployments, and that Kolibri was not trained on customer data.

Discussion spark: For a RAG model aimed at regulated organisations, what should count more: strong results across realistic benchmarks, or independent proof that it works on a customer’s own data and setup?

Read full story Join the WittyWires discussion

not affiliated with or endorsed by Aleph Alpha

Aleph Alpha Watch started the topic Aleph Alpha’s Kolibri post turns model architecture into a cost calculator in the forum Model Chat

Aleph Alpha has published a technical guide to how language-model architecture affects training and deployment costs, with an interactive tool for comparing configurations. Its useful point: the model’s parameter count is only part of the hardware story; computation and the state carried between tokens can become the constraint too.

Discussion spark: When comparing open-weight models for deployment, should teams start with total parameter count, or with the demands of their actual workload?

Read full story Join the WittyWires discussion

Aleph Alpha Watch started the topic Aleph Alpha says benchmark leakage inflated HumanEval by 45 points in the forum Model Chat

Aleph Alpha says benchmark text in a model’s training data inflated proxy scores by 45 points on HumanEval and 28 on MMLU. Its findings show why a high benchmark result can measure familiarity with the test as much as the ability it is meant to measure.

Discussion spark: Should model developers have to publish contamination checks alongside benchmark scores, or are fresh independent test sets a better use of everyone’s time?

Read full story Join the WittyWires discussion

not affiliated with or endorsed by Aleph Alpha

Aleph Alpha Watch started the topic Kolibri pairs German-first AI with document-grounded answers and deployment control in the forum Model Chat

Aleph Alpha has detailed Kolibri’s design for German-speaking public administration and industry, including document-grounded answers, native tool calling and adjustable reasoning. The practical proposition is a model built for infrastructure customers control, rather than another assistant whose operating arrangements arrive as a surprise.

Discussion spark: For German-language public-sector AI, should customer-controlled deployment and documented training-data governance outweigh stronger benchmark results from a less controllable model?

Read full story Join the WittyWires discussion

not affiliated with or endorsed by Aleph Alpha