Ai2 has released Olmo-core 3, an open training framework designed to scale mixture-of-experts models towards the trillion-parameter range. Its headline result is a reported 2.7× throughput gain over Ai2’s earlier implementation in a preliminary eight-GPU test, alongside a stack other researchers can use and adapt.
Allen Institute for AI Watch analysis
What happened
Olmo-core 3 is a redesigned training system for large mixture-of-experts (MoE) models, which route each token through only some of a model’s expert components. Ai2 says the framework combines distributed data parallelism with expert and pipeline parallelism, plus changes to routing and computation, to tackle the memory and communication costs that can blunt the efficiency of sparse models.
In one benchmark, Ai2 increased the expert pool from 8 to 128 while selecting four experts per token. Total capacity rose from 4.6 billion to 47 billion parameters, while training throughput fell by less than 5%. In a separate preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU, compared with 19,400 using Ai2’s earlier FSDP-based implementation. These are Ai2’s reported system benchmarks, not evidence that a trained model became more capable.
Our top picks
- A bigger expert pool without a steep throughput penalty
Ai2 reports less than a 5% throughput drop as total capacity grows from 4.6B to 47B parameters. - A substantial gain against Ai2’s previous stack
In a preliminary eight-B300 test, the 47B model reached 52,000 tokens per second per GPU, versus 19,400 before. - A trillion-parameter configuration has been benchmarked
Ai2 reports testing a 1.2T-parameter model across 512 B300 GPUs, with 58.36B parameters active per token. - Lower-precision training with measured trade-offs
In a controlled four-B300 test, MXFP8 delivered about 21% higher throughput than BF16 and reduced peak active memory from 103 GiB to 95 GiB. - The infrastructure is open to other builders
Researchers and developers can use the framework to train MoEs and experiment with its routing and parallelism choices.
Why it matters
MoE models promise large total capacity without activating every parameter for every token. But distributing experts across a GPU cluster brings its own costs: memory, routing and communication can eat into the savings. A public training stack that reports how it handles those costs gives researchers something more useful than a diagram and a hopeful adjective: an implementation to inspect and benchmark.
The scale claims need their labels kept firmly attached. Ai2 says the 1.2T configuration used random routing to measure system performance, not the quality of a trained model; its 2.38T DeepEP v2 result was a short-capacity test, not a full training run. Impressive infrastructure numbers are not model-quality scores, however much the comma count tempts us.
Our read
This is a meaningful release for people building or studying large MoEs. The combination of open code, concrete scaling tests and candidly described failed optimisations makes it worth a closer look. Treat the speed figures as Ai2’s benchmark results, then look for independent runs and evidence that the gains hold across different hardware and workloads.
What to watch
- Whether external researchers reproduce the reported throughput and memory results.
- How the framework performs on sustained training runs, not just capacity or system benchmarks.
- What the next-generation Olmo model trained with this stack demonstrates about model quality.
- How the results compare with other open and established MoE training systems.
Discussion spark: For open AI development, which matters more: publishing infrastructure that others can run and adapt, or waiting for independent results showing it improves trained models?
Sources and evidence
- Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs (1 October 2026, 15:01 UTC)
Independent WittyWires tracker for public updates about Allen Institute for AI. Not affiliated with or endorsed by Allen Institute for AI; this is not an official account.