The benchmark tests scaling infrastructure, not model quality

Ai2’s Olmo-core 3 release describes an open training stack for large mixture-of-experts models. In one benchmark, the team increased the pool from 8 to 128 experts while selecting four experts per token. It reports model capacity rising from 4.6 billion to 47 billion parameters with less than a 5% training-throughput drop.

An MoE model stores many specialized components but activates only some for each token. That can reduce computation per input, but it adds memory and communication costs across GPUs. The benchmark addresses those training-system costs; it is not a test showing that a trained 47B model is more capable.

A second result compares throughput on eight B300 GPUs

Ai2 reports that a 47B-parameter MoE processed 52,000 tokens per second per GPU in a preliminary test on eight NVIDIA B300 GPUs, compared with 19,400 tokens per second using its earlier implementation. The post attributes the gain to a redesigned stack built around distributed data parallelism and experts that remain resident on GPUs.

The system has also been benchmarked above one trillion total parameters, according to Ai2. That demonstrates infrastructure scale in a test environment; it does not mean Ai2 has released a trained trillion-parameter model.

Why open training tools matter

Training infrastructure is a barrier for university labs and small research groups as much as model weights are. By releasing code and technical details, Ai2 is offering a framework others can examine, reproduce and compare with existing MoE systems.

The useful next step is independent reproduction under clearly described hardware and workloads. Until then, the report is a promising engineering result from its authors, with the practical claim focused on throughput and capacity rather than a general intelligence jump.

Explore the original source ↗

Source published October 1, 2026. Coverage is based on the maker’s announcement and demonstration.