Infrastructure·
Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoE Models
Allen Institute for AI launches Olmo-core 3, a new open‑source training stack that scales mixture‑of‑experts models to trillion parameters with near‑linear GPU throughput.

Allen Institute for AI has released Olmo‑core 3, a redesigned training stack that makes mixture‑of‑experts (MoE) models practical at the trillion‑parameter scale.
The framework replaces the previous fully sharded data‑parallel approach with a distributed data‑parallel design that keeps expert weights resident on each GPU and routes only the relevant token batches to them. This eliminates the repeated weight‑gathering step that limited throughput on large clusters.
Why it matters for GPU / AI infrastructure
In a benchmark on eight NVIDIA B300 GPUs, a 47‑billion‑parameter MoE achieved 52,000 tokens per second per GPU — roughly 2.7× the throughput of the earlier FSDP‑based implementation — while expanding the expert pool from 8 to 128 without increasing active parameters per token.
Olmo‑core 3 combines expert parallelism, tensor parallelism, and pipeline parallelism with optimized routing kernels, enabling near‑linear scaling as the model grows. The open‑source release gives researchers and smaller labs a production‑grade path to train massive sparse models without proprietary tooling.
- aigpu
- ai gpu
- ai gpu cloud
- aigpu dubai
- moe training
- large language models
- gpu cluster
- open source
By AiGpu Editorial · Editorial rewrite based on public reporting (Hugging Face Blog)
← All articles