
Native omnimodal agentic model with 1M-token context and 15B active parameters.
About
MiMo-V2.5 is Xiaomi's natively omnimodal open-source model, a 310B-parameter Mixture-of-Experts architecture with 15B active parameters. It pairs a 729M-parameter vision transformer using hybrid window attention with a 1M-token context window, handling text and image understanding in a single unified model without a separate vision adapter.