MiMo-V2.6 Pro and Flash launch: access, prices, and limits

Xiaomi's MiMo-V2.6 Pro and Flash offer multimodal understanding and open weights. Compare hosted access, API costs, UltraSpeed, and self-hosting limits.

Quick answer

Xiaomi announced MiMo-V2.6 Pro and Flash on September 22, with hosted services and downloadable model weights. Both understand text, images, video, and audio, with a documented one-million-token context window. For a reader choosing an AI app, the important distinction is what the chosen interface actually exposes. This is a source-based release explainer, not a hands-on model ranking.

Download Chat AI Opens the official App Store or Google Play for your device.

Pro and Flash target different workloads

Xiaomi positions Pro for complex, extended projects and Flash for frequent calls and larger volumes of work. Its release emphasizes reinforcement learning across several task types. That describes how the models were trained; it does not establish that a normal chat session retrains the model or makes it continuously improve for you.

Sources: Xiaomi MiMo, Xiaomi MiMo

Where can you try MiMo-V2.6?

MiMo Studio's September 21 version 3.4.0 notes list Pro and Flash in MiMo Chat for conversation and visual understanding. The September 22 model announcement also names Xiaomi's API and Desktop client; Desktop supports a membership route or configuration with your own API key.

  • For ordinary chat, inspect the model selector and supported attachments in MiMo Studio before starting a task.
  • For an integration, Xiaomi lists the API identifiers mimo-v2.6-pro, mimo-v2.6-flash, and mimo-v2.6-pro-ultraspeed.
  • A provider release does not mean every third-party app has added the model. Check the exact version, plan, and tools in the service you intend to use.

Sources: Xiaomi MiMo, Xiaomi MiMo

Multimodal understanding is not every kind of output

The model cards list text, image, video, and audio capabilities. Xiaomi's API catalog specifies text generation, a 1M-token context window, and a maximum output of 128K tokens for Pro and Flash. Speech synthesis is documented separately under the MiMo-V2.5-TTS family. Do not assume that selecting V2.6 in a chat interface also enables spoken replies or video generation.

Sources: Xiaomi MiMo on Hugging Face, Xiaomi MiMo on Hugging Face, Xiaomi MiMo, Xiaomi MiMo

Flash costs less on the API; UltraSpeed has separate rates

Xiaomi's overseas real-time API pricing, checked September 22, lists the following US-dollar rates per million tokens. These are usage charges, not consumer app subscription prices.

  • Pro: $0.435 for uncached input, $0.0036 for cached input, and $0.87 for output.
  • Flash: $0.14 for uncached input, $0.0028 for cached input, and $0.28 for output.
  • Pro UltraSpeed: $4.35 for uncached input, $0.036 for cached input, and $8.70 for output.

Sources: Xiaomi MiMo

Downloadable weights do not make this a phone-local model

Xiaomi publishes MiMo-V2.6-Pro-RL and MiMo-V2.6-Flash-RL checkpoints on Hugging Face, with both model cards marked MIT. Pro's card describes 1.02 trillion total parameters and 42 billion activated per token; Flash lists 309 billion total and 15 billion activated. The smaller activated count is not the size of the complete model you need to host.

Sources: Xiaomi MiMo on Hugging Face, Xiaomi MiMo on Hugging Face

A small exercise before trusting a mixed-media summary

If your chosen service accepts both an image and an audio input, try this fictional exercise. It is an editorial proposal with a hand-written reference answer, not a recorded MiMo or Chat AI test. A text-only version can check conflict handling but cannot test image or audio understanding.

  1. Create a simple image saying: Workshop — Tuesday, 2 pm, Room 4.
  2. Record a short voice note saying: Correction to the workshop notice: it is now Wednesday at 3 pm. The room stays the same. Registration has not opened.
  3. Ask the model to report the current day, time, room, and registration status, identifying which source supports each detail. Ask it to mark missing information instead of guessing.

Frequently asked questions

What readers usually ask

Does the one-million-token context window guarantee a complete summary?

No. Context capacity is not an accuracy guarantee. Check whether the interface accepted the full input, then verify important details against the original material.

Does open-weight MiMo-V2.6 mean the hosted API is free?

No. Downloadable checkpoints and hosted inference are different routes. Xiaomi publishes separate API usage rates, and self-hosting requires computing resources.

Are MiMo-V2.6 Pro and Flash verified in Chat AI?

No. Neither exact model is verified in Chat AI's published catalog. This article explains Xiaomi's release and does not promise access through Chat AI.

Evidence

Sources

  1. MiMo-V2.6 release announcementXiaomi MiMo · Primary source
  2. MiMo Studio version 3.4.0 release notesXiaomi MiMo · Primary source
  3. API model capabilities and limitsXiaomi MiMo · Primary source
  4. Audio understanding inputs and restrictionsXiaomi MiMo · Primary source
  5. Pay-as-you-go API pricingXiaomi MiMo · Primary source
  6. MiMo-V2.6-Pro-RL model cardXiaomi MiMo on Hugging Face · Primary source
  7. MiMo-V2.6-Flash-RL model cardXiaomi MiMo on Hugging Face · Primary source