How to Use Xiaomi MiMo v2.6 Pro on AIHubMix: Setup Guide

AIHubMix6 min read
How to Use Xiaomi MiMo v2.6 Pro on AIHubMix: Setup Guide

In short: point the official openai Python SDK at https://aihubmix.com/v1, use model ID mimo-v2.6-pro, and authenticate with your AIHubMix key. The API is OpenAI-compatible (callable in the same format as OpenAI's), so no special client is required. Setup takes about five minutes.

Key points

  • OpenAI-compatible API — change base_url, nothing else
  • Paid mimo-v2.6-pro · free xiaomi-mimo-v2.6-pro-free [1] [2]
  • $0.48/M input · $0.96/M output [1]
  • Free tier 5 requests/min · 100/day · 1M tokens/day [2]
  • Context 1.05M tokens · max output 131K tokens [1]
  • Requirements: Python 3.9+, an AIHubMix API key

Reference values

Base URLhttps://aihubmix.com/v1
ProtocolOpenAI Chat Completions
Paid model IDmimo-v2.6-pro
Free model IDxiaomi-mimo-v2.6-pro-free
Pricing (paid)$0.48/M input · $0.96/M output · $0.0038/M cache read [1]
Free tier limits5 req/min · 100 req/day · 1M tokens/day [2]
Context window1.05M tokens [1]
Max output131K tokens [1]
Input modalitiesText, vision, audio, video — text output only [1]
Listed featuresStreaming, tool calling, structured outputs, prompt caching [1]

Should you use v2.6 Pro or v2.5 Pro?

Use v2.6 Pro. It is the same price and accepts more kinds of input.

mimo-v2.5-promimo-v2.6-pro
InputText only [5]Text, vision, audio, video [1]
Price in / out$0.48/M · $0.96/M [5]$0.48/M · $0.96/M [1]
Context / max output1.05M · 131K [5]1.05M · 131K [1]
AvailabilityRetires 21 October 2026 [5]Current [1]

Xiaomi's release notes say V2.6 "adopts the same API pricing as the V2.5 series" and give V2.6-Pro 46 points on the Artificial Analysis Intelligence Index — the top open-source score at the time of writing, above Kimi K3 and Qwen3.8 Max and below Claude Fable 5.1 and GPT-6 Astra [4]. The clearest generation-over-generation line in the notes is about the cheaper model: V2.6-Flash "has comprehensively outperformed MiMo-V2.5-Pro" [4].

Two things to keep in mind. Xiaomi has not published a direct V2.6-Pro vs V2.5-Pro benchmark table, so the comparison above is specs and price. And mimo-v2.5-pro is scheduled to retire on 21 October 2026 with mimo-v2.6-pro named as the migration target [5], so anything new built on v2.5 would need moving almost immediately.

Why AIHubMix

  • You only change base_url. The gateway speaks the OpenAI Chat Completions format, so the official SDK works unmodified. Anthropic-format /v1/messages (beta) and native Gemini calls are available on the same key [6].
  • One key covers many vendors. GET /v1/models returned 416 model IDs on the day of setup, 21 of them MiMo variants [3]. Changing model means changing one string in your script.
  • You can test for free first. xiaomi-mimo-v2.6-pro-free costs $0 within 5 requests/min, 100/day and 1M tokens/day [2] — enough to prove the setup works before you spend anything.
  • No subscription. Billing is pay-as-you-go with no membership or monthly fee [6], and the paid MiMo rate matches the model page at $0.48/M in and $0.96/M out [1].
  • Debugging is faster. If a call fails, point the same script at another vendor's model. A success there means your key, URL and request shape are fine, and the problem is the model — one call instead of an afternoon.

How to set it up

1. Create the project

mkdir mimo-demo && cd mimo-demo
python3 -m venv .venv
.venv/bin/pip install "openai>=1.40" "python-dotenv>=1.0"

2. Store your key

printf 'AIHUBMIX_API_KEY=your_key_here\n' > .env
chmod 600 .env
printf '.env\n.venv/\n__pycache__/\n' > .gitignore

Never hardcode the key in your script, and never commit .env.

3. Create the client

Save as chat.py:

import os
import sys

from dotenv import load_dotenv
from openai import OpenAI

load_dotenv()

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

prompt = sys.argv[1] if len(sys.argv) > 1 else "Introduce yourself in one sentence."

stream = client.chat.completions.create(
    model="mimo-v2.6-pro",
    messages=[{"role": "user", "content": prompt}],
    max_tokens=1024,
    stream=True,
)

for chunk in stream:
    if not chunk.choices:
        continue
    piece = chunk.choices[0].delta.content
    if piece:
        print(piece, end="", flush=True)
print()

Keep both checks inside the loop. Some chunks arrive with an empty choices array, which raises IndexError if indexed; delta.content is None on the first and last frames, which print would output as the literal text None.

4. Run it

.venv/bin/python chat.py "Introduce yourself in one sentence."
I'm MiMo, a large language model developed by Xiaomi's LLM Core Team, here to assist with answering questions, writing, and a wide range of helpful tasks.

Expect roughly six seconds to the first tokens, against the 6.4s listed on the model page [1].

5. Adjust as needed

For a non-streaming call, drop stream=True and read response.choices[0].message.content. To switch models, change the model argument. Raise max_tokens for longer output — the default of 1024 above is well below the model's limit [1].

Checklist

  • [ ] Virtualenv created; openai and python-dotenv installed
  • [ ] .env contains your real key, chmod 600
  • [ ] .env listed in .gitignore
  • [ ] base_url is https://aihubmix.com/v1 — the /v1 is required
  • [ ] Model ID is mimo-v2.6-pro (paid) or xiaomi-mimo-v2.6-pro-free (free)
  • [ ] Streaming loop guards chunk.choices and delta.content
  • [ ] A test call returns text

FAQ

Is the API really OpenAI-compatible? Yes — standard Chat Completions. Any OpenAI-compatible client works by setting the base URL to https://aihubmix.com/v1 and passing your AIHubMix key as the bearer token. AIHubMix's documentation notes that coding-xiaomi-mimo-v2.6-flash supports OpenAI-compatible formats only [3].

Which model ID should I use? mimo-v2.6-pro for paid, xiaomi-mimo-v2.6-pro-free for free. The two tiers use different naming conventions — free variants carry the xiaomi- prefix and paid ones do not, which is easy to get wrong.

What other variants are available? mimo-v2.6-pro-ultraspeed, mimo-v2.6-flash, mimo-v2.5-pro, and coding-tuned IDs such as coding-xiaomi-mimo-v2.6-pro [3]. List them from the API:

curl -s https://aihubmix.com/v1/models \
  -H "Authorization: Bearer $AIHUBMIX_API_KEY" \
  | python -c "
import json,sys
for i in sorted(m['id'] for m in json.load(sys.stdin)['data']):
    if 'mimo' in i: print(i)
"

I get 429 — what does it mean? It depends on the message, and the two cases need opposite handling. "rate limited by provider" is a per-minute throttle: wait and retry. "reached the limit of the free model quota" means your account's free quota is exhausted and shared across free models — retrying will not help, so switch to mimo-v2.6-pro.

I get 400 "cannot be served at the moment". First confirm the ID against the API:

curl -s https://aihubmix.com/v1/models \
  -H "Authorization: Bearer $AIHUBMIX_API_KEY" \
  | python -c "import json,sys; print('mimo-v2.6-pro' in [m['id'] for m in json.load(sys.stdin)['data']])"

If that prints True, repeat the chat call with curl and look for "code": "no_available_channel". That means the model is listed but has no upstream capacity at that moment — a provider-side condition that no local change will fix. Wait, or contact AIHubMix support with the request ID from the response.

How can I tell whether a problem is mine or the provider's? Call a different vendor's model with the same key and script — for example gpt-4o-mini. If it succeeds, your key, base URL and request format are correct and the problem is that specific model. If it fails too, check your key and billing.

Why is my output cut off mid-sentence? You reached max_tokens, not a model limit. Raise it.

How do I check my account status?

curl -s https://aihubmix.com/v1/dashboard/billing/subscription \
  -H "Authorization: Bearer $AIHUBMIX_API_KEY"

Look at has_payment_method and the limit fields.

Next steps

The script above is a working foundation. MiMo v2.6 Pro's listed capabilities include tool calling, structured outputs, vision and audio input, and prompt caching [1] — all through the same OpenAI-compatible interface, so extending the client is mostly a matter of adding standard parameters. Start with the free tier to confirm your setup, then switch the model ID to mimo-v2.6-pro for regular use.


Sources