Our website uses necessary cookies to enable basic functions and optional cookies to help us to enhance your user experience. Learn more about our cookie policy by clicking "Learn More".
Accept All Only Necessary Cookies
NEWSREVIEWSHOWTO

Xiaomi MiMo-V2.6 Pro: Every Benchmark and Model Ranked

Candida Corkery

Xiaomi MiMo-V2.6-Pro tops the open weights leaderboard at 46 on the Artificial Analysis Index, with 1T parameters, MIT license, and $0.13 per task.

Catelog

    Xiaomi released MiMo-V2.6 on September 21, and the Pro model walked straight to the top of the open weights leaderboard. It scored 46 on the Artificial Analysis Intelligence Index, a 20-point jump over its own predecessor, and it did that at a price that undercuts nearly everything around it. If you run models locally or route work through an API, this is the release worth understanding.

    Ranking a model family is trickier than ranking a phone. You can't just order the numbers, because a model that wins one benchmark loses another. So this list sorts by what each result means for a real workload: coding, agent tasks, cost, and the license that decides whether you can actually self-host.

    What actually matters when you're choosing? Cost and the license, more often than raw scores.

    The criteria here: only numbers we can verify through Artificial Analysis or Xiaomi's own technical report, plus the cost data that determines practical use. No vendor marketing claims without a source. Here's how MiMo-V2.6 stacks up across the dimensions that matter.

    Xiaomi MiMo-V2.6 Pro Benchmarks, Ranked

    RankDimensionResultWhy It's Here
    1Cost per task$0.13 on the Pareto frontierBest intelligence per dollar in the index
    2Open weights rank46, top open model globallyHighest score of any downloadable model
    3Agent tasksAutomationBench 53.1, Terminal Bench 2.1 89.9Beats closed rivals on tool use
    4CodingDeepSWE 71.9, ProgramBench 26.5Strong agentic coding, weak competition coding
    5Visual agent workMiMo Visual Coding 72.3Ahead of Claude Opus 5
    6CybersecurityCyberGym 94.0, ExploitBench 47.9Mixed, offensive security trails badly

    Cost: Why $0.13 per Task Changes the Calculus

    This is the result that matters most for anyone paying per token. At $0.13 per Intelligence Index task, MiMo-V2.6-Pro sits on the Intelligence-versus-Cost Pareto frontier. What that means in plain terms: no other model offers more intelligence per dollar, and no cheaper model offers more intelligence.

    Compare it to the closed frontier and the gap is stark. The top-scoring models on the same index cost many times more per task. If you're running an agent loop that fires dozens of calls per request, the difference between $0.13 and several dollars per task is the difference between a prototype and a product you can afford to ship.

    That's the whole story in one number. Cheap enough to be careless with.

    That's not a hypothetical benefit. Agent frameworks burn tokens on retries, tool calls, and verification passes. A model that performs at frontier level for a fraction of the cost changes what you can build, not just how much you pay.

    Open Weights: Xiaomi MiMo Open Weights Reached Number One

    The xiaomi mimo open weights release is the headline for anyone who self-hosts. MiMo-V2.6-Pro scored 46 on the Artificial Analysis Intelligence Index, the highest of any openly available model in the world, and the highest in the 114-model class Artificial Analysis groups it into.

    Before this release, the open weights crown was shared by Z.AI's GLM-5.3 and Moonshot's Kimi K3, both at 44. Xiaomi didn't just edge past them; it jumped two points and did it while its predecessor MiMo-V2.5-Pro sat at 26. A 20-point single-generation jump is rare, and it happened in a model family that was mid-tier last cycle.

    Xiaomi published the weights under an MIT license, along with the technical report, its reinforcement-learning environments, and training code. MIT matters because it's permissive. You can use the model commercially, modify it, and redistribute it without the restrictions some open licenses attach.

    If you want to test it locally, the weights are available to download, and the MIT license means you can run them on your own hardware without asking permission. Expect to need serious GPU memory for the 1.02T model, since the full parameter set has to fit somewhere even if only 42 billion activate at a time. Quantized builds shrink that requirement, and the community usually ships them within days of a release this size.

    Agent Tasks: Where MiMo-V2.6-Pro Punches at the Frontier

    Agent work is where this model genuinely competes with the closed labs. On AutomationBench v1.0.6, it scored 53.1, beating all three closed rivals it was tested against, which posted 50.3, 45.8, and 46.2. On Terminal Bench 2.1, it hit 89.9, the best result in its comparison table.

    That's the real story. It beats models that cost far more.

    It tied Claude Opus 5 on Agents' Last Exam at 31.6 and posted 1673 on GDPval-AA 2.1 against Opus 5's 1708, close enough to count as competitive. On OSWorld-Verified it scored 82.0, and on Toolathlon-Verified 76.9, both within a few points of the leaders without overtaking them.

    The one clear gap in this category is Terminal Bench 4.0, where it scored 34.9 against Opus 5's 49.0. The hardest terminal-based agentic work still belongs to the closed models, and Xiaomi's own report is upfront about that.

    Coding: Strong at Agentic Work, Weak at Competition Problems

    Coding splits into two very different skills, and MiMo-V2.6-Pro shows up differently on each.

    On DeepSWE v1.1, it scored 71.9, close behind Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0, and ahead of Fable 5 at 70.0. That's agentic coding, the kind where the model edits files, runs tests, and iterates. It's the skill most working developers actually need.

    On ProgramBench, the picture changes. MiMo-V2.6-Pro scored 26.5 against Opus 5's 37.0, a wide gap. ProgramBench leans toward competitive-style algorithm problems, and this is where the model trails. If your work is mostly building and debugging real codebases, the DeepSWE number is the one to trust. If you need a model for algorithm contests, look elsewhere.

    There's a practical reason the two numbers diverge. Agentic coding rewards long-horizon planning and tool use, the exact skills reinforcement learning can sharpen with enough environment scaffolding. Competitive programming rewards tight, single-shot reasoning under constraints, which is a harder thing to train into a model through interaction. Xiaomi optimized for the first and the benchmark data shows it.

    For most teams, that trade is the right one. Real development work looks more like the DeepSWE task than a contest problem, so the weaker ProgramBench score rarely shows up in daily use.

    Visual Agent Work and Computer Use

    On MiMo Visual Coding, the model scored 72.3, ahead of Claude Opus 5 at 70.0 and Fable 5 at 69.1, just behind GPT-5.6 Sol at 73.4. Visual agent tasks cover things like reading a screen and acting on it, which is the foundation for computer-use features.

    That result pairs with the OSWorld-Verified score of 82.0 to make a clear point: this model is good at looking at a screen and doing something useful with it. The company describes MiMo-V2.6 as an omnimodal model across text, image, video, and audio, with a 1M-token context window. Long context plus vision is the combination that makes computer-use agents practical rather than demo-grade.

    Cybersecurity: The Real Weak Spot

    The security results are mixed, and honesty requires saying so. MiMo-V2.6-Pro scored 94.0 on CyberGym and 80.2 on Xiaomi's own MiMo Cyber Bench. Those are strong numbers.

    Then the offensive side collapses.

    Then the offensive security benchmarks pull the other way. On ExploitBench it scored 47.9 against GPT-5.6 Sol's 78.5, and on SEC Bench Pro it posted 66.3 against Sol's 79.1. Defensive and general security knowledge look fine. Finding and exploiting vulnerabilities isn't a strength, and it trails badly there.

    If you're building a security tool, this isn't the model for the offensive half of the job. For general security question-answering and defensive analysis, the CyberGym number suggests it holds up.

    Architecture, License, and the Flash Variant

    MiMo-V2.6-Pro is a sparse mixture-of-experts model with 1.02 trillion total parameters and 42 billion active per token. The active parameter count is what keeps inference affordable. You pay for the 42 billion that fire on each token, not the full trillion sitting in memory.

    Xiaomi released two variants, Pro and Flash, both omnimodal and both advanced through scaled reinforcement learning. The company says the Pro model performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks, with stronger coding, computer use, 3D reasoning, and creative capabilities than the previous generation. Recent comparisons place the Pro model at general availability on September 22, 2026, one day after the weights went public.

    The Flash variant is the smaller option for teams that can't host a trillion-parameter model. On the same index, Flash lands well below Pro, which makes the choice straightforward: Pro for quality, Flash for footprint.

    Both models are omnimodal, meaning they handle text, image, video, and audio inputs without a separate vision or audio pipeline bolted on. That matters for agents that need to look at a screen or read a document and then act. A single model covering every input type cuts the number of moving parts in your stack, which is the quiet benefit of the omnimodal design.

    Who Should Actually Use MiMo-V2.6-Pro

    Three groups get the most out of this release. First, developers running agent pipelines where cost per task decides viability. Second, teams that need an open weights model for compliance or self-hosting reasons and can't accept a closed API. Third, anyone doing visual agent or computer-use work who needs long context and vision in one model.

    Who should skip it: competitive programming users, offensive security work, and anyone whose workload depends on the top of the closed frontier on the hardest reasoning tasks. Xiaomi built a model that trades wins with rivals costing many times more, but it concedes ground on the toughest benchmarks, and pretending otherwise wouldn't help you pick the right tool.

    One caution on reading any single score. Artificial Analysis revises its index regularly, and models move position as new evaluations get folded in. A ranking today isn't a permanent verdict, and the leaderboard that put MiMo-V2.6-Pro on top has shifted several times this year as Chinese labs traded the open weights title back and forth.

    There's one more angle Xiaomi raised: the company frames MiMo-V2.6 as a step toward recursive self-improvement research, where models help improve the next generation. Whether that pans out is genuinely unknown. What's concrete is the score, the price, and the license, and all three hold up on their own.

    The mimo v2.6 benchmark picture is consistent across sources: this is a frontier-class agentic model at open weights pricing. Check results through Artificial Analysis and Xiaomi's technical report, then run your own workload against it, because a benchmark score is a starting point rather than a verdict. If it fits your stack, the cost advantage alone justifies the test, and the MIT license means the only thing you risk is the time it takes to try.

    If you're looking for a daily assistant with a mature ecosystem, Gemini remains the handy choice to pair with this Android Drop.
    Google Gemini APK

    Google Gemini is a productivity and creativity app that supports live conversations and multimodal creation.

    ProductivityAI Chatbot

    ProductivityAI Chatbot


    Back to top

    Featured lists

    NEWSNEWSREVIEWSREVIEWSHOWTOHOWTO
    Latest Reviews
    What Is Claude Sonnet 5.5? Benchmarks, Pricing, and What's New in the 2026 Model
    What Is the September 2026 Android Security Update? 180 Vulnerabilities Explained
    Transformers: Eternal War Review 2026: Is This Mobile RPG Worth Playing?
    What Is the Anthropic Claude ART Enzyme System? Discovery Explained
    Top Reviews
    Back Alley Tales: A Unique Blend of Surveillance and Storytelling
    Best New Features in the Minecraft 26.50 Update, Ranked
    Kingdom Rush 6: Genesis TD Review: Worth Buying vs Earlier Games?
    Roblox VNG vs Global Version: Which Roblox Should You Play?
    Does Cat Resolution Pro Work? Review of the Stretched Screen App for Mobile Phones
    Summer Life in the Countryside: A Charming Visual Novel Experience
    EA SPORTS FC Soccer Mobile 27 Update: Every New Season Feature Ranked
    GTA 5 Mobile: The Ultimate Open-World Experience on Your Fingertips
    Subscribe to APKPure
    Be the first to get access to the early release, news, and guides of the best Android games and apps.
    No thanks
    Sign Up
    Subscribed Successfully!
    You're now subscribed to APKPure.