Our website uses necessary cookies to enable basic functions and optional cookies to help us to enhance your user experience. Learn more about our cookie policy by clicking "Learn More".
Accept All Only Necessary Cookies
NEWSREVIEWSHOWTO

What Is Kimi? Moonshot AI's K3 Model and Chatbot App Explained

Candida Corkery

What is Kimi? Learn how Moonshot AI's 2.8T parameter K3 model works, what the Kimi chatbot app can do, agent cluster capabilities, pricing, and why it matters for AI users.

Catelog

    Kimi is an AI assistant built by Moonshot AI, a Beijing-based startup founded in March 2023 by Yang Zhilin, a Carnegie Mellon PhD who previously worked at Google Brain. The product launched as a simple long-text chatbot in October 2023 and has since grown into one of China's most widely used AI tools, with a web platform, iOS and Android apps, developer APIs, and a coding assistant. Its latest model, Kimi K3, was released on July 16, 2026, with full open weights published on July 27. K3 packs 2.8 trillion parameters into a Mixture-of-Experts architecture, supports a 1-million-token context window, and ranks first on the Frontend Code Arena benchmark.

    If you've been hearing the name "Kimi" in AI coverage and want to understand what it's about, how it works, and whether the Android app is worth your time, this guide covers all of it. Think of this as your primer on Kimi, the China AI assistant that's now competing with the biggest names in the field.

    2. What Is Kimi and Who Made It?

    Kimi is the consumer-facing AI assistant product from Moonshot AI, a company whose Chinese name translates to "dark side of the moon," a reference to the Pink Floyd album. The company was founded in March 2023 by Yang Zhilin and a small team of Tsinghua University and Carnegie Mellon alumni. Yang holds a bachelor's degree from Tsinghua and a PhD in computer science from Carnegie Mellon, where he worked under Ruslan Salakhutdinov (now Apple's AI chief). He co-authored two foundational NLP papers, Transformer-XL and XLNet, and is among the most cited researchers in China's NLP field under 35.

    The company's Chinese name comes from Yang's favorite Pink Floyd album. He and co-founder Zhou Xinyu played in a campus rock band at Tsinghua. That unconventional streak carries into Moonshot's technical choices. While competitors chased proprietary model development behind closed APIs, Moonshot bet on open-weight releases from early on.

    Kimi started as a free chatbot that could process 200,000 Chinese characters in a single prompt. By early 2024, that limit jumped to 2 million characters. Big jump. The product expanded from a text-only web chat to a multimodal AI assistant with coding, research, and agent capabilities available across mobile, web, and desktop platforms. Moonshot is backed by Alibaba and Tencent. Its latest valuation sits around $3.3 billion.

    3. Kimi K3: The 2.8-Trillion Parameter Model

    Kimi K3 is Moonshot's most capable model to date. It was announced on July 16, 2026, with full model weights open-sourced on July 27. Here's what the specs look like:

    DeveloperMoonshot AI
    Total parameters2.8 trillion (2.8T)
    Active parameters per token~104 billion
    ArchitectureMoE + Kimi Delta Attention + Attention Residuals
    Experts896 (16 activated per token)
    Context window1,048,576 tokens (1 million)
    Input modalitiesText, images, video frames
    Output modalityText
    Open weightsYes (CC-BY-NC-SA 4.0)
    API model IDkimi-k3

    The model uses a Sparse Mixture-of-Experts design. Out of 896 expert modules, only 16 are activated for any given token. That keeps inference costs manageable. It also means the model doesn't need to run all 2.8 trillion parameters for every single word it generates.

    Two architectural innovations distinguish K3 from its predecessors. The first is Kimi Delta Attention (KDA), a hybrid linear attention mechanism that replaces some quadratic attention calculations. This makes processing 1 million tokens economically feasible. Not theoretically possible. Actually affordable. The second is Attention Residuals, which improve information flow across deep layers of the network and stabilize training at this scale. Together, these give K3 roughly 2.5 times the scaling efficiency of K2. Same compute budget, meaningfully more capable model.

    K3 also includes a vision encoder called MoonViT-V2, trained from scratch using next-token prediction rather than contrastive pre-training. This means the model's visual understanding is baked into the core architecture, not bolted on as a separate module. You can hand it a screenshot, a chart, or a UI mockup. It reads the content directly.

    On benchmark performance, K3 ranks first on Arena.ai's Frontend Code Arena leaderboard, ahead of Claude Fable 5, GPT-5.6, and GLM-5.2. On the Artificial Analysis Intelligence Index, it ranks fourth globally, behind only Claude Fable 5, GPT-5.6 Sol, and Claude Opus 4.8. Does one benchmark prove overall superiority? No. But it shows an open-weight Chinese model is now competing at the top tier alongside proprietary systems from OpenAI and Anthropic.

    4. What Can the Kimi Chatbot Do?

    The Kimi chatbot has grown well beyond simple Q&A. It now covers four main use areas:

    • Long-document analysis. This has been Kimi's signature strength since launch. With a 1-million-token Kimi long context window, you can upload entire codebases, research papers, legal contracts, or financial reports and ask questions across all of them in a single conversation. A 1-million-token window roughly equals 750,000 English words. For context, that's about 10 full-length novels or a complete enterprise documentation portal.
    • Coding and development. Kimi Code is a separate developer toolset that includes a CLI agent and a VS Code plugin. The underlying K3 model handles code generation, debugging, refactoring, and multi-step development tasks. Cognition, the company behind the AI coding agent Devin, integrated K3 into both its desktop client and CLI tools. The company stated that K3 was the first open-weight model they tested that approached frontier-level performance on coding benchmarks.
    • Deep research. Kimi's research mode uses what Moonshot calls "thinking mode" to break down complex questions into sub-tasks, search the web, synthesize findings, and produce structured reports. The Kimi-Researcher agent, first introduced with the K2 Thinking model, scored 44.9% on Humanity's Last Exam (HLE), surpassing GPT-5's 41.7% at the time. On BrowseComp, which tests autonomous web browsing and information gathering, K2 Thinking scored 60.2% compared to the human average of 29.2%.
    • Visual understanding. K3 can read images, screenshots, diagrams, and video frames. Upload a UI screenshot and it can reproduce the interface as working frontend code. Hand it a chart and it can extract the data and explain the trends. The multimodal capability is native, meaning the model processes images and text through the same architecture rather than routing images through a separate vision module.

    5. Kimi Agent Cluster: 300 Sub-Agents Working in Parallel

    One of K3's most distinctive features is its Agent Swarm capability. When you give Kimi a complex task, it doesn't have to work through it sequentially. It can spin up to 300 sub-agents that work in parallel, with support for over 4,000 concurrent tool calls.

    Here's what that means in practice. Say you ask Kimi to research the competitive dynamics of a specific industry across 20 companies. Instead of processing one company at a time, Kimi can dispatch sub-agents to research multiple companies simultaneously, run web searches in parallel, cross-reference findings, and compile the results into a single report. Moonshot claims this can reduce task completion time by up to 10x for large-scale research, multi-language translation, or cross-domain analysis work.

    The Agent Swarm was first introduced with K2.5 and scaled up significantly with K3. The system uses Parallel Agent Reinforcement Learning (PARL) for training, which teaches the model how to decompose tasks, assign sub-tasks to sub-agents, and synthesize their outputs. This is different from simpler agent frameworks that chain tools linearly. The Kimi app on kimi.com and in the mobile app offers this as a feature called "K3 Cluster" mode.

    For developers, Moonshot also offers Kimi Claw, a zero-deployment cloud automation platform with a library of over 5,000 skills that can be chained together for automated workflows. And Kimi Work, a local agent for Mac and Windows desktop, uses Kimi Code as its core engine for running scheduled tasks and installed skills.

    6. Kimi Pricing and How to Access It

    Kimi is available through several channels, each with different pricing:

    Free web and mobile app. The Kimi chatbot at kimi.com and the Android/iOS app offer free unlimited chat with the K2.6 model. Free users get a limited quota of K3 and Agent tasks per day. No payment required.

    Consumer subscriptions. Moonshot introduced a tiered subscription model with musical tempo names:

    TierPriceWhat you get
    Adagio (free)¥0Unlimited K2.6 chat, limited Agent quota
    Andante¥49/monthMore K3 quota, deep research, API credits
    Moderato¥99/monthHigher K3 limits, parallel deep research
    Vivace (international)$199/monthPeak priority access, full feature set

    Note: As of late July 2026, Moonshot temporarily paused new consumer subscriptions due to compute capacity limits. Existing subscribers are unaffected. New subscriptions are waitlisted.

    Developer API. The Kimi open platform (platform.moonshot.cn) offers pay-as-you-go API access:

    ModelInputOutputCache hit
    K3¥20/MTok¥100/MTok¥2/MTok
    K2.7 Code¥6.50/MTok¥27/MTok¥1.30/MTok
    K2.6¥6.50/MTok¥27/MTok¥1.10/MTok

    The API is OpenAI-compatible, so existing tools built for GPT APIs can switch to Kimi with minimal changes.

    7. Kimi vs Other AI Assistants: What Makes It Different?

    Kimi occupies a specific niche among AI assistants. Here's how it compares:

    • Vs. ChatGPT. Kimi's advantage is its 1-million-token context window (ChatGPT's longest is 1 million tokens with GPT-5.6, but at much higher API cost). Kimi's K3 API is priced at roughly ¥20 per million input tokens, while GPT-5.6 commands significantly more. The trade-off: ChatGPT has a larger library of plugins, integrations, and third-party tools built around it.
    • Vs. Claude. Anthropic's Claude Fable 5 edges out K3 on the Artificial Analysis Intelligence Index, but K3 beats it on the Frontend Code Arena. Kimi's open weights mean you can download and run the model locally; Claude is API-only. For users who want to customize a model for specific domains or keep data on their own infrastructure, that's a meaningful difference.
    • Vs. DeepSeek. DeepSeek's API is roughly 17 times cheaper than K3 for output tokens. DeepSeek is the budget option. K3 targets the performance frontier, competing with closed-source models on capability rather than competing on price. DeepSeek V4 tops out at 1.6 trillion parameters; K3 is nearly double that.

    What makes Kimi distinct isn't any single feature. It's the combination of open weights, frontier-level performance, a 1-million-token context window, native multimodal input, and the Agent Swarm architecture. No other Kimi open source model competitor currently offers all five.

    8. How the Wider Industry Is Using Kimi

    Several major companies have adopted Kimi models in their production stacks:

    • Cloudflare adopted Kimi K2.5 for its automated code review system in March 2026. The company published a blog post detailing the economics: its AI agents process over 7 billion tokens per day. Switching from a leading closed-source model to Kimi cut costs by 77%, from an estimated $2.4 million per year to roughly $550,000. Cloudflare chose Kimi based on the price-performance ratio, not sentiment.
    • Cognition (maker of Devin, an AI coding agent) integrated K3 into its desktop client and CLI tools. The company stated that K3 was the first open-weight model they tested that approached frontier-level coding performance on their internal benchmarks.
    • Cursor, the AI-powered code editor, was found to be using Kimi K2.5 under the hood for its Composer 2 feature. When developers discovered the model ID, it sparked discussion about how deeply open-weight models had penetrated production developer tools.
    • Perplexity CEO Aravind Srinivas publicly praised Kimi K2 on launch. He noted strong internal evaluation results. Perplexity began post-training work with Kimi models. It became one of the few open-weight Chinese models integrated by a major US AI search company.

    On the infrastructure side, Huawei's Ascend 950 series, Alibaba Cloud's Zhenwu M890 super-node instances, Nebius, Baseten, and Fireworks all announced Day-0 support for K3 deployment. Hugging Face CEO Clem Delangue reported that K3's open-weight release garnered over 4,000 likes in 30 minutes, calling it the fastest growth he'd seen for a model release.

    9. Kimi's Founder and NVIDIA GTC 2026

    Yang Zhilin, Moonshot's founder and CEO, was the only Chinese independent AI company representative invited to speak at NVIDIA's GTC 2026 conference in San Jose in March 2026. NVIDIA CEO Jensen Huang mentioned Kimi K2.5 multiple times in his keynote address.

    Yang's GTC talk focused on three technical concepts: Token Efficiency, Long Context, and Agent Swarms. He argued that as high-quality training data becomes a finite resource, improvements in token efficiency (getting more learning per token of training data) become the key lever for pushing model intelligence higher. He introduced the Muon optimizer, a second-order optimizer that doubles token learning efficiency compared to standard AdamW.

    His broader argument: if open-source and proprietary models reach similar capability levels, open systems will ultimately dominate by unlocking wider developer ecosystems, more applications, and greater overall token output. Closed models hold market share today, but open platforms may generate more economic value by enabling more participants to build on top.

    10. Is the Kimi App Worth Downloading?

    If you want to try Kimi on your phone, the Android app is available on APKPure. The app gives you the full Kimi experience: chat with K2.6 or K3, use Agent mode for complex tasks, upload images for visual analysis, and run deep research queries.

    For users who mainly need a free AI chatbot for everyday questions, Kimi's free tier covers that. The subscription tiers add value if you need heavier K3 access, parallel agent tasks, or API credits for development work.

    The app is lightweight, supports both Chinese and English input, and syncs conversation history across devices. It also includes access to Kimi's Agent features. You can build websites, create presentations, process documents and spreadsheets, and run batch research tasks from your phone.

    If you're a developer, the Kimi open platform and Kimi Code CLI give you API access and a local coding agent. The API is OpenAI-compatible. Switching existing projects over requires minimal effort.

    11. Summary

    Kimi started as a long-text chatbot from a Beijing startup named after a Pink Floyd album. Three years later, it's the product behind the world's largest open-weight AI model, a chatbot with agent cluster capabilities that can dispatch 300 sub-agents in parallel, and a tool that companies like Cloudflare and Cursor have integrated into production systems. The K3 model's 2.8-trillion-parameter architecture, 1-million-token context window, and first-place ranking on the Frontend Code Arena benchmark put it in direct competition with the best models from OpenAI and Anthropic, with the added advantage of open weights you can download and run yourself.

    The Kimi Android app brings all of this to your phone, and you can grab the latest version from APKPure to try it out.

    Kimi APK

    Kimi is a productivity and research app that coordinates tools to create documents, presentations, spreadsheets, and websites.

    Tools

    Tools



    Back to top

    Featured lists

    NEWSNEWSREVIEWSREVIEWSHOWTOHOWTO
    Latest Reviews
    Honkai Star Rail 4.6 Phase 2 Ranked: Blade Rerun and Every Event
    Applied Thaumaturge Explained: The Thaumcraft Addon That Talks to AE2
    What Is Claude Sonnet 5.5? Benchmarks, Pricing, and What's New in the 2026 Model
    What Is the September 2026 Android Security Update? 180 Vulnerabilities Explained
    Top Reviews
    Back Alley Tales: A Unique Blend of Surveillance and Storytelling
    Kingdom Rush 6: Genesis TD Review: Worth Buying vs Earlier Games?
    Best New Features in the Minecraft 26.50 Update, Ranked
    One State RP: The Ultimate Open World Role-Playing Experience
    EA SPORTS FC Soccer Mobile 27 Update: Every New Season Feature Ranked
    Does Cat Resolution Pro Work? Review of the Stretched Screen App for Mobile Phones
    Roblox VNG vs Global Version: Which Roblox Should You Play?
    Sober: 7 Features That Make Roblox Work Natively on Linux
    Subscribe to APKPure
    Be the first to get access to the early release, news, and guides of the best Android games and apps.
    No thanks
    Sign Up
    Subscribed Successfully!
    You're now subscribed to APKPure.