Our website uses necessary cookies to enable basic functions and optional cookies to help us to enhance your user experience. Learn more about our cookie policy by clicking "Learn More".
Accept All Only Necessary Cookies
NEWSREVIEWSHOWTO

What Is dots3-note Preview? 280B Long-Range Agent Explained

Dorothy Morgan

dots3-note Preview is a 280B MoE model from Xiaohongshu AI with 512K context, multimodal reasoning, and TEMPO reinforcement learning for long-range agent tasks.

Catelog

    Unlike most open-source releases that prioritize benchmark scores, the dots3-note Preview model focuses on a harder challenge: staying coherent over tasks that take hours, days, or even weeks to finish. If you're tracking the shift from chatbots to autonomous agents, this is one of the first open-source models built specifically for that gap.

    1. What dots3-note Preview Brings to AI Development

    dots3-note Preview is the first open-weight release from dots studio, the Xiaohongshu AI research team. It's a Mixture-of-Experts model with 280 billion total parameters that activates only 16 billion per token, supports a 512K context window, and processes text, images, video, and audio in a single forward pass. Released on August 14, 2026 under the Apache 2.0 license, the model targets a problem that most large language models still struggle with: completing tasks that unfold over hours or days rather than seconds. Most open-source models compete on benchmark scores for math and coding.

    The dots3 model takes a different angle: long-horizon tasks where the user's intent isn't fully clear at the start, and the environment keeps changing.

    Think of planning a two-week trip, coordinating a wedding, or running a small business for a quarter. These tasks don't have a single correct answer, and they stretch across days or weeks with shifting constraints. dots3-note Preview is the lightest member of the dots3 family, with heavier variants named jazz and aria planned for later release. The same series produced dots-note-3.0, which scored a full 42 out of 42 at the 2026 International Mathematical Olympiad, the first AI to achieve that score under official supervision.

    2. What's in 280B MoE Model Architecture

    The architecture behind dots3-note Preview uses a Mixture-of-Experts design that keeps inference costs low despite the large total parameter count. Here's how the numbers break down from the official GitHub repository:

    • 280 billion total parameters
    • 16 billion activated per token
    • 256 routed experts plus 1 shared expert with top-8 routing
    • 46 layers (1 dense plus 45 MoE)

    The attention mechanism mixes 13 Dedicated Self-Attention layers with 33 Sliding Window Attention layers in roughly a 1:3 ratio. Sliding Window Attention handles local token relationships at lower memory cost, while Dedicated Self-Attention with Top-2048 token selection maintains access to important long-range dependencies. That hybrid design is what makes the 512K context window feasible without running out of memory. A Multi-Token Prediction head with 1.13 billion parameters enables speculative decoding, which can cut token generation latency by more than half. The 280B MoE model supports both BF16 and FP8 precision, and the FP8 checkpoint fits on a single 8-GPU node.

    3. Multimodal Reasoning Across Text, Images, and Audio

    dots3-note Preview handles four input types in one model: text, images, video, and audio.

    The vision encoder is a Mixture-of-Experts ViT with 7 billion total parameters (1.2 billion activated), and the audio encoder is a dense 800-million parameter network. This means the model can read a document, look at a floor plan, watch a video clip, and listen to a voice memo without switching to a separate model for each input. For long-range agent tasks, this matters because real-world information doesn't arrive in clean text.

    A user planning a trip might share screenshots of flight options, a voice note about dietary restrictions, and a video walkthrough of a hotel room. The model processes all of these in a single conversation thread within its 512K context window. This multimodal reasoning capability is what separates dots3-note Preview from text-only models that need external adapters to handle images or audio.

    4. TEMPO Reinforcement Learning for Long-Range Agent Tasks

    The training method that sets dots3-note Preview apart is called TEMPO, which stands for Test-time-scaled Value Estimation with Macro-step Policy Optimization.

    Traditional reinforcement learning for language models gives a reward at the end of a full rollout, which works fine for short tasks like solving a math problem. But when a task runs for hundreds of steps across hours, the credit assignment problem becomes severe: which specific step caused the failure? TEMPO reinforcement learning addresses this by splitting long trajectories into stages, then running self-evaluation at each stage boundary.

    The model temporarily stops executing and uses extra compute to assess its current state and estimate future success probability before adjusting its strategy. In experiments on ARC-AGI-3, TEMPO improved average scores by 31.5% over the baseline checkpoint and 20.6% over GRPO. The approach trains both the action model and the self-evaluation model together, so the agent learns to judge its own progress during deployment rather than waiting for a final score.

    How to Track AI Costs with OpenRouter Activity Dashboard and Analytics API

    How To

    How to Track AI Costs with OpenRouter Activity Dashboard and Analytics API

    Learn how to use the OpenRouter Activity dashboard and beta Analytics API to track AI cost per agent, monitor model spending, and drill into individual requests.

    Learn More

    5. Real-World Use Cases and Open Evaluation Benchmarks

    dots studio released two evaluation benchmarks alongside the model, both focused on real-world long-horizon tasks.

    • VibeSearchBench

    VibeSearchBench tests multi-round search where the user's intent is disclosed gradually across 20 domains and 200 tasks, scored with Triplet F1. The current best score belongs to Claude Opus 5 at 31.14, and every tested model sits far below practical usability.

    • VibeLifeBench

    VibeLifeBench is harder. It simulates real timelines with 10 domains, 20 tasks, each spanning 20 to 30 stages, checked against 1,247 atomic criteria for state consistency, tool execution, and final delivery. Seven top-tier models were tested, and none reached the passing threshold.

    In practical demonstrations, dots3-note Preview has planned multi-day trips, managed wedding logistics, played through Slay the Spire II with real-time strategy adjustments, cracked ARC-AGI-3 puzzles, and completed end-to-end visionOS app development from requirements to running code. These aren't lab benchmarks. They show the model acting as a long-range agent that can observe, decide, and correct course over extended timeframes. The multimodal capabilities also map to scenarios that current phone-based assistants can't handle: reading a screenshot of a restaurant menu, cross-referencing it with a voice memo about food allergies, and suggesting what to order requires exactly the kind of processing this model is built for.

    6. Strengths and Limitations of the dots3 Model

    The strengths are clear from the architecture and evaluation data. The MoE design keeps per-token compute at 16 billion activated parameters, making it cheaper to run than dense models of similar total size. The 512K context window holds long conversations and large documents without chunking. Multimodal input means the model works with real-world data formats out of the box. The Apache 2.0 license allows commercial use without restrictions. And the TEMPO training method gives the model a self-correction mechanism that most agent models lack.

    The limitations are equally real. No model passed VibeLifeBench's passing threshold, including dots3-note Preview itself, which tells you that long-range agent tasks remain an unsolved problem. The model outputs text only, so it can't generate images, audio, or video. Running the FP8 checkpoint requires 8 GPUs, which puts it out of reach for individual developers without cloud compute. The Transformers and SGLang integration PRs are still under review, so vLLM is the only framework with native support right now. And while the model can play games and write code, these capabilities haven't been independently verified by third-party evaluators. Compared to other open-source models like DeepSeek-V3 or Qwen3, dots3-note Preview trades raw benchmark performance for agent-specific training, which is a bet on where AI development is heading rather than where it is today.

    7. How to Access and Deploy dots3-note Preview

    The model weights are available on HuggingFace (dots-studio/dots3-note-prev and dots-studio/dots3-note-prev-fp8) and ModelScope. The GitHub repository at studio-dots-ai/dots3-note-prev contains the full documentation, quickstart code, and deployment recipes. For serving, vLLM has native support on its main branch, with an FP8 deployment recipe targeting eight NVIDIA H100 GPUs with tensor parallelism of 8 and expert parallelism of 8. SGLang support is available through a development Docker image (lmsysorg/sglang:dev-dots3-note). The model also exposes an API endpoint at dots.ai/platform for developers who want to try it without hosting infrastructure. The MTP speculative decoding feature can generate multiple candidate tokens per forward pass and verify them in parallel, which reduces decoding latency and makes the model practical for interactive applications where response time matters.

    If you're evaluating AI models for agent-based workflows, dots3-note Preview is worth testing against your own long-horizon tasks. The open weights and benchmarks let you verify the claims yourself rather than relying on marketing materials. Download the checkpoint from HuggingFace, run the VibeLifeBench evaluation on your own infrastructure, and compare results against your specific use case. The dots3 model family is still in its early stages, and the preview release is a research artifact more than a production tool. But for developers building autonomous agents, it's one of the few open-source options that takes the long-range problem seriously.

    Back to top

    Featured lists

    NEWSNEWSREVIEWSREVIEWSHOWTOHOWTO
    Latest Reviews
    What Is Claude Sonnet 5.5? Benchmarks, Pricing, and What's New in the 2026 Model
    What Is the September 2026 Android Security Update? 180 Vulnerabilities Explained
    Transformers: Eternal War Review 2026: Is This Mobile RPG Worth Playing?
    What Is the Anthropic Claude ART Enzyme System? Discovery Explained
    Top Reviews
    Back Alley Tales: A Unique Blend of Surveillance and Storytelling
    Best New Features in the Minecraft 26.50 Update, Ranked
    Kingdom Rush 6: Genesis TD Review: Worth Buying vs Earlier Games?
    EA SPORTS FC Soccer Mobile 27 Update: Every New Season Feature Ranked
    Does Cat Resolution Pro Work? Review of the Stretched Screen App for Mobile Phones
    Roblox VNG vs Global Version: Which Roblox Should You Play?
    GTA 5 Mobile: The Ultimate Open-World Experience on Your Fingertips
    One State RP: The Ultimate Open World Role-Playing Experience
    Subscribe to APKPure
    Be the first to get access to the early release, news, and guides of the best Android games and apps.
    No thanks
    Sign Up
    Subscribed Successfully!
    You're now subscribed to APKPure.