Moonshot AI released Kimi K3 on July 16, 2026, a 2.8 trillion-parameter mixture-of-experts model with a one-million-token context window and native vision capabilities. It is the largest open-weight model ever built. Weights are scheduled for public release by July 27.
On the Artificial Analysis Intelligence Index, K3 scores 57, ranking fourth overall. Claude Fable 5 leads at 60, followed by GPT-5.6 Sol at 59 and Sol (xhigh) at 58. The gap between K3 and the top proprietary models is smaller than any previous open-weight model has achieved. Moonshot itself acknowledges K3 “still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol,” according to its technical blog.
The model’s significance is less about raw intelligence scores and more about where it leads and what it costs.
Architecture and Scaling Efficiency
K3 is built on two architectural innovations Moonshot calls Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). KDA handles long-context efficiency, while AttnRes lets each layer retrieve selected outputs from earlier layers. Moonshot credits these changes with roughly 25% better training efficiency for under 2% additional computational cost, according to Towards AI’s analysis.
The model uses an expanded mixture-of-experts architecture, activating 16 out of 896 experts per token via a Stable LatentMoE framework. Combined with refined training recipes, Moonshot claims an approximate 2.5x improvement in overall scaling efficiency compared to its predecessor, Kimi K2. The raw weight files occupy roughly 1.5 terabytes in MXFP4 format, and Moonshot recommends at least 64 accelerators for deployment.
For builders evaluating self-hosting, the scale represents a real barrier. The key-value cache adds substantial memory overhead beyond the base weights. This is not a model that runs on a single node.
Coding and Agent Capabilities
K3’s strongest benchmark results cluster around coding and agentic tasks. On the Artificial Analysis Coding Index, it scores 76.24, within 0.25 points of Claude Fable 5. On AutomationBench-AA, which measures agentic SaaS workflows, K3 leads at 52.71%.
Moonshot’s own blog documents several capabilities that go beyond benchmark scores. During development, an early version of K3 handled the majority of the team’s GPU kernel optimization work. K3 built MiniTriton, a compact compiler with its own tile-level intermediate representation layer, optimization passes, and a PTX code-generation pipeline. Across supported benchmarks, MiniTriton delivers performance on par with or better than Triton and torch.compile.
In a 48-hour autonomous run, K3 designed a chip to serve a nano model built on its own architecture. Using open-source EDA tools, it closed timing at 100 MHz, packed 1.46 million standard cells and 0.277 MB of SRAM within 4 mm², and sustained over 8,700 tokens per second decode throughput in simulation. These are demonstration results, not production deployments, but they illustrate the length of autonomous engineering sessions K3 can sustain.
The weakest benchmark lane is knowledge reliability. K3’s AA-Omniscience score of 18 trails Fable’s 40 and Opus 4.8’s 27. For agent builders, this means K3 handles code execution and tool orchestration well but is less reliable when agents need to recall or verify factual claims.
The Economics
K3’s API pricing is $3.00 per million input tokens and $15.00 per million output tokens, with a cache hit price of $0.30 per million tokens, according to Kimi’s API platform. Those per-token prices are 40% to 50% below OpenAI and Anthropic’s comparable tiers.
The per-task economics tell a more complicated story. Artificial Analysis measures K3 at $0.95 per weighted Intelligence Index task, only 8% below Sol’s $1.04. The reason: K3 generates 54% more output tokens per task than Sol. Cheaper per token, but chattier per response.
Speed is another constraint. K3’s first-party endpoint generates 35.2 output tokens per second, according to Artificial Analysis, and completes a standardized 500-token request in about 75.6 seconds end-to-end. One advantage: K3 streams its first reasoning token in approximately four seconds, while Fable and Sol’s maximum-effort endpoints stay silent for around two minutes. For interactive agent workflows where perceived latency matters, the faster first token is meaningful even if total completion time is longer.
The pricing also represents a step up from Moonshot’s own prior models. K3’s per-token costs are roughly three times higher than Kimi K2.6, which offers input at $0.95 and output at $4.00 per million tokens. The jump reflects the model’s dramatically larger parameter count and computational demands.
Demand Overwhelmed Supply
Within 48 hours of launch, demand pushed Moonshot’s GPU infrastructure near capacity. The company paused new subscriptions on July 19 to protect service quality for existing customers. Towards AI reports this confirms that “inference supply already binds adoption.”
The subscription pause is a reminder that open weights do not eliminate the compute bottleneck. They shift it from API access to inference infrastructure. When K3’s weights release on July 27, teams that want to self-host will need to provision hardware at a scale most organizations do not maintain.
Where K3 Sits in the Competitive Landscape
K3 marks the third trillion-scale model Moonshot has released in twelve months, accelerating a cadence where, according to Towards AI, Kimi models have “set the upper bound of open-model sizes” for nine of the past twelve months.
The Artificial Analysis leaderboard snapshot as of July 23 shows how compressed the frontier has become:
| Model | Intelligence Index | Cost per Task |
|---|---|---|
| Claude Fable 5 | 60 | $2.75 |
| GPT-5.6 Sol (max) | 59 | $1.04 |
| GPT-5.6 Sol (xhigh) | 58 | $0.68 |
| Kimi K3 | 57 | $0.95 |
| GPT-5.6 Sol (high) | 56 | $0.45 |
| Claude Opus 4.8 (max) | 56 | $1.80 |
K3 occupies a narrow band: three points below Fable on intelligence, roughly one-third the cost per task. For agent workloads where K3’s coding and automation scores match or exceed the leaders, the price gap is the argument. For workloads requiring factual reliability, the knowledge deficit is the counter-argument.
The stock market registered the competitive pressure. AI stocks sold off on July 18 in what Towards AI described as “a smaller DeepSeek moment, with K3 one driver among several.” The sell-off reflects investor anxiety about how quickly open-weight models are closing the gap that justifies premium API pricing.
The Distillation Question
Towards AI’s analysis directly addresses the industry’s open debate about whether Chinese labs are distilling from US frontier models. The assessment: “of course they are, and so is everyone else.” Frontier intelligence enters training pipelines through multiple routes, from direct outputs to synthetic examples to purchased datasets that used Claude or GPT to generate tasks and evaluations. Elon Musk has acknowledged xAI’s use of Claude-generated data.
The analysis argues this does not diminish K3’s achievement. Using distilled data effectively requires a strong base model, sophisticated reinforcement-learning infrastructure, and extensive original research. Moonshot’s novel architecture work (KDA, AttnRes, Stable LatentMoE) and the demonstrated ability for K3 to autonomously build compilers and optimize kernels reflect engineering that goes well beyond data collection.
The Open-Weight Pricing Pressure
K3’s significance for agent builders is structural, not just technical. If a near-frontier open-weight model costs meaningfully less per task than closed-API alternatives, the economic logic for agent fleets shifts. Longer inference chains, more aggressive system prompts, and persistent multi-agent deployments all become cheaper when the underlying model’s intelligence-per-dollar ratio improves.
Moonshot has not yet proven K3’s economics are sustainable at scale, and the subscription pause suggests they are not, at current infrastructure levels. But the benchmark results demonstrate that the intelligence gap between open and closed models has narrowed to single digits on standardized evaluations. The question for US frontier labs is whether a three-point lead on the Intelligence Index justifies a three-to-one premium on cost per task, especially in the agentic workloads where K3 competes most effectively.
The weights release on July 27 will test that question directly. Every team that provisions 64 or more accelerators will be running their own cost analysis.