Kimi K3 Review 2026: Is Moonshot AI's 2.8T Open-Weight Model Worth It?

Kimi K3 is Moonshot AI's 2.8-trillion-parameter open-weight model with a 1M-token context window that reportedly leads Arena.ai's Frontend Code Arena. Here's the pricing, the benchmarks, and who it's actually for.

⚠️ Affiliate Disclosure: This article contains affiliate links. We may earn a commission if you purchase through our links, at no extra cost to you. Read our full disclosure.
Kimi K3 review featured graphic showing a dark mode AI dashboard interface with code and chat panels in purple branding

Quick Verdict

Kimi K3, released by Moonshot AI on July 16, 2026, is the largest open-weight model to ship to date — a 2.8-trillion-parameter Mixture-of-Experts system that landed on the Artificial Analysis Intelligence Index above Claude Opus 4.8 and reportedly leads Arena.ai's Frontend Code Arena leaderboard. It ships under a Modified MIT license, supports a 1,000,000-token context window with native vision, and is available now through Moonshot's website and API — though the full downloadable weights aren't out yet; Moonshot has promised those "by July 27, 2026."

This isn't a hands-on usage review — Kimi K3 became publicly available less than 48 hours before this article was written. What follows is based on Moonshot AI's launch materials, OpenRouter's listed pricing, independent benchmark tracking from Artificial Analysis and Simon Willison's testing notes, and coverage from outlets like Tom's Hardware, CNBC, and Fortune. We'll update this page with first-hand testing notes once we've put the model through real workloads.

Our early take: 8.2/10 — a genuinely frontier-level open-weight release with an unusual amount of mainstream financial-press attention behind it, offset by the fact that the open weights themselves aren't downloadable yet and pricing has jumped sharply from Moonshot's previous model.

What Is Kimi K3?

Kimi K3 is the newest flagship model from Moonshot AI, the Beijing-based lab behind the Kimi assistant. It's a sparse Mixture-of-Experts (MoE) model with roughly 2.8 trillion total parameters spread across 896 experts, of which only 16 are activated per token — about 1.8% of the total pool, which is how Moonshot keeps inference costs manageable despite the model's overall size.

Architecturally, K3 is built on what Moonshot calls Kimi Delta Attention (KDA), a hybrid linear-attention design paired with "Attention Residuals," aimed at making the 1-million-token context window computationally practical rather than just a marketing number. The model ships with native vision understanding and shipped in two variants at launch: K3 Max, positioned for chat and general agent tasks, and K3 Swarm Max, aimed at large-scale parallel workloads.

The release got an unusual amount of coverage outside the usual AI press — Fortune, CNBC, and Yahoo Finance all framed it as a "DeepSeek shock" moment, referencing the market reaction to China's DeepSeek-R1 release in early 2025. Whether that framing holds up is a separate question from whether the model is good, which is what the rest of this review focuses on.

Kimi K3 Pricing: API Costs

Kimi K3 is priced per token through Moonshot's API and mirrored on aggregators like OpenRouter. Notably, this is a real price increase over Moonshot's prior flagship, Kimi K2.6.

Access methodPriceNotes
Moonshot API (direct)$3.00 / 1M input tokens, $15.00 / 1M output tokensListed rate as of launch; matches OpenRouter's model page
OpenRouter (aggregator)Same listed rate, routed across providersAdds redundancy/uptime; useful for comparing K3 against other models in one dashboard
Kimi K2.6 (predecessor, for comparison)$0.95 / 1M input tokens, $4.00 / 1M output tokensK3 costs roughly 3x more per token than K2.6

At $3/$15 per million tokens, K3's pricing now sits in the same range as Anthropic's Claude Sonnet-tier models — a meaningful jump for a Chinese open-weight lab, and reportedly the most expensive model any Chinese AI company has released to date. That's a departure from the usual playbook of undercutting Western labs on price, and it's worth factoring into any cost comparison against GLM-5.2 or LongCat-2.0, both of which remain significantly cheaper per token.

Check current Kimi K3 pricing and try the model

View Kimi K3 on the Kimi API Platform →

Kimi K3 Benchmarks: How It Stacks Up

The benchmark figures below come from Moonshot's own launch materials, cross-referenced against independent tracking from Artificial Analysis (whose Intelligence Index is one of the more widely cited cross-model comparisons) and Arena.ai's community-voted leaderboards. As with any vendor-published benchmark, treat these as a starting point, not a substitute for testing on your own workload.

BenchmarkKimi K3Claude Opus 4.8GPT-5.6 SolClaude Fable 5
Artificial Analysis Intelligence Index57.155.758.959.9
Arena.ai Frontend Code ArenaLeads leaderboard (independent, community-voted)
Context window1,000,000 tokensVaries by tierVaries by tierVaries by tier
Total / active parameters~2.8T total, 16 of 896 experts activeNot disclosedNot disclosedNot disclosed
LicenseModified MIT (open weights, pending July 27 release)Closed/proprietaryClosed/proprietaryClosed/proprietary

The pattern that emerges: on the independent Artificial Analysis Index, K3 lands ahead of Claude Opus 4.8 but behind GPT-5.6 Sol and Claude Fable 5. Moonshot's own self-reported benchmarks show a similar shape — K3 mostly beating Claude Opus 4.8 (max reasoning setting) and GPT-5.5 (high setting), while trailing Claude Fable 5 and GPT-5.6 Sol on the same tasks. The standout result is the Frontend Code Arena lead, since that's an independent, community-voted benchmark rather than a lab-published number — though "wins a code-focused arena" and "is the best coding model overall" aren't the same claim, and we haven't run our own side-by-side tests yet.

Key Features

  • 1M-token context window: Built on Kimi Delta Attention (KDA), a hybrid linear-attention architecture designed to make million-token context practical for long-horizon coding and agent workloads, not just technically possible.
  • Massive but sparse MoE architecture: 2.8 trillion total parameters with only 16 of 896 experts active per token (~1.8%), which is how Moonshot keeps a model this large usable for real-time chat and agent tasks.
  • Native vision: K3 handles image understanding natively rather than through a bolted-on vision adapter, per Moonshot's launch documentation.
  • Two purpose-built variants: K3 Max for general chat and agent work, and K3 Swarm Max for large-scale parallel processing workloads — a split that suggests Moonshot is targeting both individual developers and higher-throughput enterprise use cases.
  • Open license, weights pending: K3 ships under a Modified MIT license and is described by Moonshot as open-source, but as of this writing the downloadable model weights haven't been published yet — Moonshot has committed to releasing them by July 27, 2026.

How to Access Kimi K3

Kimi K3 is available through a couple of channels right now, with a self-hosting option coming later this month:

  • Kimi.com: Chat directly with K3 through Moonshot's consumer-facing web interface.
  • Kimi API Platform: Pay-per-token access at the rates above; see platform.kimi.ai for current documentation and pricing, since rates on new model releases tend to shift in the first few months.
  • OpenRouter: A unified API endpoint that routes requests to K3 alongside other models, useful for direct cost and output comparisons.
  • Self-hosting (coming soon): Moonshot has said full open weights will be published by July 27, 2026, at which point self-hosting becomes possible for teams with the GPU capacity to run a 2.8T-parameter model.

Compare Kimi K3's pricing and providers in one place

See Kimi K3 on OpenRouter →

Pros and Cons

ProsCons
Leads Arena.ai's independent Frontend Code Arena leaderboardOpen weights aren't downloadable yet — only promised by July 27, 2026
Scores above Claude Opus 4.8 on the Artificial Analysis Intelligence IndexRoughly 3x more expensive per token than Moonshot's own previous model, Kimi K2.6
1M-token context window built on an architecture (KDA) designed for real efficiency, not just a headline numberReleased July 16, 2026 — no long-term track record yet, no independent hands-on review from us at time of writing
Native vision support built inTrails GPT-5.6 Sol and Claude Fable 5 on the same independent intelligence index
Modified MIT license once weights ship — self-hostable for compliance-sensitive teamsUsing the hosted Kimi API routes data through a China-based provider, which may matter for data-residency-sensitive teams

Who Should Use Kimi K3?

Kimi K3 is worth evaluating if you are:

  • A developer doing frontend or agentic coding work — the Frontend Code Arena leadership and 1M-token context window suggest real strength here, though we recommend testing on your own codebase before committing.
  • Already comparing frontier-tier closed models — K3's Intelligence Index score puts it in the same conversation as Claude Opus 4.8, GPT-5.6 Sol, and Claude Fable 5, worth benchmarking against whatever you're currently paying for.
  • Planning ahead for a self-hosted deployment — if your team can wait for the July 27 open-weight release and has the GPU capacity for a 2.8T-parameter model, this is worth tracking closely.

Kimi K3 is probably not the right fit yet if you:

  • Need to self-host today — the downloadable weights haven't been released as of this writing.
  • Are cost-sensitive on a per-token basis — at $3/$15 per million tokens, K3 is no longer the budget option Moonshot's earlier models were, and cheaper open-weight alternatives like GLM-5.2 or LongCat-2.0 exist.
  • Have strict requirements against routing data through China-based infrastructure and can't wait for the self-hostable weights.
  • Want a model with a longer public track record before committing production workloads — K3 is, at the time of writing, about two days old.

Frequently Asked Questions

Is Kimi K3 free?

Not currently. The hosted API is pay-per-token ($3.00/$15.00 per million input/output tokens as of launch), and the open weights — which would allow free self-hosting — haven't been published yet. Moonshot has said they'll release the full weights by July 27, 2026. You can also chat with K3 for free (with usage limits) through Kimi.com.

Is Kimi K3 actually better than GPT-5.6 Sol or Claude Fable 5?

Not across the board. On the Artificial Analysis Intelligence Index, K3 (57.1) trails both GPT-5.6 Sol (58.9) and Claude Fable 5 (59.9), while beating Claude Opus 4.8 (55.7). K3's standout result is leading Arena.ai's independent Frontend Code Arena, which suggests real strength specifically on frontend coding tasks — but a leaderboard win on one benchmark doesn't mean it outperforms every closed frontier model on every task, and we haven't run our own side-by-side testing yet.

Can I self-host Kimi K3?

Not yet. K3 is licensed under a Modified MIT license and described as open-source, but the downloadable model weights hadn't been published as of this writing. Moonshot has committed to releasing them by July 27, 2026.

How big is Kimi K3's context window?

1,000,000 tokens, built on Moonshot's Kimi Delta Attention (KDA) hybrid linear-attention architecture, which is designed to make that context length computationally practical rather than just a marketing spec.

Why is Kimi K3 more expensive than Moonshot's previous models?

At $3.00/$15.00 per million input/output tokens, K3 costs roughly 3x more than Kimi K2.6 ($0.95/$4.00). Moonshot hasn't published a detailed explanation for the increase, but it puts K3 in the same pricing tier as Anthropic's Claude Sonnet-class models, a departure from the aggressive undercutting typical of Chinese open-weight labs.

Final Verdict

Kimi K3 is a significant release — the scale of the model (2.8T parameters), the independent Frontend Code Arena win, and an Intelligence Index score ahead of Claude Opus 4.8 are hard to wave away, especially from a lab that's iterating this fast. The financial-press attention (Fortune, CNBC, Yahoo Finance all running "DeepSeek shock" framing) is a signal of how seriously the market is taking Chinese open-weight labs right now, even if that framing says more about market psychology than about the model itself.

The honest caveats: the open weights that would make this a truly self-hostable, cost-free option aren't out yet, pricing has jumped sharply versus Moonshot's last flagship, and this review is based on launch-week specs and benchmarks rather than our own extended testing, because the model has been public for less than 48 hours. We'll revisit this review with hands-on notes once the weights ship and we've run K3 against real workloads.

Early rating: 8.2/10 — a genuinely frontier-level open-weight model with real independent-benchmark credibility, held back only by the not-yet-released weights and a real price increase over its predecessor.