China’s Kimi K3 Is Out—And Beats Claude Fable and GPT 5.6 Sol on Key Benchmarks

by shayaan
US Government Says China's Best AI Models Lag Behind. Experts Aren't So Sure

In short

  • Moonshot AI released Kimi K3 on July 16: a 2.8 trillion-parameter open-weight model that beats US labs on specific specialized benchmarks.
  • K3 has the same price as Claude Sonnet 5 ($3 per million input tokens, $15 per million output tokens) and scores closer to Fable 5.
  • The full model weights (the files that allow anyone to run, refine, or build on the model locally) will be released under a modified MIT license on July 27, making K3 the largest freely available AI model in history.

Moonshot AI just released the largest Chinese open source model ever released, and it surpassed Claude Fable 5 in scripting.

Towards AIs write Elo– a benchmark where models write real scripts that are judged blindly against published versions, scored using the same Elo system that ranks chess players – puts Kimi K3 at 2,840, above Fable 5 (max) at 2,760. That’s a ranking that the Anthropic team has historically dominated.

K3 also claimed the top spot on Arena AI’s Frontend Code Leaderboard – a ranking constructed from thousands of pairwise human votes on code generation tasks, again with an Elo score – with 1,679 to Fable 5’s 1,631. First place in six of the seven frontend domains.

The Artificial Analysis Intelligence Index– a score constructed from nine independent evaluations covering coding, reasoning, agentic work and knowledge, rated from 0 to 100 – puts K3 at 57, with Claude Fable 5 at 60, GPT-5.6 Sol at 59 and Claude Opus 4.8 at 56. That places K3 as the third most capable model on the composite, while Fable 5 only beats it by 3%.

If you want an idea of ​​what it can do, this is one zero shot result of a prompt asking the model to build an iOS clone. For comparison: this is the best approach shared on social media using GPT 5.6 Sol and a much verbose prompt.

See also  Which Platform Builds the Best AI Agents? We Test ChatGPT, Claude, Gemini and More

What this thing actually is

K3 contains 2.8 trillion parameters (the numerical values ​​that store a model’s knowledge) in a mix of expert architecture. A mix of experts splits these parameters into 896 ‘expert’ subnetworks and activates only a fraction for a given task. This way you get borderline intelligence without melting the server space.

“It is the world’s first open source model in the 3 trillion parameter class, designed for groundbreaking intelligence scenarios including long-horizon coding, knowledge work and reasoning,” said Moonshot AI. say. That’s not marketing theater: DeepSeek’s V4-Pro scores on 1.6 trillion parameters, Moonshot’s own K2 on one trillion. K3 roughly doubles the nearest open weight competitor on the size chart.

It comes with a context window of one million tokens – tokens are the basic unit of information that an AI processes, each about three-quarters of a word – native image and video understanding, and always-on reasoning.

Two architectural techniques underlie the efficiency gains. Kimi Delta Attention accelerates long string decoding, up to 6.3x faster on contexts of millions of tokens. Attention Residuals selectively routes information across model layers instead of uniformly collecting it, adding approximately 25% training efficiency at less than 2% additional computational cost, which together deliver approximately 2.5x better scaling efficiency than K2.

Benchmarks are fun, prizes are more fun

Kimi K3 costs $3 per million input tokens and $15 per million output tokens – the same rate as Claude Sonnet 5, Anthropic’s mid-range model. The difference is that Sonnet 5 is Anthropic’s mid-range offering; K3 is three points below Fable 5 on the Artificial Analysis composite. Per job in this nine-benchmark suite, K3 costs $0.94 versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8.

See also  XRP Price Surges Above Key Level, Bulls Take Full Control

In other words, this model offers top performance at mid-level prices.

As Decrypt reported in May, the price gap between Chinese and US border AI was 15 to 30x earlier this year. K3 doesn’t undercut DeepSeek rates (its prices are comparable to a Western mid-range model), but it delivers almost groundbreaking performance at that level. For teams building on the API, this means a major cost improvement.

If Anthropic follows through with its intention to make Fable 5 available via API only, K3 will become the closest open-weight alternative to any model currently ranked second in the industry – at half the cost per job of Opus 4.8. That’s the scenario the benchmark hunters are already working on.

The launch of K3 is the argument that proponents of US chip export controls don’t want to have. The US restricted exports of Nvidia’s H800 GPUs to China in late 2023; Moonshot confirmed that it has trained previous models on those chips. K3’s own benchmark documentation refers to H200s and what the company calls “an alternative vendor GPGPU” – broadly interpreted as Huawei Ascend hardware – without specifying where that hardware is located.

Moonshot AI President Yutong Zhang framed the restriction directly in Davos this year, per Silicon Republic: “We knew we didn’t have the luxury of simply scaling up computing power… That forced us to focus on fundamental research and efficiency.” Bank of America analysts wrote in a post-launch note that K3 proves that “pre-training scaling, combined with architectural innovation, can still deliver incremental gains for Chinese flagship models under those constraints.”

See also  Microsoft, Aptos Labs To Create AI-Powered Web3 Assistant To Simplify User Transition From Web2 - Microsoft (NASDAQ:MSFT)

Moonshot is one of the so-called AI Tiger startups that collectively changed the global modeling landscape without access to the chips Washington said it needed. Whether that is an argument for stricter export controls or an argument that they don’t work is a policy question that Washington has not yet resolved.

You have to read the asterisk

K3’s hallucination rate on AA-Omniscience – a metric that measures how often a model confidently makes up an answer it doesn’t know – rose from 39% to 51% compared to predecessor K2.6. More correct answers overall; also more invented. The model also acknowledges in its own documentation that it can be “excessively proactive,” making unexpected decisions on behalf of a user during long autonomous tasks.

For teams using the Kimi K2.6-based tooling and looking to upgrade, K3 is a meaningful step forward on most fronts, but that hallucination delta is worth testing before trusting it with anything that needs to be accurate.

If you want to try it for free, you can. It is available on Kimi’s official website. But good luck: the servers are so full that tasks are constantly interrupted due to traffic restrictions, making it barely usable. A better alternative is to pay for a subscription or use it via an API.

Weights will be released on July 27th. These will be available to large corporations and businesses. No domestic GPU, no matter how large, can currently handle a model of this size.

Daily debriefing Newsletter

Start every day with today’s top news stories, plus original articles, a podcast, videos and more.



Source link

You may also like

Latest News

Copyright © Sovereign Wealth Signals