I made a video on the Kimi K3 launch because the headline and the receipts point in different directions. Moonshot AI shipped it on July 16, 2026, a 2.8 trillion parameter model with a 1 million token context window. Within hours it hit number one on the Frontend Code Arena, ahead of Claude Fable 5 and GPT 5.6, and the Arena post crossed 18 million views. Then I looked at what real testers found.
The leaderboard is one leaderboard
In the video I walk through where Kimi actually stands. It ranked first in six of seven front-end categories on the Arena, but on Artificial Analysis it scored 57, near Opus 4.8 and GPT 5.5 and behind Fable 5 and GPT 5.6. Moonshot’s own launch article says K3 still trails Fable and Soul overall. So my read is that Kimi won some coding and front-end tests, not that it caught the best closed models across the board.
Same prompt, very different bill
The number that made me want to record this: Chase AI gave Kimi, Fable, and Soul the same complex front-end task. Kimi used about 21.5M tokens and ran 1 hour 33 minutes. Fable used about 3.5M tokens in 17 minutes. Soul used about 5.5M tokens in 25 minutes. The final results looked close, the routes did not. Jeremy Chone saw the same in a smaller test, Kimi about five times slower than Fable.
Other testers back it up. AI Coding Daily ran K3 through 25 prompts across five projects and it scored 22.5 out of 25, the best from a Chinese model, but the run took the whole day. Mehul Mohan built a real product with it over 24 hours, called its first infrastructure plan mostly wrong, and gave it a 6 out of 10 as a daily driver. My takeaway in the video is one line: price per token is not cost per completed task.
Can you actually run it yourself?
Full weights release July 27, which usually implies self-host and save. The arithmetic gets in the way, and I do the math on camera. At a 4-bit format, 2.8 trillion parameters is about 1.4TB just for the weights, before runtime and context. Moonshot recommends a super node with 64 or more accelerators for efficient inference. So open is not local. You can inspect and host the weights with data-center hardware, but most people will still call a hosted endpoint.
What’s genuinely new
There is real engineering here. K3 is a mixture-of-experts model with 896 experts that fires only 16 per token, and it adds Kimi Delta Attention, which claims up to 6.3x faster decoding at million-token context. Listed API price is $15 per million output tokens with cached input at 30 cents per million, well under Fable’s output price. But three days after launch Moonshot paused new subscriptions because demand hit capacity.
My verdict in the video: test K3 against your own prompts and measure cost per correct answer, not the Arena rank. Watch it for every figure and the source posts, and subscribe for more practical AI engineering breakdowns.
I specialize in local AI and cloud architecture calls like this one. Book a call at cloudyeti.io/meet.