Models·3 min read·Moonshot AI

Moonshot's Kimi K3 Is the Largest Open-Weight Model Ever — and It Just Beat Fable 5 at Frontend Code

China's Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight mixture-of-experts model — the biggest ever — that took the #1 spot on Arena.ai's Frontend Code leaderboard at 1,679 points, ahead of Claude Fable 5. It has a 1M-token context, activates just 16 of 896 experts per token, and full weights are promised on Hugging Face by July 27.

MOONSHOT AI · OPEN WEIGHTS Kimi K3 The largest open-weight model ever — now #1 at frontend code, ahead of Fable 5 #1 Frontend Code Arena · 1,679 pts · ahead of Fable 5 2.8T · MoE 1M context 16/896 active Weights Jul 27 Open weights on Hugging Face · native multimodal · $0.30/$3 in · $15 out per million tokens
Share:

China's open-model push just cleared its most symbolic bar yet. On July 16, Moonshot AI released Kimi K3, a 2.8-trillion-parameter mixture-of-experts model it calls the largest open-weight AI system ever built — and it promptly took the #1 spot on Arena.ai's Frontend Code leaderboard at 1,679 points, edging out Anthropic's Claude Fable 5. For a downloadable model out of China to top a live coding arena, ahead of the most capable model the American labs sell to the public, is a milestone the open-weight camp has been chasing for a year.

The scale is the headline. At 2.8 trillion total parameters, K3 dwarfs the previous "largest open" claims — including the 975-billion-parameter Inkling that Thinking Machines shipped days earlier — yet it's engineered to stay cheap to run. As a mixture-of-experts model it activates just 16 of its 896 experts per token, roughly 1.8% of the pool, so the compute cost of any single forward pass is a fraction of the parameter count. It pairs that with a 1-million-token context window, native multimodal input, and maximum reasoning enabled at launch.

On the benchmark that made the news, the jump is dramatic. Kimi K3 didn't just place — it climbed 17 spots from K2.6's #18 to first, and led six of the seven frontend domains Arena measures, from layout and styling to interactive components. That's a pointed result: front-end code generation has been one of the areas where closed frontier models like Fable 5 and GPT-5.6 held a comfortable lead, and it's exactly the kind of everyday developer task that decides which model teams actually reach for.

There's an honest asterisk. As of launch the full weights aren't public — Moonshot has committed to releasing them on Hugging Face by July 27, so today K3 is "open-weight" in promise more than in practice, usable through the API while the download is pending. Pricing undercuts the Western frontier sharply: about $0.30 per million cache-hit input tokens, $3 on cache misses, and $15 per million output tokens. Combined with self-hosting once the weights land, that puts frontier-class frontend coding within reach of anyone willing to run the model themselves.

The strategic message is louder than any single score. Built under US export controls that have squeezed China's access to top-end compute, K3 is proof that Chinese labs can still ship frontier-scale open models — and give them away — even as Washington tries to slow them down. It extends a wave that already includes DeepSeek, Zhipu's GLM, and Moonshot's own Kimi line, and it sharpens the industry's central fault line: as American leaders like Anthropic and OpenAI defend closed, premium models, China is betting that the future of AI is something you can download for free — and, with Kimi K3, something that now wins.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

SPACEXAI · CODING & AGENTS Grok 4.5 Opus-class coding — at fast-model speed and a fraction of the cost AVG. OUTPUT TOKENS · SWE-BENCH PRO TASK Opus 4.8 67k Grok 4.5 16k 4.2× fewer tokens $2 / $6 per 1M tokens 80 TPS output speed #1 SWE Marathon Trained alongside Cursor · Terminal-Bench 2.1 83.3% · SWE-Bench Pro 64.7% · free for now in Grok Build & Cursor
Models

SpaceXAI's Grok 4.5 Bets on Efficiency Over the Crown — Opus-Class Coding at 4.2× Fewer Tokens

PRISMML · ON-DEVICE AI Bonsai 27B A 27B multimodal model that runs on your phone — fully offline ~54 GB · FP16 1-BIT QAT multimodal · 262K context 3.9 GB · 1-bit Native low-bit QAT · keeps >90% of full precision · 11 tok/s on iPhone 17 Pro · Apache 2.0 on Hugging Face
Models

PrismML's Bonsai 27B Squeezes a 27-Billion-Parameter Model Onto Your Phone — at 3.9GB, Fully Offline

THINKING MACHINES LAB · MIRA MURATI Inkling An open-weight, multimodal foundation model — built to be customized OPEN WEIGHTS · APACHE 2.0 975B MoE · 41B active 1M token context 45T tokens · 4 modalities Text · Image · Audio · Video  •  download on Hugging Face, fine-tune on Tinker
Models

Mira Murati's Thinking Machines Releases Inkling — a 975B Open-Weight, Multimodal Model Built to Be Customized