Models·2 min read
By BitsMindsSource: MarkTechPost

Moonshot AI Releases Kimi K2.6: Open-Source Giant with 1T Parameters and 300-Agent Swarms

Moonshot AI has released Kimi K2.6, a 1-trillion-parameter open-source model capable of coordinating 300 parallel sub-agents across 4,000 steps — available on Hugging Face under a Modified MIT License.

Moonshot AI Releases Kimi K2.6: Open-Source Giant with 1T Parameters and 300-Agent Swarms
Share:

Moonshot AI has released Kimi K2.6, the latest iteration of its Kimi open-source model series, marking a significant leap in what open-weight models can accomplish for agentic and coding tasks. Released on April 20, 2026, K2.6 is available across Kimi.com, the Kimi App, API, and as downloadable weights on Hugging Face under a Modified MIT License.

The model uses a Mixture-of-Experts architecture with 1 trillion total parameters, activating 32 billion per token. It employs 384 experts with 8 activated per prompt, a 256K token context window, and a 400-million-parameter MoonViT vision encoder that processes both images and video natively. Multi-head Latent Attention (MLA) reduces memory requirements without sacrificing performance, and a SwiGLU activation function improves training stability and hardware efficiency across 61 transformer layers.

The headline capability is Kimi K2.6's agent swarm architecture, which allows the model to deploy up to 300 parallel sub-agents executing across 4,000 coordinated steps simultaneously — a threefold improvement over K2.5's limit of 100 agents. In one published case study, the model autonomously optimized a financial matching engine over 13 hours, achieving a 185% medium throughput increase and 133% performance gain through systematic code modifications with minimal human oversight.

On the HLE-Full benchmark — one of the most demanding agentic evaluations — Kimi K2.6 scores 54.0, edging out GPT-5.4 at 52.1 and Claude Opus 4.6 at 53.0. It also scores 58.6 on SWE-Bench Pro, which measures real-world GitHub issue resolution. Moonshot says the model shows particular strength in Rust development and front-end generation from natural language descriptions.

For deployment, the model runs on vLLM, SGLang, or KTransformers with two inference modes: Thinking mode for complex multi-step reasoning and Instant mode for lower-latency interactive use. A new "claw groups" feature enables seamless human-AI task handoffs within the same agentic loop. With open weights and competitive benchmark performance, Kimi K2.6 cements Moonshot AI's position as a major challenger to closed frontier models from OpenAI and Anthropic.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Many kinds of input, one Qwen context On a deep violet field, a filmstrip, an audio waveform, a landscape photograph and a text page curve toward a single luminous sphere bearing the official purple Qwen star. A label below reads “1M context”, representing text, images, audio and video sharing one million-token context window. Aa 1M context BITSMINDS.COM
Models

Qwen3.8-Omni-Flash Undercuts Gemini on Audio and Video

Hill Climb: GPT-6 Astra versus Claude Fable 5.1 A teal desert buggy climbs a sandstone ridge on the left. An orange jeep climbs a green hill on the right, under an arc of gold coins. The two landscapes meet at a diagonal divide. 7 GPT-6 ASTRA CLAUDE FABLE 5.1 VS HILL CLIMB BITSMINDS.COM
Models

GPT-6 Astra vs Claude Fable 5.1: Hill Climb

Gemini 3.8 Live: thinking while the conversation continues An editorial illustration in a dark blue and violet room. A carefully drawn studio microphone and a tilted smartphone flank a translucent speech bubble carrying the multicoloured Gemini star. A continuous luminous audio waveform travels between them. Above the bubble, a separate arc connects small search, reasoning and completion symbols, representing background work continuing during a spoken conversation. The phone screen and visual paths are conceptual, not a reproduction of Google's actual interface or internal reasoning. Original vector illustration for BitsMinds, gemini-3-8-live-extended-thinking-voice. 16 September 2026. GEMINI 3.8 LIVE EXTENDED THINKING Gemini 3.8 LIVE The conversation continues Thinking. Still talking. BITSMINDS.COM
Models

Gemini 3.8 Live Tops Voice AI and Undercuts GPT