AI News

Latest updates from the world of artificial intelligence

Shieldstral Mistral's 3B open-weights safety guard — you write the policy at runtime policy.txt "Flag content that promotes violence or self-harm." plain language, at inference 3B Apache 2.0 open weights calibrated score 0.97 BLOCK one forward pass, yes/no logits 84.9 F1 text · 83.8 F1 multimodal · 12 languages · runs on one 16 GB GPU BITSMINDS.COM
Models

Mistral's Shieldstral: A 3B Guard That Reads Your Policy

Mistral's new open-weights safety classifier takes the moderation policy as a plain-language prompt at inference time — no retraining — and matches guard models seven times its size while running on a single 16 GB GPU.

Mistral AI

Read more →
MAI PLAYGROUND · HIDDEN EARLY ACCESS MAI-Realtime does not take turns. It listens and speaks together. OUT · SPEAKS IN · LISTENS hears you barge in never waits for a pause SAME INSTANT VOICES Victoria · Grant LANGUAGES 17, switched mid-sentence STATUS unannounced, no pricing BITSMINDS.COM
Models

Microsoft's MAI-Realtime Aims at OpenAI's Voice Slot

A hidden entry in Microsoft’s MAI Playground reveals MAI-Realtime, a full-duplex speech model that listens and speaks simultaneously across seventeen languages, with two configurable ways of deciding whose turn it is. Microsoft has confirmed nothing — but the component it would displace inside Azure’s Voice Live service is OpenAI’s GPT-Realtime.

TestingCatalog

Read more →
QWEN3.8-MAX Weights next week.
Models

Alibaba Opens Qwen3.8-Max to the World — and Says the Weights Come Next Week

The 2.4-trillion-parameter flagship went broadly available today through Alibaba Cloud's APIs and the new QwenWork beta, with open weights promised for next week. The parameter count is the least interesting part: this is Alibaba reversing course after keeping recent flagships closed. The performance claim behind it — second only to Claude Fable 5 — still has no independent benchmark, and would require a 14-point jump in one generation.

South China Morning Post

Read more →
AI & MATHEMATICS · OPENAI ‘ASTRA’ Ten Problems, Decades Old Theorem. machine-checked ×10 This time the proof comes with a receipt.
Models

OpenAI Names Its Next Model Family Astra — and Says It Solved Ten Decade-Old Math Problems

An internal version of Astra resolved ten open problems across group theory, high-dimensional geometry, coding theory, quantum complexity, lattice cryptography and extremal combinatorics, OpenAI says — with the proofs formalised in Lean as machine-checkable certificates. That is a genuine methodological upgrade over May's Erdős result, which rested on expert sign-off. It also leaves the one question a proof checker cannot answer.

The Decoder

Read more →
OPEN MODELS · DEEPSEEK V4 FLASH They Didn't Make It Bigger same model, same price what it scores Nothing about it grew. It got better anyway.
Models

DeepSeek Gained Ten Index Points Without Changing the Model — and Undercut Luna a Day After Its 80% Cut

DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, up from 40 in April, with identical architecture, the same 13B active parameters and the same price list — the gain came entirely from post-training. It now sits one point behind GPT-5.6 Luna at roughly 60% lower cost per task, with a 98% cache discount. The nuance most coverage will skip: accuracy was unchanged, and the jump came mainly from a reduced hallucination rate.

Artificial Analysis

Read more →
GPT-5.6 · API PRICING Prices Just Fell 80% Only one of them got cheaper.
Models

OpenAI Cut Luna's Price 80% and Left Sol Untouched — the GPT-5.6 Range Is Now 25× Wide

GPT-5.6 Luna drops to $0.20 / $1.20 per million tokens from $1 / $6, and Terra to $2 / $12 from $2.50 / $15. Sol is unchanged, and its new Fast option costs twice as much for about 2.5× the speed. OpenAI credits serving efficiency — but a roughly 20% cost improvement explains Terra's cut and comes nowhere near explaining Luna's, which leaves the gap between the cheapest and best model five times wider than it was a month ago.

OpenAI

Read more →
BENCHMARKS · ARC-AGI-3 Same Model. Same Test. retained reasoning compaction same benchmark Two switches. Three times the score.
Models

OpenAI Says Two Settings Tripled Its ARC-AGI-3 Score — ARC Prize Says That Isn't the Same Test

OpenAI reported that enabling retained reasoning and compaction in its Responses API took GPT-5.6 Sol from a 13.3% baseline to 38.3% on ARC-AGI-3's public set, using roughly six times fewer output tokens — a figure above Claude Opus 5's 30.2% record. ARC Prize says its official scores use a standardised setup with no provider-specific settings, and François Chollet allows general-purpose API settings only if clearly reported. Both sides have a point, and the comparison still isn't apples to apples.

OpenAI

Read more →
MOONSHOT AI · KIMI K3 2.8 Trillion. Now Free. Weights live on Hugging Face Now find a machine that can run it
Models

The Largest Open Model Ever Just Went Free: Kimi K3's 2.8-Trillion-Parameter Weights Are Live

Moonshot AI published Kimi K3's full weights on Hugging Face at 00:00 UTC on July 27, exactly as promised — 2.8 trillion parameters under Apache 2.0, the biggest open-weight release in history. It lands three days after 50 companies signed a letter defending open models, and days after the White House accused Moonshot of distilling an Anthropic model to build it.

Moonshot AI

Read more →
CLAUDE OPUS 5  vs  GPT-5.6 One Dial vs Three Models Claude Opus 5 vs GPT-5.6 Sol Terra Luna So which lineup actually wins?
Models

Claude Opus 5 vs GPT-5.6: One Dial vs Three Models

OpenAI split GPT-5.6 into Sol, Terra and Luna. Anthropic shipped one model with an effort dial. Put both on the same cost-independent index and the settings line up exactly — including one result that should change how you configure your API calls.

BitsMinds Analysis

Read more →
BITSMINDS ANALYSIS · CLAUDE LINEUP Opus 5 vs Fable 5 vs Sonnet 5 What actually changed — on price, specs, and the benchmarks Claude Fable 5 $10 / $50 33.7% Claude Opus 5 $5 / $25 43.3% Claude Sonnet 5 $3 / $15 not measured Opus 4.8 · legacy $5 / $25 18.7% Bars: Frontier-Bench v0.1 (coding), as published by Anthropic All three current models share a 1M-token context · prices per million input / output tokens
Models

Claude Opus 5 vs Fable 5 vs Sonnet 5 vs Opus 4.8: The Benchmark Comparison

Opus 5 costs the same as the Opus 4.8 it replaces, beats the double-priced Fable 5 on Anthropic's own coding charts, and burns a fraction of the tokens doing it. We lined up every published number — specs, prices, benchmarks — and flagged exactly where the comparisons don't hold.

BitsMinds Analysis

Read more →
ANTHROPIC · OFFICIAL LAUNCH Claude Opus 5 Is Here The rumors were right — here’s what actually changed
Models

Claude Opus 5 Is Officially Here — 1M Context, a New 'xhigh' Mode, and the Rumors Were Right

Anthropic launched Claude Opus 5 on July 24, replacing Opus 4.8 as its workhorse model: a 1-million-token context window, a new 'xhigh' reasoning-effort level above today's max, and unchanged $5/$25-per-million-token pricing. The 'Honeycomb' leak, the Vertex AI sightings, and the specs rumor all turned out accurate — Fable 5 remains Anthropic's top-tier model.

Anthropic

Read more →
GOOGLE · GEMINI Three Flash models ship… …while Gemini 3.5 Pro is still nowhere to be seen 3.6 Flash −17% output tokens 3.5 Flash-Lite cheap · beats Gemini 3 3.5 Flash Cyber finds & fixes bugs Gemini 3.5 Pro promised June · still not shipped → Meanwhile, Google says it has begun its biggest pretraining run yet — Gemini 4 Shipped July 21 via the Gemini API & Gemini Enterprise · the price war grinds on
Models

Google Ships Three New Gemini Flash Models — but 3.5 Pro Is Still Missing, and It's Already Pretraining Gemini 4

On July 21 Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a security-tuned 3.5 Flash Cyber — useful, cheap, efficient. But the flagship Gemini 3.5 Pro is still nowhere to be seen, months past its June target, and Google says it has already begun its most ambitious pretraining run yet: Gemini 4.

Google

Read more →
RUMOR WATCH · UNCONFIRMED Claude Opus 5 — the evidence board Cursor · Jul 9 “Claude Honeycomb EAP” — then gone Vertex AI · ×2 listing spotted twice, vanished both times 5? the next flagship? 1M context? leaked, unverified “xhigh” mode? launch “this week”?? No official API ID · no system card · every thread on this board is still a rumor
Models

Claude Opus 5 Rumor Watch: The 'Honeycomb' Leak, a Reported 1M Context, and Talk of a Launch This Week

The internet is convinced Claude Opus 5 ships any day now — a model codenamed 'Honeycomb' surfaced briefly in Cursor and a Google Vertex AI listing, with leaked talk of a 1M-token context and an 'xhigh' reasoning mode. Here's the honest version: what's actually been spotted, what's pure speculation, and why the timing rumor won't die.

Leak Reports

Read more →
MOONSHOT AI · OPEN WEIGHTS Kimi K3 The largest open-weight model ever — now #1 at frontend code, ahead of Fable 5 #1 Frontend Code Arena · 1,679 pts · ahead of Fable 5 2.8T · MoE 1M context 16/896 active Weights Jul 27 Open weights on Hugging Face · native multimodal · $0.30/$3 in · $15 out per million tokens
Models

Moonshot's Kimi K3 Is the Largest Open-Weight Model Ever — and It Just Beat Fable 5 at Frontend Code

China's Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight mixture-of-experts model — the biggest ever — that took the #1 spot on Arena.ai's Frontend Code leaderboard at 1,679 points, ahead of Claude Fable 5. It has a 1M-token context, activates just 16 of 896 experts per token, and full weights are promised on Hugging Face by July 27.

Moonshot AI

Read more →
SPACEXAI · CODING & AGENTS Grok 4.5 Opus-class coding — at fast-model speed and a fraction of the cost AVG. OUTPUT TOKENS · SWE-BENCH PRO TASK Opus 4.8 67k Grok 4.5 16k 4.2× fewer tokens $2 / $6 per 1M tokens 80 TPS output speed #1 SWE Marathon Trained alongside Cursor · Terminal-Bench 2.1 83.3% · SWE-Bench Pro 64.7% · free for now in Grok Build & Cursor
Models

SpaceXAI's Grok 4.5 Bets on Efficiency Over the Crown — Opus-Class Coding at 4.2× Fewer Tokens

Grok 4.5, trained alongside Cursor, doesn't top the leaderboards — Fable 5 and GPT-5.5 still edge it on raw coding benchmarks. But it resolves SWE-Bench Pro tasks in about 15,954 output tokens versus ~67,020 for Opus 4.8, is served at 80 tokens/sec, and costs just $2/$6 per million tokens. xAI's pitch: Opus-class results at a fraction of the time and cost.

xAI

Read more →
PRISMML · ON-DEVICE AI Bonsai 27B A 27B multimodal model that runs on your phone — fully offline ~54 GB · FP16 1-BIT QAT multimodal · 262K context 3.9 GB · 1-bit Native low-bit QAT · keeps >90% of full precision · 11 tok/s on iPhone 17 Pro · Apache 2.0 on Hugging Face
Models

PrismML's Bonsai 27B Squeezes a 27-Billion-Parameter Model Onto Your Phone — at 3.9GB, Fully Offline

PrismML has open-sourced Bonsai 27B, a 1-bit build of Qwen3.6-27B that shrinks a model needing ~54GB at full precision down to 3.9GB — small enough to run on an iPhone at 11 tokens/sec while keeping more than 90% of full-precision performance. It's multimodal, handles a 262K-token context, and ships under Apache 2.0.

PrismML

Read more →
THINKING MACHINES LAB · MIRA MURATI Inkling An open-weight, multimodal foundation model — built to be customized OPEN WEIGHTS · APACHE 2.0 975B MoE · 41B active 1M token context 45T tokens · 4 modalities Text · Image · Audio · Video  •  download on Hugging Face, fine-tune on Tinker
Models

Mira Murati's Thinking Machines Releases Inkling — a 975B Open-Weight, Multimodal Model Built to Be Customized

Thinking Machines Lab's first broadly available model is open. Inkling is a 975B-parameter mixture-of-experts model (41B active), trained on 45T tokens of text, image, audio, and video, with a 1M-token context — released under Apache 2.0 and tuned on the lab's Tinker platform.

Thinking Machines Lab

Read more →
BYTEDANCE · IMAGE GENERATION Seedream 5.0 Pro China’s new flagship reasons about prompts — and renders native 2K Native 2K 14 languages $0.075 / image GPT Image 2-level Prompt reasoning · point-and-lasso editing · on-image text · native 2K output
Models

ByteDance's Seedream 5.0 Pro Pushes China to the Front of AI Image Generation

ByteDance's new flagship image model reasons about prompts, edits with point-and-lasso precision, writes on-image text in 14 languages, and outputs native 2K — at $0.075 an image. Observers already put its quality at GPT Image 2 level.

TestingCatalog

Read more →
Anthropic Extends Free Fable 5 Again — the Deadline That Keeps Moving
Models

Anthropic Extends Free Fable 5 Again — the Deadline That Keeps Moving

Anthropic pushed free Fable 5 access to July 19 — its third extension in 18 days — as GPT-5.6 undercuts it on price and Grok 4.5 goes live. An Opus 5 leak in Cursor fuels the theory that the extensions are a bridge to the next flagship.

Anthropic

Read more →
GPT-5.6 Launches After a Government Delay — and Sol Tops the Coding Charts
Models

GPT-5.6 Launches After a Government Delay — and Sol Tops the Coding Charts

OpenAI shipped GPT-5.6 on July 9 as a three-tier family — Sol, Terra, Luna — after a US-government delay. Sol tops TerminalBench 2.1 at 91.9% and is ~54% more token-efficient on agentic coding, though Anthropic's Fable 5 still leads SWE-Bench Pro.

OpenAI

Read more →
MEITUAN · OPEN-WEIGHT CODING MODEL LongCat-2.0 1.6-trillion-parameter MoE Trained and served end-to-end on ≈50,000 domestic chips — zero Nvidia GPUs. SWE-bench Pro 59.5 1M context From $0.30 / 1M 50,000 DOMESTIC CHIPS 0 × NVIDIA BITSMINDS.COM
Models

LongCat-2.0: A 1.6T Coding Model Trained on Chinese Chips

Meituan open-sourced LongCat-2.0 on June 30 — a 1.6-trillion-parameter mixture-of-experts coding model trained and served end to end on more than 50,000 Chinese-made chips, with no Nvidia hardware in the loop. After two months quietly topping OpenRouter under the alias “Owl Alpha,” the MIT-licensed model claims 59.5% on SWE-bench Pro at a fraction of GPT-5.5’s price.

VentureBeat

Read more →
GEMINI 3.5 PRO · GENERAL AVAILABILITYGoogle’s June deadline slips to JulyJUNE 202630TARGET MISSEDJULY 2026NEW WINDOWGoogle cites token-efficiency and long-horizon tuning after early-tester feedback.BITSMINDS.COMSource: TechTimes
Models

Gemini 3.5 Pro Slips to July as Google Misses Deadline

Google let its self-imposed June deadline for Gemini 3.5 Pro pass without a launch, and now says the frontier model will reach general availability in July. The company points to token-efficiency and long-horizon agentic tuning — even as a wave of senior DeepMind researchers heads for the exits.

TechTimes

Read more →
ANTHROPIC · NEW MODELClaude Sonnet 5Near-Opus coding at intro $2 / $10 per million tokens58.1Sonnet 4.663.2Sonnet 569.2Opus 4.8Agentic coding benchmark · SWE-bench Pro (%)BITSMINDS.COM
Models

Claude Sonnet 5 Lands, Closing In on Opus 4.8

Anthropic released Claude Sonnet 5 on June 30 — its most agentic mid-tier model yet, scoring 63.2% on SWE-bench Pro and matching Opus 4.8 on some knowledge work, with introductory pricing of $2/$10 per million tokens. It is now the default model on Claude’s Free and Pro plans.

Anthropic

Read more →
Ornith-1.0 cracks the open-weights coding race SWE-Bench Verified (% resolved) · MIT license · 9B / 31B / 35B / 397B 87.6 Opus 4.8 82.4 Ornith-1.0 397B 80.8 Opus 4.7 OPEN WEIGHTS BITSMINDS.COM
Models

Ornith-1.0: Open Coding Models That Self-Scaffold

DeepReinforce has open-sourced Ornith-1.0, an MIT-licensed family of coding models (9B to 397B) that learn to write their own agent scaffolds during reinforcement training. The 397B flagship resolves 82.4% of SWE-Bench Verified issues — second only to Claude Opus 4.8 — while the 9B model runs on a single 80GB GPU, putting frontier-class agentic coding within reach of self-hosters.

MarkTechPost

Read more →
GPT-5.6 ARRIVES IN THREE TIERS OpenAI ships Sol, Terra and Luna — under U.S. government limits FLAGSHIP Sol Frontier coding & security $5 / $30 in / out per 1M tokens BALANCED Terra High-volume business tasks $2.50 / $15 in / out per 1M tokens EFFICIENT Luna Fast, low-cost everyday work $1 / $6 in / out per 1M tokens BITSMINDS.COM
Models

GPT-5.6 Is Here: OpenAI Ships Sol, Terra and Luna

OpenAI released GPT-5.6 on June 27 as a three-model family — the flagship Sol, the balanced Terra and the fast, cheap Luna — but only as a limited preview to about 20 U.S. government-approved partners. Sol leads agentic-coding benchmarks (91.9% on Terminal-Bench 2.1) and matches prior security results using roughly a third of the output tokens, with general availability promised in the coming weeks.

OpenAI

Read more →
BYTEDANCE LAUNCHES SEED 2.1 Doubao Seed 2.1 Pro & Turbo — built for the coding and agent era #8 WORLDWIDE CODE ARENA: FRONTEND 1539 Draws level with Claude Opus 4.6 JUN 24 “comparable to GPT-5.5” on Volcano Engine Pro ¥6/¥30 · Turbo ¥3/¥15 BITSMINDS.COM
Models

ByteDance’s Seed 2.1 Pro Pulls Level With Claude Opus 4.6

ByteDance launched its Doubao Seed 2.1 Pro and cheaper Turbo models on June 24, built for the “coding and agent era.” Seed 2.1 Pro scored 1539 on the Code Arena: Frontend leaderboard — eighth in the world and level with Claude Opus 4.6 — while ByteDance claims three core capabilities are comparable to GPT-5.5.

AIbase

Read more →
MODEL WATCH · ANTHROPIC · RUMORED JUN 24 Claude Sonnet 5 — rumored. 'Fennec' is floated for the week of June 23 — but Anthropic has confirmed nothing. 60% LAUNCH ODDS BitsMinds estimate for this week — not a market WHAT'S RUMORED Codename 'Fennec', from a leaked log SWE-Bench around 82–92% — estimated Feb's 'Fennec' shipped as Sonnet 4.6 BITSMINDS.COM Rumor · Vertex AI leak · Crypto Briefing
Models

Claude Sonnet 5 Rumors Resurface — Still Nothing Confirmed

Reports this week floated Claude Sonnet 5 (codename 'Fennec') for the week of June 23, alongside GPT-5.6 — but Anthropic has confirmed nothing. And the same 'Sonnet 5' leak in February shipped as Sonnet 4.6, a reminder to read codenames with caution.

Crypto Briefing

Read more →
SAKANA · FUGU One Model to Orchestrate the Rest T THINKER W WORKER V VERIFIER S SYNTHESIZER FUGU BITSMINDS.COM
Models

Sakana AI's Fugu Turns Model Orchestration Into a Product

The Tokyo lab made Fugu generally available on June 22 — a foundation model trained to command a pool of other AIs through one OpenAI-compatible API, pitched as a hedge against single-vendor lock-in.

Sakana AI

Read more →
Z GLM-5.2 by Z.ai · MODEL REVIEW 62.1 SWE-bench Pro #1 open model · beats GPT-5.5 74.4 FrontierSWE ≈ Opus 4.8 (75.4) Open-weight · MIT · 744B MoE · 1M context BITSMINDS.COM
Models

GLM-5.2 Review: The Open Model That Out-Codes GPT-5.5

GLM-5.2 is the strongest open-weight coding model yet: MIT-licensed, 744B MoE, 1M context. It beats GPT-5.5 on SWE-bench Pro and trails Opus 4.8 by a point on FrontierSWE — at a fraction of the cost. Our review, with benchmark charts.

The Decoder

Read more →
GPT-5.6 WHAT WE KNOW SO FAR — STATUS: UNCONFIRMED NOW Pre-train Canary Red-team Early access Launch CODEX ROUTING LOG ▸ route: gpt-5.6 · seen May 13 ▸ route: gpt-5.5 · then reverted Rumored · 1.5M context Rumored · ~1/3 cheaper vs Fable 5 Odds 83% · Jun 22–28 BITSMINDS.COM
Models

GPT-5.6: Everything We Know — Now Launched as Sol, Terra & Luna

GPT-5.6 is here: OpenAI shipped it on July 9, 2026 as a three-tier family — Sol, Terra and Luna — after a rare US-government delay, with Sol topping several agentic-coding charts. Here's the confirmed launch plus the full pre-launch rumor trail that got us here.

WaveSpeed Blog

Read more →
MiMo-V2-Pro Xiaomi · API-only reasoning model 49 INTELLIGENCE INDEX 1M-token context #10 worldwide $1 / $3 per 1M tokens MI BITSMINDS.COM
Models

Xiaomi's MiMo-V2-Pro Cracks the Global Top 10 for AI Reasoning — at a Fraction of the Price

Xiaomi's reasoning model scores 49 on the Artificial Analysis Intelligence Index — #10 worldwide, between Kimi K2.5 and GLM-5 — while running its whole benchmark suite for just $348.

Artificial Analysis

Read more →
MODEL WATCH · OPENAI · LEAKS, UNCONFIRMED JUN 17 GPT-5.6 looks days away — if the leaks hold. Prediction markets put a late-June launch near 83% — but OpenAI has confirmed nothing. 83% LAUNCH ODDS Polymarket · June 22–28 window RUMORED SPECS Context window up to ~1.5M tokens Stronger agentic coding, a leaker claims Pricing rumored near a third of Fable 5 BITSMINDS.COM Source: Polymarket · Cryptopolitan · leaks
Models

GPT-5.6 Rumors Reach a Fever Pitch: Prediction Markets Bet on a Late-June Launch, Leaks Claim a 1.5M-Token Window

No OpenAI announcement exists yet, but developer leaks and prediction markets now point to a GPT-5.6 launch in late June — Polymarket prices the June 22–28 window near 83% — with rumored upgrades including a ~1.5M-token context and stronger agentic coding. Treat it all as unconfirmed.

Cryptopolitan

Read more →
ZHIPU / Z.AI · OPEN MODEL JUN 16 GLM-5.2 goes open under MIT. A 744B-parameter MoE built for long-horizon agentic coding. 744B TOTAL · 40B ACTIVE (MoE) 1M CONTEXT TOKENS 131K MAX OUTPUT TOKENS MIT OPEN-WEIGHTS LICENSE No benchmarks published at launch — performance is a vendor claim. glm-5.2[1m] · 28.5T training tokens · agentic coding BITSMINDS.COM Source: Z.ai
Models

Zhipu Releases GLM-5.2: a 744B-Parameter, 1M-Token Coding Model Under a Full MIT License

Zhipu (Z.ai) has launched GLM-5.2, a 744B-parameter Mixture-of-Experts model (40B active) with a 1M-token context window, built for agentic coding and released under a permissive MIT license. The catch: no benchmarks were published at launch, so performance claims remain vendor assertions for now.

Z.ai

Read more →
Moonshot AI Ships Kimi K2.7-Code, an Open-Weight Coding Model It Says Uses 30% Fewer Reasoning Tokens
Models

Moonshot AI Ships Kimi K2.7-Code, an Open-Weight Coding Model It Says Uses 30% Fewer Reasoning Tokens

The Beijing lab’s fifth K2 release in a year is a coding-first, 1-trillion-parameter open-weight model under a modified MIT license — but independent testers say the headline benchmarks are Moonshot’s own and don’t all hold up.

Crypto Briefing

Read more →
BITSMINDS.COM
Models

Fable 5 vs Opus 4.8, One Prompt Each: Four Builds, Dead Level at 2–2

We pasted identical prompts into Claude Opus 4.8 and Claude Fable 5 — an animated highway interchange, a one-file Space Invaders, a pure-SVG aquarium, and a 3D rocket launch — one shot each, no retries. Everything runs live inside the article. Four builds in, it is dead level at 2–2 — and the launch round flips the very pattern the first three revealed.

BitsMinds Lab

Read more →
BITSMINDS VERDICT Claude Fable 5 The most capable public model yet — for a premium 4.6 OUT OF 5 ★★★★½ SWE-Bench Pro 80.3 Humanity's Last Exam 59.0 Blueprint-Bench 2 38.6 Fable 5 headline benchmarks — higher is better BITSMINDS.COM
Models

Review: Claude Fable 5 Is the Most Capable Public Model Yet — for a Premium

Anthropic's Claude Fable 5 is the most capable model the public can use today — topping SWE-Bench Pro and excelling at vision and long tasks. It is also the priciest major model and ships with hard safety guardrails. Our first-look verdict.

BitsMinds Reviews

Read more →
CLAUDE FABLE 5MYTHOS-CLASS, NOW PUBLICSWE-BENCH PRO80.3agentic coding · vs 69.2 OpusPRICE / 1M TOKENS$10 · $50input / outputCYBER ATTACK SUCCESS5.4%guardrails on · lower is saferAVAILABLE NOWAPI · CopilotPro · Max · Team · EnterpriseBITSMINDS.COMSource: Anthropic
Models

Anthropic Launches Claude Fable 5: Its Most Powerful Public Model Yet, With the Cyber Edges Sanded Off

Anthropic released Claude Fable 5 — the first publicly available model in its powerful Mythos class — scoring 80.3 on SWE-Bench Pro (vs 69.2 for Opus 4.8), priced at $10/$50 per million tokens, with classifier guardrails that block cyber, bio, and distillation requests and fall back to Opus 4.8.

Anthropic

Read more →
RUMOREDCLAUDEFABLE 5PUBLIC MYTHOS?claude-fable-5 · rumoredUNVERIFIEDMODEL CARD · LEAKEDclaude-fable-5ArchitectureClaude 5Long contextextendedMulti-turnimprovedCyber toolingExploit devUNLOCKINGwith guardrailsBITSMINDS.COMReporting: Alex Heath
Models

Claude Fable 5: The Internet Is Sure It Ships Today. Anthropic Hasn't Said a Word.

Prediction markets put it near 94%, a tech reporter says it lands June 9, and internal checkpoints named claude-fable-5 have surfaced — but Anthropic has confirmed nothing. Here is what the noise around Claude Fable 5, the rumored public release of its Mythos model, actually amounts to.

Reporting: Alex Heath; Polymarket

Read more →
Xiaomi Pushes a 1-Trillion-Parameter Model Past 1,000 Tokens a Second — on Eight Off-the-Shelf GPUs
Models

Xiaomi Pushes a 1-Trillion-Parameter Model Past 1,000 Tokens a Second — on Eight Off-the-Shelf GPUs

MiMo-V2.5-Pro-UltraSpeed broke the 1,000 tokens-per-second barrier on a one-trillion-parameter model using a single eight-GPU commodity node — roughly 15x faster than GPT-5.5 or Claude Opus, and without a line of custom silicon.

MarkTechPost

Read more →
OPEN MODELS · CODING · 1M CONTEXTMINIMAX M3 · JUNE 1MiniMax M3Open-weight · native multimodal · MSA attention~1/20 compute at 1M ctx · 9× faster prefill · 15× decodeBENCHMARK SCORECARD · MINIMAX-REPORTEDSWE-Bench Pro59%Terminal-Bench 2.166%SWE-fficiency34.8%BrowseComp83.5BROWSECOMPBeats Opus 4.783.5 vs 79.3Weights + technical report promised within 10 days · scores not yet independently verifiedBITSMINDS.COMSource: The Decoder · MiniMax
Models

MiniMax M3 Lands as an Open-Weight, Million-Token Coding Model That Claims to Edge Out GPT-5.5

The Chinese lab says its new open-weight model pairs frontier coding, a 1M-token context window and native multimodality on a sparse-attention architecture that cuts long-context compute 20x — but the weights and technical report are still days away, so every number is company-reported.

The Decoder

Read more →
REPORTEDLY LEAKEDCLAUDEOCEANUSANTHROPIC · MYTHOS LINEclaude-oceanus-v1-p · surfaced June 3 · unconfirmedBITSMINDS.COMSource: Cyber Security News
Models

Anthropic's Unreleased 'Claude Oceanus' Reportedly Leaked Through a Chinese Proxy Hours After Reaching Red Teamers

A model identifier, claude-oceanus-v1-p, surfaced in Anthropic's console on June 3 and went to vetted red-team testers — then, within hours, an unknown actor allegedly resold access through a Chinese API proxy. None of it is officially confirmed, but it offers an early look at the next model in the Mythos line.

Cyber Security News

Read more →
Aaheadline · in-image textSUBJECTLOGO#hex paletteIMAGE GENERATION · JUNE 3, 2026LAYOUT-NATIVEReve 2.0 + Ideogram 4Layout control is the new promptReve 2.0 — proprietary, 4K, layout editingIdeogram 4 — open-weight, 9.3B, in-image textBoth chase OpenAI GPT Image 2 and Google GeminiBITSMINDS.COMSource: Reve · Ideogram · Latent Space
Models

Reve 2.0 and Ideogram 4 Land the Same Day, Both Betting Image AI's Next Move Is Layout Control

On June 3, two image-AI startups shipped models built on the same wager: that the future of generation is precise layout control, not ever-longer prose prompts. Reve's proprietary 2.0 chases 4K and "images you can touch"; Ideogram's open-weight 4 nails in-image text and bounding-box placement. Both are gunning for OpenAI's GPT Image 2 and Google's Gemini.

Latent Space

Read more →
OPENAI · LIFE SCIENCESJUNE 2026 UPDATEGPT-RosalindRebuilt on GPT-5.5, tuned for biologyDrug discovery · Genomics · Medicinal chemistryOutperforms GPT-5.5 on the new LifeSciBench evalBITSMINDS.COMSource: OpenAI
Models

OpenAI Rebuilds GPT-Rosalind on GPT-5.5 and Widens Access — Its Science Model Now Beats the General One on Biology

OpenAI's life-sciences model just got a major upgrade. Rebuilt on GPT-5.5, it now outscores the general model on a new benchmark called LifeSciBench across drug discovery, genomics, and medicinal chemistry — and OpenAI is widening access beyond launch partners like Moderna, Amgen, and Novo Nordisk.

OpenAI

Read more →
BUILD 2026 · MICROSOFT TAKES AIM AT CLAUDE LOCKED ON MICROSOFT AI MAI-Thinking-1 FIRST IN-HOUSE REASONING MODEL 35B active · ~1T MoE · 256K ctx trained with zero distillation Claude ENTERPRISE CODING DEFAULT SWE-Bench Pro ≈ 53% — level with Opus 4.6 BITSMINDS.COM Microsoft's claims, pending independent benchmarks
Models

Microsoft Aims MAI-Thinking-1 Straight at Claude: a 35B Reasoning Model It Says Beats Sonnet 4.6 and Matches Opus 4.6 on Code

At Build 2026, Microsoft’s AI Superintelligence Team unveiled MAI-Thinking-1, its first in-house reasoning model — a sparse Mixture-of-Experts design with 35B active parameters (~1T total) and a 256K context window. Microsoft says human raters prefer it to Claude Sonnet 4.6 and that it matches Opus 4.6 on the SWE-Bench Pro coding benchmark, and it pointedly trained the model from scratch with zero distillation on commercially licensed data — no OpenAI involved.

Microsoft AI / TechTimes

Read more →
OPEN SOURCE · MIXTURE-OF-EXPERTS · APACHE 2.0JETBRAINS MELLUM2 · JUN 2Mellum212B total parameters · ~2.5B active per tokenMoE · 131K context · 2x faster inference6 variants · Base · Instruct · Thinking (RLVR)THE FOCAL-MODEL THESISFast, specialized parts orchestrated by frontier modelsROUTER · 8 OF 64 EXPERTS FIREOnly ~21% of parameters fire per tokenBITSMINDS.COMSource: JetBrains AI Blog · Hugging Face
Models

JetBrains Open-Sources Mellum2, a 12B Mixture-of-Experts Model Built to Be a Fast "Focal" Part, Not a Frontier Rival

JetBrains released Mellum2 under Apache 2.0 on June 2 — a 12B MoE model that activates just 2.5B parameters per token for 2x-faster inference, pitched as a fast specialized component for multi-model AI pipelines.

JetBrains AI Blog

Read more →
Microsoft Will Unveil Its Own GitHub Copilot Coding Model at Build — a Direct Shot at Claude Code
Models

Microsoft Will Unveil Its Own GitHub Copilot Coding Model at Build — a Direct Shot at Claude Code

According to The Information, Microsoft plans to debut a homegrown coding model for GitHub Copilot at its Build conference on June 2–3, part of a wider suite of in-house MAI models led by Mustafa Suleyman. The push to reduce reliance on OpenAI, Anthropic and Google comes as Copilot has lost ground to Anthropic’s Claude Code among developers.

The Information / TestingCatalog

Read more →
FRONTIER MODEL SHOWDOWN · WHO WINS? Three labs. Three strongest models. One fight. ANTHROPIC Claude Opus 4.8 AUTONOMY OPENAI GPT-5.5 “Spud” AGENTS GOOGLE Gemini 3.1 Ultra REASONING VS VS BITSMINDS.COM BitsMinds original analysis
Models

Claude Opus 4.8 vs GPT-5.5 vs Gemini 3.1 Ultra: Benchmarks

A BitsMinds analysis. We put each lab’s strongest model through the benchmark portfolio that actually separates frontier systems in 2026 — GPQA Diamond, ARC-AGI-2, AIME, SWE-bench, BFCL and more — then weigh price, context and speed. The short version: Gemini 3.1 Ultra is the sharpest reasoner and the best value, Claude Opus 4.8 owns real coding and agentic work, and GPT-5.5 is the strong all-rounder that everyone already has. Here is the full scorecard.

BitsMinds Analysis

Read more →
MODEL RELEASE · AGENTIC AI · ANTHROPICCLAUDE OPUS 4.8 · MAY 28Opus 4.8Anthropic’s most honest model yetSame price as 4.7 · $5 / $25 per M tokensFast mode · 2.5× speed · 3× cheaperEffort dial · plan-and-verify subagentsDYNAMIC WORKFLOW · PARALLEL SUBAGENTSPLANVERIFY ✓running 100s in parallel · one sessionBITSMINDS.COMSource: Anthropic · TechCrunch · Axios
Models

Anthropic Ships Claude Opus 4.8: Dynamic Workflows Spawn Hundreds of Parallel Subagents, and It’s the “Most Honest” Claude Yet

Anthropic released Claude Opus 4.8 on May 28, 41 days after 4.7 — adding Dynamic Workflows that plan and run hundreds of parallel subagents per session, a user-facing effort dial, a 2.5× faster fast mode at 3× lower cost, and its “most honest” self-review behavior yet, at 4.7’s price.

Anthropic / TechCrunch

Read more →
ANTHROPIC ® INTRODUCING OPUS 4.8 LEAK FORBIDDEN-STRINGS FILE · NPM SOURCE MAP · MARCH 31, 2026 BITSMINDS Source: WaveSpeed AI · decodethefuture.org
Models

Anthropic's Leaked Forbidden-Strings File Names Opus 4.8, Sonnet 4.8 and a New Tier Above Opus Called Capybara

Inside the 512K-line Claude Code source map that npm shipped on March 31 was a feature literally called Undercover Mode — and inside Undercover Mode was a hard-coded list of names it was supposed to scrub from public builds. That list spells out Opus 4.8, Sonnet 4.8, a never-released Sonnet 4.7 and animal codenames Capybara and Tengu — with Capybara described internally as a new tier sitting above Opus.

WaveSpeed AI / ChaoBro / AI Nexus Daily

Read more →
Alibaba's Qwen3.7-Max Lands With a 1M-Token Context, AA Index 56.6, and a 35-Hour Agent Run
Models

Alibaba's Qwen3.7-Max Lands With a 1M-Token Context, AA Index 56.6, and a 35-Hour Agent Run

Qwen3.7-Max-Preview ships with a 1M-token context, extended-thinking mode, and benchmark gains across CritPt, Humanity's Last Exam, and Terminal-Bench Hard — clearing Gemini 3.5 Flash on the AA Intelligence Index.

Qwen (Alibaba Cloud)

Read more →
Google's Gemini 'Omni' Video Model Leaks Days Before I/O 2026, Pointing to a Unified Multimodal System
Models

Google's Gemini 'Omni' Video Model Leaks Days Before I/O 2026, Pointing to a Unified Multimodal System

Test UI strings inside Gemini reveal a new Google video model that lets users remix and edit clips directly in chat, with a likely debut at I/O 2026 on May 19-20.

9to5Google

Read more →
Mira Murati's Thinking Machines Unveils TML-Interaction-Small, a Full-Duplex Voice Model That Beats GPT and Gemini
Models

Mira Murati's Thinking Machines Unveils TML-Interaction-Small, a Full-Duplex Voice Model That Beats GPT and Gemini

Mira Murati's Thinking Machines Lab released a research preview of TML-Interaction-Small, a 276B mixture-of-experts model that responds in under 0.4 seconds and listens while it speaks — outpacing GPT-realtime-2 and Gemini 3.1 Flash Live on the FD-bench latency test.

TechCrunch

Read more →
OpenAI Ships GPT-Realtime-2, Translate, and Whisper, Bringing GPT-5 Reasoning Into Voice Apps
Models

OpenAI Ships GPT-Realtime-2, Translate, and Whisper, Bringing GPT-5 Reasoning Into Voice Apps

OpenAI rolled out three new voice models on May 7 — a reasoning agent with a 128K context window, a 70-language live translator at $0.034 a minute, and a streaming Whisper that transcribes as you speak.

OpenAI

Read more →
Google Ships Gemini 3.1 Flash-Lite to GA at $0.25 per Million Input Tokens
Models

Google Ships Gemini 3.1 Flash-Lite to GA at $0.25 per Million Input Tokens

Gemini 3.1 Flash-Lite is now generally available on Vertex AI and Gemini Enterprise, pitched as the cheapest, fastest member of the Gemini 3 family for high-volume agentic, classification, and tool-use workloads.

Google Cloud Blog

Read more →
Mistral Medium 3.5 Lands With Cloud Coding Agents and 77.6% on SWE-Bench
Models

Mistral Medium 3.5 Lands With Cloud Coding Agents and 77.6% on SWE-Bench

Mistral fuses chat, reasoning, and coding into a single 128B dense model and pairs it with Vibe — async cloud-based coding agents that hand back work as pull requests instead of terminal output.

Mistral AI

Read more →
DeepSeek V4 Preview Closes Gap With Frontier Models at a Fraction of the Price
Models

DeepSeek V4 Preview Closes Gap With Frontier Models at a Fraction of the Price

Chinese lab DeepSeek released two open-weight V4 preview models with 1M-token context windows, matching closed-source rivals on reasoning while undercutting them on price.

TechCrunch

Read more →
NVIDIA Unveils Nemotron 3 Nano Omni: Open Multimodal Model with 9x Throughput
Models

NVIDIA Unveils Nemotron 3 Nano Omni: Open Multimodal Model with 9x Throughput

NVIDIA's new 30B-A3B mixture-of-experts model unifies vision, audio, and language into a single open system, delivering up to nine times the throughput of comparable open omni models for AI agents.

NVIDIA Blog

Read more →
OpenAI Launches GPT-5.5: Smarter, Faster, and Built for Agentic Work
Models

OpenAI Launches GPT-5.5: Smarter, Faster, and Built for Agentic Work

OpenAI's GPT-5.5, codenamed "Spud," brings 82.7% Terminal-Bench 2.0 performance and autonomous multi-step task execution — a step toward OpenAI's unified AI super app.

TechCrunch

Read more →
DeepSeek Releases V4: Open-Source Frontier Model at a Fraction of the Price
Models

DeepSeek Releases V4: Open-Source Frontier Model at a Fraction of the Price

DeepSeek's V4-Pro and V4-Flash arrive with 1.6 trillion parameters, a 1M-token context window, and Apache 2.0 licensing — matching near-frontier performance at 21x less cost than Claude Opus 4.7.

Simon Willison

Read more →
Google Launches Gemini 3.1 Pro: Tops ARC-AGI-2 Benchmark with 2M Token Context and Agentic Coding
Models

Google Launches Gemini 3.1 Pro: Tops ARC-AGI-2 Benchmark with 2M Token Context and Agentic Coding

Google's Gemini 3.1 Pro hits a verified 77.1% on the ARC-AGI-2 reasoning benchmark — more than double its predecessor — and arrives with a 2-million-token context window, multi-step agentic coding, and support across Gemini API, Vertex AI, and NotebookLM.

Google

Read more →
Meta Launches Muse Spark: Multimodal AI with Parallel Agents, Health Insights, and Shopping Intelligence
Models

Meta Launches Muse Spark: Multimodal AI with Parallel Agents, Health Insights, and Shopping Intelligence

Meta Superintelligence Labs unveils Muse Spark, a fast multimodal model that deploys parallel AI subagents, analyzes medical images, generates functional websites, and surfaces social-powered shopping recommendations.

Meta

Read more →
xAI Launches Grok 4.3 Beta: Native Video Input, Document Generation, and Desktop Automation
Models

xAI Launches Grok 4.3 Beta: Native Video Input, Document Generation, and Desktop Automation

xAI has launched Grok 4.3 in beta with native video processing, PDF and spreadsheet generation, and tighter Grok Computer integration — behind a $300/month SuperGrok Heavy paywall ahead of a broader May rollout.

TechSifted

Read more →
Moonshot AI Releases Kimi K2.6: Open-Source Giant with 1T Parameters and 300-Agent Swarms
Models

Moonshot AI Releases Kimi K2.6: Open-Source Giant with 1T Parameters and 300-Agent Swarms

Moonshot AI has released Kimi K2.6, a 1-trillion-parameter open-source model capable of coordinating 300 parallel sub-agents across 4,000 steps — available on Hugging Face under a Modified MIT License.

MarkTechPost

Read more →
Google Gemma 4: Open Models That Outperform Systems 20x Their Size
Models

Google Gemma 4: Open Models That Outperform Systems 20x Their Size

Google DeepMind released Gemma 4 on April 2, a family of four open-weight multimodal models under Apache 2.0 that bring frontier-level reasoning and agentic capabilities to phones, edge devices, and developer environments.

Google DeepMind

Read more →
Anthropic's Claude Mythos 5: The First 10-Trillion-Parameter AI Model
Models

Anthropic's Claude Mythos 5: The First 10-Trillion-Parameter AI Model

Anthropic has revealed Claude Mythos 5, the first publicly acknowledged 10-trillion-parameter AI model, engineered for cybersecurity, advanced coding, and deep academic reasoning — but not yet available to the public.

Anthropic / AI Analytics Diaries

Read more →
Microsoft Launches Three MAI Models to Rival OpenAI and Google in Speech, Voice, and Image
Models

Microsoft Launches Three MAI Models to Rival OpenAI and Google in Speech, Voice, and Image

Microsoft unveils MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 through its Foundry platform, marking the company's boldest push yet to build AI infrastructure independent of its OpenAI partnership.

TechCrunch

Read more →
Anthropic Releases Claude Opus 4.7 With Major Coding and Vision Upgrades
Models

Anthropic Releases Claude Opus 4.7 With Major Coding and Vision Upgrades

Claude Opus 4.7 is now generally available, bringing a 13% coding benchmark improvement, 3x higher image resolution support, and a new xhigh effort setting -- at the same price as Opus 4.6.

Anthropic

Read more →
NVIDIA Launches Ising: Open AI Models to Supercharge Quantum Computing
Models

NVIDIA Launches Ising: Open AI Models to Supercharge Quantum Computing

NVIDIA unveils Ising, the world's first family of open-source AI models built to accelerate fault-tolerant quantum computing through AI-powered calibration and error correction.

NVIDIA Newsroom

Read more →
OpenAI Launches GPT-6: 40% Capability Jump, 2M Token Context, and Super-App Integration
Models

OpenAI Launches GPT-6: 40% Capability Jump, 2M Token Context, and Super-App Integration

OpenAI releases GPT-6 with a 40% performance leap, a 2 million token context window, and a unified super-app merging ChatGPT, Codex, and the Atlas browser into a single agent experience.

OpenAI

Read more →
Google Gemini 3.1 Ultra Breaks Reasoning Records with 94.3% on GPQA Diamond
Models

Google Gemini 3.1 Ultra Breaks Reasoning Records with 94.3% on GPQA Diamond

Google's flagship Gemini 3.1 Ultra achieves a verified 94.3% on GPQA Diamond and 77.1% on ARC-AGI-2, setting new state-of-the-art benchmarks for complex scientific and logical reasoning.

Google DeepMind

Read more →
OpenAI Releases GPT-5.4 Mini and Nano: Purpose-Built for the Subagent Era
Models

OpenAI Releases GPT-5.4 Mini and Nano: Purpose-Built for the Subagent Era

OpenAI's new small models are designed to function as parallel subagents inside larger AI workflows — fast, cheap, and capable enough to handle the bulk of agentic work.

OpenAI

Read more →
NVIDIA Unveils Nemotron 3: Nano, Super, and Ultra Open Models for Agentic AI
Models

NVIDIA Unveils Nemotron 3: Nano, Super, and Ultra Open Models for Agentic AI

NVIDIA debuted the Nemotron 3 family at GTC 2026, releasing Nano immediately and previewing Super (49B) and Ultra (253B) variants, along with three trillion tokens of pre-training data.

NVIDIA Newsroom

Read more →
Meta Releases Llama 5: Open-Source 600B Model Claims Frontier Dominance
Models

Meta Releases Llama 5: Open-Source 600B Model Claims Frontier Dominance

Meta CEO Mark Zuckerberg unveiled Llama 5, a 600-billion-parameter open-weight model trained on 500,000 Blackwell GPUs that Meta says outperforms GPT-5 and Gemini on key benchmarks.

FinancialContent / Meta AI

Read more →
Google Releases Gemma 4: Open Models That Outperform Rivals 20x Their Size
Models

Google Releases Gemma 4: Open Models That Outperform Rivals 20x Their Size

Google launches Gemma 4, a family of open models ranging from 2B to 31B parameters that deliver unprecedented intelligence-per-parameter, outperforming models 20 times their size on key benchmarks.

Google Blog

Read more →
Meta Unveils Muse Spark: First AI Model from $14B Superintelligence Labs
Models

Meta Unveils Muse Spark: First AI Model from $14B Superintelligence Labs

Meta debuts Muse Spark, a proprietary multimodal model built by Alexandr Wang's elite superintelligence team, marking a dramatic shift away from the company's open-source Llama strategy.

TechCrunch

Read more →
Anthropic's Claude Mythos 5: A 10-Trillion-Parameter Step Change in AI Capability
Models

Anthropic's Claude Mythos 5: A 10-Trillion-Parameter Step Change in AI Capability

Anthropic has begun limited enterprise testing of Claude Mythos 5, a 10-trillion-parameter model that the company calls a 'step change' in AI performance, excelling at cybersecurity, complex coding, and academic reasoning.

Fortune

Read more →
Google Releases Gemini 3.1 Pro with Massive 2 Million Token Context Window
Models

Google Releases Gemini 3.1 Pro with Massive 2 Million Token Context Window

Google's Gemini 3.1 Pro sets a new industry benchmark with a 2 million token context window and 77.1% on ARC-AGI-2, enabling AI to reason over entire codebases and lengthy documents in one pass.

MarkTechPost

Read more →
Google Gemini 2.0: The Model That Unifies Everything
Models

Google Gemini 2.0: The Model That Unifies Everything

Google unveils Gemini 2.0 with advanced multimodal capabilities and deep integration across all Google services.

Google AI Blog

Read more →
Anthropic Releases Claude 4: Deeper AI Understanding
Models

Anthropic Releases Claude 4: Deeper AI Understanding

Anthropic launches Claude 4 with dramatic improvements in text comprehension, coding, and mathematical reasoning.

Anthropic

Read more →
OpenAI Launches GPT-5: The Next Generation of AI
Models

OpenAI Launches GPT-5: The Next Generation of AI

OpenAI has announced GPT-5, its most advanced model yet, with enhanced reasoning capabilities and multimodal features.

OpenAI Blog

Read more →