Research·3 min read
By BitsMindsSource: arXiv

Frontier Models Recover a Paper’s Idea 3–15% of the Time

A new blind benchmark hands models only the references a researcher cited before publishing and asks them to reinvent the paper. Across 643 papers in six fields, the best single model managed 13.3%.

RECONSTRUCTION BENCHMARK 3–15% SINGLE MODEL 36% PANEL OF FOUR BITSMINDS.COM
Share:

A new benchmark called Reconstruction asks a blunt question: if you hand a model only the list of papers a researcher had cited before publishing, can it reinvent the idea they published? Across 643 papers in six scientific fields, frontier models did it between 3% and 15% of the time. The paper, posted to arXiv in mid-August by a nine-author team including Shaolong Chen, Rahul Thapa, Qingqing Mao and Ritankar Das, is a preprint and has not been peer reviewed.

The design is built around leakage, which is the failure mode that has quietly wrecked most attempts to measure machine ideation. Reconstruction withholds the seed paper itself plus all contemporaneous and future literature, strips references down to anonymous IDs, freezes each paper's bibliography, and enforces a temporal citation cutoff so a model cannot recognize the target from a later citation of it. The model proposes hypotheses; an independent large language model judge decides whether any of them match the real paper's ideas. The domains are machine learning, astronomy, chemistry, materials, medicine and physics.

Seven models were tested in July 2026, and they land in a tight, low band. Claude Opus 4.8 led at 13.3% average match rate, followed by GPT-5.6 Sol Pro at 12.8%, Kimi K3 at 10.0%, GLM 5.2 at 9.3%, Gemini 3.1 Pro Preview at 8.9%, DeepSeek-V4-Pro at 6.3% and Qwen3.7-Max at 5.9%. The top-to-bottom spread is 2.3x, which sounds like a real gap until you notice that every model in the list fails the task the overwhelming majority of the time.

Ensembling helps considerably more than picking a better model. A reference-only pipeline that runs cross-model review and then Swiss-tournament selection over aligned hypothesis slots — with no external web search — reached 36.0% overall using the top four models, roughly 2.4x the best single-model score. Per domain the multi-agent pipeline ranged from 22.9% on machine learning to 41.6% on medicine, with chemistry at 38.4% and materials at 40.1%.

Machine learning being the worst domain is the most awkward number in the table. It is the field these models have almost certainly read the most about, and it is where they recovered the fewest ideas. The error bars are wide enough (roughly ±6 to ±11 points depending on domain) that the per-domain ranking should not be read too hard, but the ML result sits well below the rest, and a benchmark whose anti-leakage machinery works would be expected to produce exactly this shape: familiarity with a literature stops substituting for the ability to extend it.

Two caveats deserve equal weight to the headline. The judge is itself a language model, so "match" is a model's opinion of a match, and the whole protocol measures retrodiction rather than discovery — the answer exists, it is simply outside the model's context. That makes it an easier task than genuine novelty, which is what gives the low numbers their bite. It also puts a specific ceiling under the automated-science pitch: Jeff Dean left Google to automate the discovery loop, and models have pushed real results when a sharply framed problem is handed to them, as when Claude improved a Riemann zeta bound. Framing the problem in the first place is the part Reconstruction says is still mostly out of reach, and it lines up with the Nature study that found human scientists still ahead of top agents on complex research work.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Mythos cracks Rejetto HFS's random signing key A large ivory die on a cream field, the official Anthropic mark inlaid in clay on its front face. Its top face has split along a crack, and a brass key is rising out of it: the session signing key recovered from Math.random. Faint leaked random numbers drift in from the left; faint cookie fragments sit on the right. 0.73418 0.11902 0.58361 0.92047 0.30775 admin=1 sig:9f3c keygrip xs128+ BITSMINDS.COM
Research

A Bug Claude Mythos Found Was Exploited Within a Day

Meta Muse Spark's six math papers A fan of research manuscripts on a deep blue field. Five sheets behind carry gold check marks for the five open problems answered; the front sheet carries the official Meta mark, inlaid. Faint mathematical symbols float on either side. ∫ ∑ ψ |G| = 384 ∂ₜu λ ≥ 0 ℚₚ ≠ MUSE SPARK · THINKING 6 PAPERS · 5 OPEN PROBLEMS BITSMINDS.COM
Research

Meta Says Muse Spark Helped Crack Five Open Math Problems

arXiv's manuscript meter: two submissions per month On a burgundy desk, a tall, uneven stack of manuscripts and loose research sheets meets an imagined cream-and-burgundy submission meter bearing the official arXiv wordmark. A large physical counter reads 2 / MONTH, PER SUBMITTER. On the other side, a shallow brass-and-ivory tray holds exactly two illustrated manuscript sheets. Paper diagrams, fold corners, a retaining arm and a slim desk pen make the scene tactile. The meter is an editorial metaphor for the limit of two submissions per calendar month by the person uploading them, across all subject areas. The sheets are not represented as peer-reviewed or approved; rejected submissions also consume the monthly quota. The separate limit of three active submissions is not depicted, and the large stack is symbolic of submission volume rather than one person's active moderation queue. Original vector illustration for BitsMinds, 2 October 2026. arxiv-caps-submissions-two-per-month-ai-papers. Self-contained SVG with the project arXiv logo; no raster art or external resources. 2 / MONTH PER SUBMITTER SUBMISSIONS BITSMINDS.COM
Research

arXiv Caps Submissions at Two a Month as AI Papers Pile Up