DeepMind Alumni’s 27B Agent Beats Opus 4.8 on Replication
London lab Inherent says Faraday, its research agent running on a 27-billion-parameter Qwen model, reproduced published scientific findings more reliably than Claude Opus 4.8 and GPT-5.5 — and that how it was trained matters more than the win.
Inherent, a London lab founded by four Google DeepMind alumni, says its research agent Faraday beat Claude Opus 4.8 and GPT-5.5 at a deceptively hard task: independently reproducing the findings of published scientific papers, without being handed the answers first. Faraday runs on Qwen 3.6, a 27-billion-parameter open model — a fraction of the size of either system it outscored.
Replication is a good test precisely because it is unglamorous. A paper tells you what was found, not the dozen judgement calls that produced it: which control to run, which preprocessing step matters, when a null result means the method is wrong versus the implementation. “Many PhD students actually start by doing this,” chief scientist Edward Hughes told TechCrunch, and that is roughly the level being probed — can an agent recover a result the way a first-year would, by rebuilding the experiment rather than by pattern-matching the abstract.
Inherent’s framing of the result is unusually modest about the result itself. “What was most interesting to us about this was not so much the result of beating those frontier agents — which of course we liked — but was actually the way we went about building this,” Hughes said. The method is reinforcement learning aimed at what the team calls research taste: rather than teaching the model the scientific method as a set of rules, they reward good experimental outcomes and let an instinct for which experiments are worth running emerge from that. “We’re always guided by that north star of building an AI scientist agent and imbuing our agents with taste,” Hughes said.
The behaviour they are aiming for is a colleague rather than a tool — an agent that comes back with “I got curious about this, and I went off and I did these experiments. What do you think of these results?” That is a meaningfully different product shape from a coding agent that waits for instructions, and it is the part of the pitch that a size-versus-score comparison does not capture.
What is missing is the number. Inherent has not published a paper, named the benchmark, or released per-model scores, so “outperformed Opus 4.8 and GPT-5.5” is currently a vendor claim about an unspecified evaluation. That matters more than usual for replication work, where task selection does most of the heavy lifting: a set of papers chosen for clean, computational, small-data reproductions flatters a specialised agent, and a set drawn from wet-lab or large-compute work would flatter nobody. Until the harness is public, the honest reading is that a 27B model with task-specific RL can compete with much larger general agents on a narrow slice — which is a real finding, and a different one from a frontier-model ranking.
It also lands next to a much less cheerful datapoint. The Reconstruction benchmark, published earlier this month, found frontier models could recover a paper’s core idea from its bibliography alone just 3 to 15 percent of the time. The two tasks are not the same — one is rebuilding a stated finding, the other is inventing the finding from context — but read together they sketch where automated science actually is: agents are getting usable at re-running known work, and remain nowhere near originating it.
Inherent emerged from stealth in May 2026 with a $50 million seed round. Hughes is joined by cofounders Louis Kirsch, Kaloyan Aleksiev and Tantum Collins, all with DeepMind backgrounds; the company is based in King’s Cross and plans to reach 20 to 25 staff by the end of the year, with world models as the next line of work. For a team that size, picking replication as the proving ground is a shrewd choice — it is a task where a small model with good habits can plausibly beat a big model with none, and one where the customers, academic labs drowning in unverified literature, already know they have a problem.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.