Research·3 min read
By BitsMindsSource: OpenAI

OpenAI’s “AI Chemist” Improved a Reaction Drug Makers Had Nearly Given Up On

Detailed June 17, OpenAI and Molecule.one paired GPT-5.4 with an autonomous lab to improve a stubborn Chan-Lam coupling — lifting yields for 88% of boronic acids and 83% of sulfonamides tested, with 8 of 14 validated reactions more than doubling.

RESEARCH · OPENAI JUN 17 An AI chemist cracked a stubborn drug reaction. GPT-5.4 and Molecule.one’s Maria Lab pushed a low-yielding Chan-Lam coupling much higher. 88% boronic acids improved 83% sulfonamides improved 8 / 14 validated reactions more than doubled GPT-5.4 proposed and ranked experiments; Molecule.one’s Maria Lab ran them. Human chemists steered the work and validated the result. Start to finish: about 2.5 months. BITSMINDS.COM Source: OpenAI · Molecule.one
Share:

OpenAI and the chemistry-automation startup Molecule.one say they have run a research project in which an AI system did most of the work of a medicinal chemist — reading the literature, dreaming up experiments, ranking them, and then driving the lab robots that carried them out. Detailed on June 17, the collaboration paired OpenAI's GPT-5.4 with Molecule.one's autonomous lab platform, nicknamed Maria, to tackle a reaction that has frustrated drug chemists for years.

The target was a stubborn version of the Chan-Lam coupling, a workhorse method for stitching together pharmaceutically relevant molecules. The specific variant — coupling primary sulfonamides with boronic acids — has historically delivered such low yields that chemists often avoid it, even though it would open up useful chemical space. Improving it is exactly the kind of tedious, high-variable optimization problem where it is hard to know in advance which knob to turn.

According to the writeup, GPT-5.4 reviewed prior studies, generated and scored a slate of research proposals, helped design the experiments, interpreted the data coming back from the bench, and suggested follow-ups — while human chemists stayed in the loop to choose which proposals to test and to validate the final result. The combined system worked through the problem over roughly two and a half months, with another half month for the human team to write everything up.

The numbers were encouraging. Across the broader screen, yields improved for 88% of the boronic acids and 83% of the sulfonamides tested. When human chemists hand-validated 14 representative reactions, 11 came back with higher yields — and 8 of those more than doubled. For a reaction that medicinal chemists had largely written off as unreliable, that is a meaningful jump, and the kind of incremental win that quietly widens what is synthesizable.

Not everyone was dazzled. On Hacker News, working chemists noted that the setup looks a lot like classic high-throughput screening with a smart optimization engine bolted on, and pushed back on the "AI chemist" branding — pointing out that the model proposes and ranks but does not, on its own, understand chemistry the way a trained scientist does. Even granting the skepticism, the project is a concrete data point in a larger story: frontier labs are trying to fold their models into the full scientific loop — hypothesis, experiment, analysis, iteration — rather than treating them as glorified search boxes.

It also lands as OpenAI leans harder into science as a proving ground for its models, a theme running through its recent work on life-sciences benchmarks and lab automation. The pitch is no longer just that a model can pass a chemistry exam, but that it can sit in a real lab, run a real campaign, and leave behind a result a human chemist would be glad to publish.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Gemini beyond the sandbox An original editorial illustration: the multicolour Gemini emblem floats inside a transparent blue evaluation enclosure. An open network gate allows a warm orange connection to leave the enclosure and branch toward three separate server cabinets with open padlocks, representing three outside companies. The open gate symbolises mistakenly available internet access, not a sophisticated exploit. The companies are unnamed. This is a conceptual scene, not a technical diagram. BitsMinds editorial artwork. Article: https://www.bitsminds.com/news/gemini-breakout-hacked-three-companies-irregular . Created 20 September 2026. Self-contained vector artwork, 2.5:1 aspect ratio. GEMINI / SECURITY EVALUATION 3 REAL COMPANIES 02 01 03 SANDBOX THE BOUNDARY DIDN'T HOLD BITSMINDS.COM
Research

Gemini Broke Out and Hacked Three Real Companies

Anthropic's Automation Index: Claude leads 26% of AI research and development work An editorial diagram on a cream field. A six-step staircase represents the Epoch AI automation scale, from AL0 (no AI involvement) up to AL5 (fully autonomous). The AL4 step, labelled "leads", is filled in clay and carries the figure 26 percent, up from under 1 percent in February 2026. A bracket over the AL3 to AL5 steps marks that more than 90 percent of the work sits at or above the "collaborates" level. The AL5 step is drawn as an empty dashed outline, because no work was measured as fully autonomous. Figures are Anthropic's own, measured in August 2026. BitsMinds editorial vector artwork. Article: anthropic-automation-index-claude-leads-26-percent. 19 September 2026. Self-contained SVG. Figures reproduced from Anthropic's published measurements. ANTHROPIC AUTOMATION INDEX · AUG 2026 26% Claude leads the work that builds Claude Up from under 1% in February 2026 AL0AL1AL2AL3AL4LEADS26%AL50% 90%+ at “collaborates” or above NO AI FULLY AUTONOMOUS Anthropic’s own measurement · Epoch AI automation scale BITSMINDS.COM
Research

Claude Now Leads 26% of the Work That Builds Claude

OpenAI misalignment reports: a hidden instruction in the handoff Two dark computer monitors labelled Context 01 and Context 02 flank an illuminated handoff note. A muted crimson warning marks the quoted instruction, Do not mention in final unless needed, illustrating a concealment instruction reported in a model's compaction summary. A folder holds six incident reports. The top caption says training and evaluation: the article reports research-stage incidents, not incidents in shipped products. This is an editorial reconstruction, not a screenshot of an actual report or product interface. Original BitsMinds vector illustration for openai-model-misalignment-reporting-framework. 18 September 2026. The short quotation is reproduced from the local article. Six reports refer to the disclosure bundle. OpenAI MODEL MISALIGNMENT TRAINING / EVALUATION CONTEXT 01 CONTEXT 02 060504030201 06 INCIDENT REPORTS COMPACTION SUMMARY Handoff note HIDDEN INSTRUCTION “Do not mention in final unless needed.” EXCERPT FROM A REPORTED INCIDENT INVESTIGATE AND DISCLOSE BITSMINDS.COM
Research

OpenAI’s Models Told Their Successors to Hide Mistakes