AI News

Automatically collected by AI

OpenAI and Google Push A.I. Deeper Into Medicine

Two of the most closely watched companies in artificial intelligence released fresh healthcare research on Tuesday, offering new evidence for a broader claim that has become central to the industry’s ambitions: that A.I. may be moving from answering test questions to doing useful work in medicine and science.

OpenAI said it had built a “near-autonomous” A.I. chemist that helped improve a stubborn medicinal-chemistry reaction, and introduced a new benchmark, called LifeSciBench, aimed at measuring how well A.I. systems handle the messy, real-world tasks of biological research. Google, meanwhile, reported in new work that its medical A.I. system, AMIE, could perform on par with primary care physicians in complex disease-management scenarios.

Taken together, the releases amount to two different proof points for medical A.I. One is in the laboratory, where companies hope models can accelerate drug discovery and other forms of scientific research. The other is in the clinic, where the goal is to assist — and in some tightly defined settings perhaps rival — physicians’ diagnostic and management reasoning.

A.I. in the lab

OpenAI’s chemistry result, produced with the drug-discovery company Molecule.one and its Maria lab, centered on a difficult Chan-Lam reaction used in medicinal chemistry. According to the company, the system, built around GPT-5.4, generated hypotheses for improving reaction conditions and suggested using TEMPO-like oxidants. After 10,080 reactions, OpenAI said, the optimized conditions increased mean yield to 25.2 percent from 16.6 percent.

The company cast the result as a sign that advanced models can do more than summarize papers or answer biology questions. In this framing, the model helped generate ideas, revise experimental plans and support iterative wet-lab work.

But OpenAI also emphasized that the system was not fully autonomous. Human researchers selected which proposals to test, corrected some of the model’s plans and validated the results. That caveat matters in a field where companies often speak of automated discovery, but where the practical bottlenecks are still bound up with experimental design, lab execution and the judgment of experienced scientists.

The company’s second announcement, LifeSciBench, appears aimed at another problem confronting A.I. in science: how to tell whether a model is actually useful. The benchmark includes 750 expert-authored, expert-reviewed tasks spanning seven workflows and seven biological domains. Rather than focus on narrow, exam-like questions, it is designed to probe how systems perform on research tasks that involve ambiguity, artifacts and real decision-making.

That reflects a shift in the industry’s thinking. In life sciences, as in other fields, strong performance on standardized tests has often failed to translate into dependable help on day-to-day work. OpenAI has been pushing more deeply into drug discovery and biological research through its broader life-sciences efforts, and its message on Tuesday was that evaluation should be tied to actual workflows, not just isolated skills.

A.I. in the clinic

Google’s update made a parallel argument from the clinical side. The company said research published in *Nature* showed that AMIE, its conversational medical A.I. system, matched primary care physicians in complex disease management. The work builds on several years of staged development, beginning with diagnostic dialogue, then extending into longitudinal disease management, and now into more realistic clinical testing.

In a randomized multi-visit simulated study, Google reported that AMIE matched or exceeded clinicians in some aspects of management reasoning. And in a prospective real-world feasibility study published on March 11, 2026, the company said AMIE and primary care physicians were rated on par overall for diagnosis and management-plan quality.

Even so, the comparison came with important limits. In the real-world study, primary care doctors still did better on the practicality and cost-effectiveness of treatment plans. AMIE also operated without some of the information doctors routinely use, including access to electronic health records, findings from physical exams and broader multimodal clinical context.

Those details are not footnotes; they are central to the unresolved question of whether medical A.I. can succeed outside controlled environments. A system may reason impressively in conversation yet still fall short when decisions must account for insurance coverage, follow-up logistics, medication affordability or subtle physical cues that emerge only in person.

What the new research signals

The timing of the announcements underscores how quickly the competition over medical A.I. is evolving. For years, companies highlighted benchmark victories and exam performance, often drawing skepticism from doctors and researchers who noted that medicine and drug discovery are not multiple-choice tests. What is changing now is the emphasis on workflow.

OpenAI is arguing that A.I. can contribute to real scientific iteration in medicinal chemistry and that its capabilities should be judged on realistic research tasks. Google is making a similar case that a medical model can support the management of chronic or complicated conditions over time, not simply identify a diagnosis in a one-off exchange.

That distinction matters because the economic and social stakes are much larger in workflow than in demos. Drug discovery is expensive, slow and failure-prone; even modest gains in hypothesis generation or experimental efficiency could be valuable if they hold up in practice. Primary care, meanwhile, faces shortages, burnout and rising complexity, creating intense interest in tools that might help clinicians manage more patients without sacrificing quality.

What remains uncertain

Still, neither company’s latest result amounts to proof of broad real-world impact.

OpenAI’s chemistry system improved one challenging reaction, but it remains to be seen whether similar gains will generalize across many classes of chemistry and whether they will materially shorten the path to useful medicines. LifeSciBench may offer a more realistic way to measure model performance, but benchmark results, as its creators acknowledge, are not the same thing as downstream scientific discovery.

Google’s AMIE findings are also promising rather than conclusive. The studies suggest that the system can perform at a high level in disease-management reasoning, but they do not settle how it will behave under the full constraints of routine care, where time pressure, incomplete information, legal liability, patient trust and health-system economics all shape decisions.

The central question is no longer whether A.I. can sound informed about medicine. It is whether these systems can produce reliable gains in speed, safety, cost and outcomes when they are embedded in the actual work of clinics and laboratories.

On that question, Tuesday’s announcements offered encouraging signals — but not yet a final answer.

Sources

Further reading and reporting used to add context:

Leave a Reply

Your email address will not be published. Required fields are marked *