pharmadog
News
when
  • Latest
  • Archive
by source
  • All Sources
  • Sources Page
Jobs
department
  • Clinical
  • Regulatory
  • Medical Affairs
  • Commercial
  • R&D / Discovery
  • Biostatistics / Data
  • Manufacturing / CMC
  • Market Access
therapeutic area
  • Oncology
  • Immunology
  • Neuroscience
  • Cardiovascular
  • Metabolic
  • Rare Disease
  • Infectious Disease
location & type
  • Remote Only
  • US Only
  • California
  • Massachusetts
  • Internships
  • Phase 3 Roles
  • All Jobs →
Sign InSubscribe
pharmadog

fetch the data · sniff the signal

Discover
  • Jobs
  • News
Hubs
  • Topics
  • Patent cliff
  • Publications
Tools
  • Compare
  • Search
  • Bookmarks
Trust
  • About
  • Sources
  • Contact
Legal
  • Privacy
  • Terms
  • Pricing

© 2026 pharmadog.xyz

made by humans and a good dog

  • home
  • jobs
  • news
  • search
STAT·9h ago·3 min read
save

STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated

In this edition of AI Prognosis: A conversation about benchmarking leading clinical chatbots, investor view on AI in biopharma, and more.

Jul 29, 2026·read at STAT ↗

STAT PlusNewsletterAI Prognosis Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated You’re reading the web edition of STAT’s AI Prognosis newsletter Manage alerts for this article Email this article Share this article By Brittany TrangJuly 29, 2026 Health Tech Reporter Brittany Trang[email protected]Brittany Trang, Ph.D., covers AI in health and medicine: Does it actually work? Who benefits, or might be harmed? She writes the weekly AI Prognosis newsletter.

Follow her on Threads, Mastodon, and Bluesky. You can reach Brittany on Signal at btrang.01. You’re reading the web edition of STAT’s AI Prognosis newsletter, our subscriber-exclusive guide to artificial intelligence in health care and medicine.

Sign up to get it delivered in your inbox every Wednesday. I saw “The Odyssey” during its opening weekend. Ever since then, I have been questioning whether I’m illiterate or whether Christopher Nolan is a poor storyteller.

This London Review of Books evaluation of the film, written by the woman whose translation of “The Odyssey” Nolan apparently read, has freed me from my wondering. (h/t to my colleague Matthew Herper)Advertisement Hot takes on Homer’s epic, or hot tips about Epic Systems: [email protected] Benchmark battle bots You might recall that in mid-June, there was a Nature Medicine study that pitted clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs. It set off a reaction in the clinical AI world like no other paper has.

“The results rang out like a gunshot,” as STAT health tech correspondent Katie Palmer describes it. The controversy surrounding the study, and everything that came after, exemplifies the problems I have with benchmarks.Advertisement Katie summed it up well when I talked to her yesterday: “The way that benchmarks have been talked about generally, and specifically in clinical AI, tends to summarize them into the headlines,” she said. “Every study needs a headline and every story needs a headline, but as we both know, and as I think most people in the industry know, an individual benchmark doesn’t mean much.” STAT+ Exclusive Story Already have an account?

Log in This article is exclusive to STAT+ subscribers Unlock this article — plus in-depth analysis, newsletters, premium events, and news alerts. Already have an account? Log in Monthly $39 Totals $468 per year $39/month Get Started Totals $468 per year Starter $30 for 3 months, then $399/year $30 for 3 months Get Started Then $399/year Annual $399 Save 15% $399/year Get Started Save 15% 11+ Users Custom Savings start at 25%!

Request A Quote Request A Quote Savings start at 25%! 2-10 Users $300 Annually per user $300/year Get Started $300 Annually per user View All Plans To read the rest of this story subscribe to STAT+. Subscribe Log In Artificial intelligence, health tech, STAT+ Submit a correction requestReprints Brittany Trang Health Tech Reporter Brittany Trang, Ph.D., covers AI in health and medicine: Does it actually work?

Who benefits, or might be harmed? She writes the weekly AI Prognosis newsletter. Follow her on Threads, Mastodon, and Bluesky.

You can reach Brittany on Signal at btrang.01. Newsletter Tech is transforming health care and life sciences. Our original reporting is here to keep you ahead of the curve.

Recommended AI Prognosis July 22, 2026 STAT Plus: What is a ‘world model’? Nabla’s Alex LeBrun explains AI Prognosis July 15, 2026 STAT Plus: The whistleblower, The Lab, and the fine print Advertisement AI Prognosis July 1, 2026 STAT Plus: The moment Anthropic convinced me it’s serious about science AI Prognosis June 24, 2026 STAT Plus: A dispatch on AI from BIOtech’s big summer bash AI Prognosis June 17, 2026 STAT Plus: Is Abridge’s ‘patient centered’ claim a bridge too far? Subscriber Picks

source

Reporting by STAT.

read at STAT ↗
607 words · retrieved 8h ago
sharex / twitterlinkedin

comments(0)

5-min edit window · permanent after that
sign in to leave a comment · permanent archive after 5 minutes
no comments yet — first sniff?

companies & drugs in this story

no entities indexed yet

related stories

  • 1h agoNSF testing new Ph.D. model giving trainees academic and industry experienceSTAT
  • 2h agoSTAT+: Latigo reports mid-stage success for would-be rival to Vertex’s pain drug JournavxSTAT
  • 4h agoSTAT+: Hims misled consumers and violated their privacy, FTC alleges in a lawsuitSTAT
  • 8h agoSTAT+: End of Medicare drug subsidy gives Democrats new attack line on rising out-of-pocket costsSTAT
  • 9h agoSTAT+: Pharmalittle: We’re reading about a fugitive who became a biotech exec, a hospital suing Lilly, and moreSTAT