Study finds clinical chatbots lag behind general AI models in accuracy, sparking debate among doctors
Hundreds of thousands of U.S. physicians are using clinical large language model tools marketed by firms such as OpenEvidence, Doximity and UpToDate as safer alternatives to generic AI chatbots.
A research team from NYU Langone Health evaluated both general purpose and clinical focused models on three sets of medical questions, and the results were published in Nature Medicine in June.
The study showed the clinical specific AI performed worse than the general models, a finding that surprised many and prompted a strong reaction on social media, including a comment from Kaiser Permanente's vice president of AI.
Developers of the clinical tools argue that current benchmarking methods do not accurately capture safety and usefulness, highlighting ongoing debate over how to assess AI in healthcare.
This writeup was produced by pharmadog from original reporting by STAT.
Original headline: “STAT+: Clinical chatbots are taking medicine by storm. Should doctors trust them?”
read at STAT ↗
comments(0)
5-min edit window · permanent after that