Skip to main content
Back to News Hub
⚙️IEEE Spectrum AI
May 13, 2026
Health

Can AI Chatbots Reason Like Doctors?

Overview

A large language model from OpenAI outperformed physicians on several clinical reasoning tasks using real emergency room records, according to a study published 30 April in Science. The findings arrive amid mixed evidence about chatbot medical advice, with some studies showing strong diagnostic performance and others documenting fabricated citations and flawed advice. The study authors expressed optimism but stressed that the results should not be read as AI replacing doctors.

Key Takeaways

  • One of the earliest stated goals for computing in medicine was to aid in clinical reasoning: the decision-making steps required to reach a diagnosis and form a treatment plan.

    And over the years, researchers have built many clinical decision support systems, which have typically been purpose-built, with painstakingly written rules about symptoms, test thresholds, and medication interactions.

  • For example, this year OpenAI introduced ChatGPT for Clinicians and ChatGPT for Healthcare .
  • How Reliable Are Chatbots on Medical Matters?

    Other researchers investigating chatbots' medical advice have recently found reason to doubt their trustworthiness .

  • Much of the research focuses on chatbots answering health questions from everyday users-the kinds of questions that a person might ask before deciding to seek medical attention.

    Using an LLM as a clinical decision-support tool for doctors is a different task entirely.

  • Researchers compared two physicians and two large language models on diagnostic tasks at multiple stages of emergency-room care.

Stats & Key Facts

  • #A large language model from OpenAI outperformed physicians on several clinical reasoning tasks using real emergency room records, according to a study published 30 April in Science.
  • #Now, a large language model (LLM) from OpenAI has outperformed physicians on several clinical reasoning tasks using real emergency room records, according to a study published 30 April in Science .
Can AI Chatbots Reason Like Doctors?

One of the earliest stated goals for computing in medicine was to aid in clinical reasoning: the decision-making steps required to reach a diagnosis and form a treatment plan. And over the years, researchers have built many clinical decision support systems, which have typically been purpose-built, with painstakingly written rules about symptoms, test thresholds, and medication interactions. As artificial intelligence capabilities develop, clinical reasoning is a natural application.

Now, a large language model (LLM) from OpenAI has outperformed physicians on several clinical reasoning tasks using real emergency room records, according to a study published 30 April in Science . The new findings arrive amid a wave of concerning evidence about medical information from chatbots, with some studies showing impressive diagnostic performance while others document fabricated citations, flawed advice, and results that shift depending on how researchers score the systems. Despite that uncertainty, products aimed towards medical professionals are already entering the market.

For example, this year OpenAI introduced ChatGPT for Clinicians and ChatGPT for Healthcare . The performance of OpenAI's o1-preview, a general-purpose model that has since been supplanted by newer models, was promising enough for the authors to recommend further testing of LLMs in real life cases, with physicians seeking second opinions on diagnosis at specific checkpoints. Mickael Tordjman , who studies AI in medical imaging at the Icahn School of Medicine in New York City, agrees that the time is right for research focused on real-world applications.

For more details please read the original article at IEEE Spectrum AI.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by IEEE Spectrum AI
Read the original