What a Reddit study reveals about ChatGPT and medical answers

In a study of 195 questions from Reddit's r/AskDocs forum, readers rated ChatGPT's answers higher than physicians' for quality and empathy. The findings suggest AI could help clinicians draft replies, but the study did not assess accuracy or fabricated information, and more research is needed.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The study suggests ChatGPT may assist clinicians, but it does not establish accuracy or patient outcomes.

What a Reddit study reveals about ChatGPT and medical answers

People who ask health questions online may value an answer that is clear, considerate and timely. A study comparing ChatGPT responses with physician replies found that readers preferred the chatbot's answers for quality and empathy. That result points to a possible supporting role for AI in patient messaging, while leaving a central question unresolved: whether the answers were accurate.

Readers favored the chatbot's replies

The cross-sectional study examined 195 randomly selected questions posted to Reddit's r/AskDocs forum. Researchers compared responses from ChatGPT with responses written by physicians. Chatbot responses were preferred, and ChatGPT received significantly higher ratings for both quality and empathy.

The study used GPT 3.5, an older version of the chatbot. The article notes that GPT-4 could provide better results, but the study's reported comparison was based on GPT 3.5. The ratings therefore describe how readers judged those responses in this particular setting, rather than establishing how every version of a chatbot performs across medical conversations.

Quality and empathy matter in an online exchange because the response is the interaction. A message that feels more understandable or attentive may be easier for a person to engage with. But favorable ratings alone cannot tell readers whether medical advice is correct, complete or suitable for an individual patient.

AI could help doctors answer messages

The researchers describe AI assistants as a possible drafting aid for clinicians. In this approach, a chatbot would prepare a response for a doctor to review and edit. That could give clinicians a starting point for answering patient questions while keeping a human involved in the final message.

The authors say the idea needs further exploration in clinical settings. Randomized trials could examine whether AI assistants improve responses, reduce clinician burnout and improve patient outcomes. Those are proposed areas for study, not results demonstrated by the comparison of forum answers.

Prompt, empathetic replies could also have wider practical value. The article says they might reduce unnecessary clinical visits and free up resources. Messaging may support patient equity for people with mobility limitations, irregular work schedules or fear of medical bills. These are potential benefits the article identifies; the forum study itself did not measure them.

Important questions remain unanswered

Questions posted to an online forum may not represent ordinary conversations between patients and physicians. The study also did not assess whether a chatbot could use personalized details from electronic health records. Without that information, its comparison cannot show how AI replies would work when a response needs to account for a patient's individual medical context.

Most significantly, the study did not evaluate the chatbot's answers for accuracy or fabricated information. A response can sound empathetic and polished while still containing a mistake. The article cautions that errors can be harder to catch than to avoid, even when a person reviews the final answer. Physicians can make mistakes too, but that does not remove the need to check AI-generated content.

The authors call for more research into clinical use and say ethical concerns must be addressed. Their recommendations include human review for accuracy and for false or fabricated information. Until those questions are examined, the ratings should be understood as a comparison of perceived response quality and empathy, not proof that chatbot medical advice is reliable.

Medical use requires a different kind of evidence

ChatGPT is not specifically optimized for medical tasks, according to the article. It also mentions Google's development of Med-PaLM 2, a large language model fine-tuned for medical purposes. Google claims it can pass medical exams and plans to start trials with professionals. These details illustrate interest in medical AI, but they do not establish that the systems are ready to answer patient questions independently.

The study offers a reason to explore AI-assisted messaging: readers preferred the chatbot's answers in one comparison. Whether that preference can translate into safe, useful clinical support remains to be tested. Research will need to consider not just how an answer feels, but whether it is accurate, appropriately personalized and improved by clinician review.