Skip to content

Assessing the Accuracy of ChatGPT in Answering Questions About Prolonged Disorders of Consciousness.

Sergio Bagnato, Cristina Boccagni, Jacopo Bonavita

Brain sciences April 13, 2025 DOI: 10.3390/brainsci15040392 via PubMed

Summary

AI-generated from the abstract

Two versions of the ChatGPT large language model (4o and o1) answered 57 open-ended questions about prolonged disorders of consciousness, written as if from a patient's relative. Both models gave predominantly correct answers (80.7-96.8% accuracy). ChatGPT 4o showed greater empathy, while ChatGPT o1 more often recommended consulting a healthcare professional (especially in Italian). English responses were more accurate than Italian only for ChatGPT 4o on clinical data. The findings suggest chatbots could help support caregivers of people with disorders of consciousness, but occasional inaccuracies mean information should be verified with a doctor.

Study at a glance

Characteristics Evaluation study Peer reviewed
Sample size 228
Population Responses from two ChatGPT models (4o and o1) to caregiver questions about prolonged disorders of consciousness
Keywords Ai in healthcare Chatgpt o1 Caregiver support Empathy Language comparison
Citations 3
Key finding Both ChatGPT models provided predominantly correct answers (80.7-96.8%) to caregiver questions about prolonged disorders of consciousness, with ChatGPT 4o showing greater empathy and ChatGPT o1 more frequently recommending professional consultation.

Abstract

Objectives: Prolonged disorders of consciousness (DoC) present complex diagnostic and therapeutic challenges. This study aimed to evaluate the accuracy of two ChatGPT models (ChatGPT 4o and ChatGPT o1) in answering questions about prolonged DoC, framed as if they were posed by a patient's relative. Secondary objectives included comparing performance across languages (English vs. Italian) and assessing whether responses conveyed an empathetic tone. Methods: Fifty-seven open-ended questions reflecting common caregiver concerns were generated in both English and Italian, each categorized into one of three domains: clinical data, instrumental diagnostics, or therapy. Each question contained a background context followed by a specific query and was submitted once to both models. Two reviewers evaluated the responses on a four-point scale, ranging from "incorrect and potentially misleading" to "correct and complete". Discrepancies were resolved by a third reviewer. Accuracy, language differences, empathy, and recommendation to consult a healthcare professional were analyzed using absolute frequencies, percentages, the Mann-Whitney U test, and Chi-squared tests. Results: A total of 228 responses were analyzed. Both models provided predominantly correct answers (80.7-96.8%), with English responses achieving higher accuracy only for ChatGPT 4o on clinical data. ChatGPT 4o exhibited greater empathy in its responses, whereas ChatGPT o1 more frequently recommended consulting a healthcare professional in Italian. Conclusions: Both ChatGPT models demonstrated high accuracy in addressing prolonged DoC queries, highlighting their potential usefulness for caregiver support. However, occasional inaccuracies emphasize the importance of verifying chatbot-generated information with professional medical advice.

Comments

No comments yet.

Log in to comment