Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2023 Aug 1;6(8):e2330320.
doi: 10.1001/jamanetworkopen.2023.30320.

Comparison of Ophthalmologist and Large Language Model Chatbot Responses to Online Patient Eye Care Questions

Affiliations

Comparison of Ophthalmologist and Large Language Model Chatbot Responses to Online Patient Eye Care Questions

Isaac A Bernstein et al. JAMA Netw Open. .

Abstract

Importance: Large language models (LLMs) like ChatGPT appear capable of performing a variety of tasks, including answering patient eye care questions, but have not yet been evaluated in direct comparison with ophthalmologists. It remains unclear whether LLM-generated advice is accurate, appropriate, and safe for eye patients.

Objective: To evaluate the quality of ophthalmology advice generated by an LLM chatbot in comparison with ophthalmologist-written advice.

Design, setting, and participants: This cross-sectional study used deidentified data from an online medical forum, in which patient questions received responses written by American Academy of Ophthalmology (AAO)-affiliated ophthalmologists. A masked panel of 8 board-certified ophthalmologists were asked to distinguish between answers generated by the ChatGPT chatbot and human answers. Posts were dated between 2007 and 2016; data were accessed January 2023 and analysis was performed between March and May 2023.

Main outcomes and measures: Identification of chatbot and human answers on a 4-point scale (likely or definitely artificial intelligence [AI] vs likely or definitely human) and evaluation of responses for presence of incorrect information, alignment with perceived consensus in the medical community, likelihood to cause harm, and extent of harm.

Results: A total of 200 pairs of user questions and answers by AAO-affiliated ophthalmologists were evaluated. The mean (SD) accuracy for distinguishing between AI and human responses was 61.3% (9.7%). Of 800 evaluations of chatbot-written answers, 168 answers (21.0%) were marked as human-written, while 517 of 800 human-written answers (64.6%) were marked as AI-written. Compared with human answers, chatbot answers were more frequently rated as probably or definitely written by AI (prevalence ratio [PR], 1.72; 95% CI, 1.52-1.93). The likelihood of chatbot answers containing incorrect or inappropriate material was comparable with human answers (PR, 0.92; 95% CI, 0.77-1.10), and did not differ from human answers in terms of likelihood of harm (PR, 0.84; 95% CI, 0.67-1.07) nor extent of harm (PR, 0.99; 95% CI, 0.80-1.22).

Conclusions and relevance: In this cross-sectional study of human-written and AI-generated responses to 200 eye care questions from an online advice forum, a chatbot appeared capable of responding to long user-written eye health posts and largely generated appropriate responses that did not differ significantly from ophthalmologist-written responses in terms of incorrect information, likelihood of harm, extent of harm, or deviation from ophthalmologist community standards. Additional research is needed to assess patient attitudes toward LLM-augmented ophthalmologists vs fully autonomous AI content generation, to evaluate clarity and acceptability of LLM-generated answers from the patient perspective, to test the performance of LLMs in a greater variety of clinical contexts, and to determine an optimal manner of utilizing LLMs that is ethical and minimizes harm.

PubMed Disclaimer

Conflict of interest statement

Conflict of Interest Disclosures: None reported.

Figures

Figure 1.
Figure 1.. Distinguishing Human-Written and AI-Generated Answers to Patient Eye Questions
The chart shows the expert ratings of human and artificial intelligence (AI)-generated answers to patient questions on whether they were definitely or probably written by a human or an AI. A higher proportion of AI responses were rated as probably or definitely written by AI.
Figure 2.
Figure 2.. Expert Ratings of the Quality of Human and AI Answers to Patient Eye Questions
For each of the 4 quality questions posed to ophthalmology expert reviewers, the proportion of responses in each category are shown for the human and artificial intelligence (AI)-generated answers to patient questions.

References

    1. Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. arXiv. Published online May 24, 2019. doi:10.48550/arXiv.1810.04805 - DOI
    1. Brown TB, Mann B, Ryder N, et al. . Language models are few-shot learners. arXiv. Published online July 22, 2020. doi:10.48550/arXiv.2005.14165 - DOI
    1. Danilov G, Kotik K, Shevchenko E, et al. . Length of stay prediction in neurosurgery with Russian GPT-3 language model compared to human expectations. Inform Technol Clin Care Public Health. Published online 2022:156-159. doi:10.3233/SHTI210882 - DOI - PubMed
    1. Selivanov A, Rogov OY, Chesakov D, Shelmanov A, Fedulova I, Dylov DV. Medical image captioning via generative pretrained transformers. Sci Rep. 2023;13(1):4171. doi:10.1038/s41598-023-31223-5 - DOI - PMC - PubMed
    1. Wang SY, Huang J, Hwang H, Hu W, Tao S, Hernandez-Boussard T. Leveraging weak supervision to perform named entity recognition in electronic health records progress notes to identify the ophthalmology exam. Int J Med Inform. 2022;167:104864. doi:10.1016/j.ijmedinf.2022.104864 - DOI - PMC - PubMed