. 2024 Jan 22:26:e51926.

doi: 10.2196/51926.

Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study

Xiaocong Liu^#^{1

2}, Jiageng Wu^#², An Shao¹, Wenyue Shen¹, Panpan Ye¹, Yao Wang¹, Juan Ye¹, Kai Jin¹, Jie Yang²

Affiliations

¹ Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
² School of Public Health, Zhejiang University School of Medicine, Zhejiang, China.

^# Contributed equally.

PMID: 38252483
PMCID: PMC10845019
DOI: 10.2196/51926

Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study

Xiaocong Liu et al. J Med Internet Res. 2024.

. 2024 Jan 22:26:e51926.

doi: 10.2196/51926.

Authors

Xiaocong Liu^#^{1

2}, Jiageng Wu^#², An Shao¹, Wenyue Shen¹, Panpan Ye¹, Yao Wang¹, Juan Ye¹, Kai Jin¹, Jie Yang²

Affiliations

¹ Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
² School of Public Health, Zhejiang University School of Medicine, Zhejiang, China.

^# Contributed equally.

PMID: 38252483
PMCID: PMC10845019
DOI: 10.2196/51926

Abstract

Background: Benefiting from rich knowledge and the exceptional ability to understand text, large language models like ChatGPT have shown great potential in English clinical environments. However, the performance of ChatGPT in non-English clinical settings, as well as its reasoning, have not been explored in depth.

Objective: This study aimed to evaluate ChatGPT's diagnostic performance and inference abilities for retinal vascular diseases in a non-English clinical environment.

Methods: In this cross-sectional study, we collected 1226 fundus fluorescein angiography reports and corresponding diagnoses written in Chinese and tested ChatGPT with 4 prompting strategies (direct diagnosis or diagnosis with a step-by-step reasoning process and in Chinese or English).

Results: Compared with ChatGPT using Chinese prompts for direct diagnosis that achieved an F₁-score of 70.47%, ChatGPT using English prompts for direct diagnosis achieved the best diagnostic performance (80.05%), which was inferior to ophthalmologists (89.35%) but close to ophthalmologist interns (82.69%). As for its inference abilities, although ChatGPT can derive a reasoning process with a low error rate (0.4 per report) for both Chinese and English prompts, ophthalmologists identified that the latter brought more reasoning steps with less incompleteness (44.31%), misinformation (1.96%), and hallucinations (0.59%) (all P<.001). Also, analysis of the robustness of ChatGPT with different language prompts indicated significant differences in the recall (P=.03) and F₁-score (P=.04) between Chinese and English prompts. In short, when prompted in English, ChatGPT exhibited enhanced diagnostic and inference capabilities for retinal vascular disease classification based on Chinese fundus fluorescein angiography reports.

Conclusions: ChatGPT can serve as a helpful medical assistant to provide diagnosis in non-English clinical environments, but there are still performance gaps, language disparities, and errors compared to professionals, which demonstrate the potential limitations and the need to continually explore more robust large language models in ophthalmology practice.

Keywords: ChatGPT; artificial intelligence; clinical decision support; large language models; retinal vascular disease.

©Xiaocong Liu, Jiageng Wu, An Shao, Wenyue Shen, Panpan Ye, Yao Wang, Juan Ye, Kai Jin, Jie Yang. Originally published in the Journal of Medical Internet Research (https://www.jmir.org), 22.01.2024.

PubMed Disclaimer

Conflict of interest statement

Conflicts of Interest: None declared.

Figures

**Figure 2**
Confusion matrices of ChatGPT for the 4 prompting strategies. BRVO: branch retinal vein occlusion; CSC: central serous chorioretinopathy; CRVO: central retinal vein occlusion; DR: diabetic retinopathy; Undiag: undiagnosed; VKH: Vogt-Koyanagi-Harada disease; wetAMD: wet age-related macular degeneration.

**Figure 3**
Diagnostic performance of humans and ChatGPT. BRVO: branch retinal vein occlusion; CSC: central serous chorioretinopathy; CRVO: central retinal vein occlusion; DR: diabetic retinopathy; VKH: Vogt-Koyanagi-Harada disease; wetAMD: wet age-related macular degeneration.

See this image and copyright information in PMC

References

1. GBD 2019 BlindnessVision Impairment Collaborators. Vision Loss Expert Group of the Global Burden of Disease Study Trends in prevalence of blindness and distance and near vision impairment over 30 years: an analysis for the Global Burden of Disease Study. Lancet Glob Health. 2021 Feb;9(2):e130–e143. doi: 10.1016/S2214-109X(20)30425-3. https://linkinghub.elsevier.com/retrieve/pii/S2214-109X(20)30425-3 S2214-109X(20)30425-3 - DOI - PMC - PubMed
1. Gahlaut N, Suarez S, Uddin MI, Gordon AY, Evans SM, Jayagopal A. Nanoengineering of therapeutics for retinal vascular disease. Eur J Pharm Biopharm. 2015 Sep;95(Pt B):323–30. doi: 10.1016/j.ejpb.2015.05.001. https://europepmc.org/abstract/MED/26022642 S0939-6411(15)00219-2 - DOI - PMC - PubMed
1. Marmoy OR, Henderson RH, Ooi K. Recommended protocol for performing oral fundus fluorescein angiography (FFA) in children. Eye (Lond) 2022 Jan;36(1):234–236. doi: 10.1038/s41433-020-01328-6. https://europepmc.org/abstract/MED/33323986 10.1038/s41433-020-01328-6 - DOI - PMC - PubMed
1. Gao Z, Pan X, Shao J, Jiang X, Su Z, Jin K, Ye J. Automatic interpretation and clinical evaluation for fundus fluorescein angiography images of diabetic retinopathy patients by deep learning. Br J Ophthalmol. 2023 Nov 22;107(12):1852–1858. doi: 10.1136/bjo-2022-321472.bjo-2022-321472 - DOI - PubMed
1. Pan X, Jin K, Cao J, Liu Z, Wu J, You K, Lu Y, Xu Y, Su Z, Jiang J, Yao K, Ye J. Multi-label classification of retinal lesions in diabetic retinopathy for automatic analysis of fundus fluorescein angiography based on deep learning. Graefes Arch Clin Exp Ophthalmol. 2020 Apr;258(4):779–785. doi: 10.1007/s00417-019-04575-w.10.1007/s00417-019-04575-w - DOI - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Medical
- MedlinePlus Health Information
Research Materials
- NCI CPTC Antibody Characterization Program

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study

Affiliations

Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Medical

Research Materials