Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced

Andrea Lorenzi^#¹, Giorgia Pugliese^#², Antonino Maniaci^{3

4}, Jerome R Lechien^{4

5}, Fabiana Allevi^{4

6}, Paolo Boscolo-Rizzo⁷, Luigi Angelo Vaira^#^{4

8

9}, Alberto Maria Saibene^#^{10

4}

Affiliations

¹ Division of Otolaryngology, Department of Surgical Sciences, Università degli Studi di Torino, Turin, Italy.
² Otolaryngology Unit, Santi Paolo e Carlo Hospital, Department of Health Sciences, Università degli Studi di Milano, Milan, Italy. giorgia.pugliese@unimi.it.
³ Faculty of Medicine and Surgery, "Kore" University of Enna, Enna, Italy.
⁴ International Federation of Otorhinolaryngological Societies (YO-IFOS) Head and Neck Research Group, Paris, France.
⁵ Department of Otolaryngology-Head and Neck Surgery, Foch Hospital, School of Medicine, University Paris Saclay, Paris, France.
⁶ Maxillofacial Surgery Unit, Santi Paolo e Carlo Hospital, Department of Health Sciences, Università degli Studi di Milano, Milan, Italy.
⁷ Department of Medical, Surgical and Health Sciences, Section of Otolaryngology, University of Trieste, Trieste, Italy.
⁸ Maxillofacial Surgery Operative Unit, Department of Medicine, Surgery and Pharmacy, University of Sassari, Sassari, Italy.
⁹ Biomedical Science PhD School, Biomedical Science Department, University of Sassari, Sassari, Italy.
¹⁰ Otolaryngology Unit, Santi Paolo e Carlo Hospital, Department of Health Sciences, Università degli Studi di Milano, Milan, Italy.

^# Contributed equally.

PMID: 38795148
PMCID: PMC11392976
DOI: 10.1007/s00405-024-08746-2

Comparative Study

Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced

Andrea Lorenzi et al. Eur Arch Otorhinolaryngol. 2024 Sep.

. 2024 Sep;281(9):5001-5006.

doi: 10.1007/s00405-024-08746-2. Epub 2024 May 25.

Authors

Affiliations

¹ Division of Otolaryngology, Department of Surgical Sciences, Università degli Studi di Torino, Turin, Italy.
² Otolaryngology Unit, Santi Paolo e Carlo Hospital, Department of Health Sciences, Università degli Studi di Milano, Milan, Italy. giorgia.pugliese@unimi.it.
³ Faculty of Medicine and Surgery, "Kore" University of Enna, Enna, Italy.
⁴ International Federation of Otorhinolaryngological Societies (YO-IFOS) Head and Neck Research Group, Paris, France.
⁵ Department of Otolaryngology-Head and Neck Surgery, Foch Hospital, School of Medicine, University Paris Saclay, Paris, France.
⁶ Maxillofacial Surgery Unit, Santi Paolo e Carlo Hospital, Department of Health Sciences, Università degli Studi di Milano, Milan, Italy.
⁷ Department of Medical, Surgical and Health Sciences, Section of Otolaryngology, University of Trieste, Trieste, Italy.
⁸ Maxillofacial Surgery Operative Unit, Department of Medicine, Surgery and Pharmacy, University of Sassari, Sassari, Italy.
⁹ Biomedical Science PhD School, Biomedical Science Department, University of Sassari, Sassari, Italy.
¹⁰ Otolaryngology Unit, Santi Paolo e Carlo Hospital, Department of Health Sciences, Università degli Studi di Milano, Milan, Italy.

^# Contributed equally.

PMID: 38795148
PMCID: PMC11392976
DOI: 10.1007/s00405-024-08746-2

Abstract

Purpose: This study evaluates the efficacy of two advanced Large Language Models (LLMs), OpenAI's ChatGPT 4 and Google's Gemini Advanced, in providing treatment recommendations for head and neck oncology cases. The aim is to assess their utility in supporting multidisciplinary oncological evaluations and decision-making processes.

Methods: This comparative analysis examined the responses of ChatGPT 4 and Gemini Advanced to five hypothetical cases of head and neck cancer, each representing a different anatomical subsite. The responses were evaluated against the latest National Comprehensive Cancer Network (NCCN) guidelines by two blinded panels using the total disagreement score (TDS) and the artificial intelligence performance instrument (AIPI). Statistical assessments were performed using the Wilcoxon signed-rank test and the Friedman test.

Results: Both LLMs produced relevant treatment recommendations with ChatGPT 4 generally outperforming Gemini Advanced regarding adherence to guidelines and comprehensive treatment planning. ChatGPT 4 showed higher AIPI scores (median 3 [2-4]) compared to Gemini Advanced (median 2 [2-3]), indicating better overall performance. Notably, inconsistencies were observed in the management of induction chemotherapy and surgical decisions, such as neck dissection.

Conclusions: While both LLMs demonstrated the potential to aid in the multidisciplinary management of head and neck oncology, discrepancies in certain critical areas highlight the need for further refinement. The study supports the growing role of AI in enhancing clinical decision-making but also emphasizes the necessity for continuous updates and validation against current clinical standards to integrate AI into healthcare practices fully.

Keywords: Artificial intelligence; Computer-assisted diagnosis; Head and neck cancer; Head and neck oncology; Large language models; Laryngeal carcinoma; Nasopharyngeal carcinoma; Oncological diagnosis; Oropharyngeal carcinoma; Parotid carcinoma; Tongue carcinoma.

PubMed Disclaimer

Conflict of interest statement

The authors have no potential conflict of interest or financial disclosures pertaining to this article.

References

1. Liu S et al (2023) Using AI-generated suggestions from ChatGPT to optimize clinical decision support. J Am Med Inform Assoc 30:1237–1245 10.1093/jamia/ocad072 - DOI - PMC - PubMed
1. Marchi F, Bellini E, Iandelli A, Sampieri C, Peretti G (2024) Exploring the landscape of AI-assisted decision-making in head and neck cancer treatment: a comparative analysis of NCCN guidelines and ChatGPT responses. Eur Arch Otorhinolaryngol 281:2123–2136 10.1007/s00405-024-08525-z - DOI - PubMed
1. Sarma G, Kashyap H, Medhi PP (2024) ChatGPT in head and neck oncology-opportunities and challenges. Indian J Otolaryngol Head Neck Surg 76:1425–1429 10.1007/s12070-023-04201-6 - DOI - PMC - PubMed
1. Saibene AM et al (2024) Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation. Eur Arch Otorhinolaryngol 281:1835–1841 10.1007/s00405-023-08372-4 - DOI - PMC - PubMed
1. Vaira LA et al (2023) Accuracy of ChatGPT-generated information on head and neck and oromaxillofacial surgery: a multicenter collaborative analysis. Otolaryngol Head Neck Surg. 10.1002/ohn.489 10.1002/ohn.489 - DOI - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Medical
- MedlinePlus Health Information

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced

Affiliations

Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced

Authors

Affiliations

Abstract

Conflict of interest statement

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Medical