Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation
- PMID: 38189967
- PMCID: PMC10943141
- DOI: 10.1007/s00405-023-08372-4
Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation
Abstract
Purpose: This study aimed to evaluate the utility of large language model (LLM) artificial intelligence tools, Chat Generative Pre-Trained Transformer (ChatGPT) versions 3.5 and 4, in managing complex otolaryngological clinical scenarios, specifically for the multidisciplinary management of odontogenic sinusitis (ODS).
Methods: A prospective, structured multidisciplinary specialist evaluation was conducted using five ad hoc designed ODS-related clinical scenarios. LLM responses to these scenarios were critically reviewed by a multidisciplinary panel of eight specialist evaluators (2 ODS experts, 2 rhinologists, 2 general otolaryngologists, and 2 maxillofacial surgeons). Based on the level of disagreement from panel members, a Total Disagreement Score (TDS) was calculated for each LLM response, and TDS comparisons were made between ChatGPT3.5 and ChatGPT4, as well as between different evaluators.
Results: While disagreement to some degree was demonstrated in 73/80 evaluator reviews of LLMs' responses, TDSs were significantly lower for ChatGPT4 compared to ChatGPT3.5. Highest TDSs were found in the case of complicated ODS with orbital abscess, presumably due to increased case complexity with dental, rhinologic, and orbital factors affecting diagnostic and therapeutic options. There were no statistically significant differences in TDSs between evaluators' specialties, though ODS experts and maxillofacial surgeons tended to assign higher TDSs.
Conclusions: LLMs like ChatGPT, especially newer versions, showed potential for complimenting evidence-based clinical decision-making, but substantial disagreement was still demonstrated between LLMs and clinical specialists across most case examples, suggesting they are not yet optimal in aiding clinical management decisions. Future studies will be important to analyze LLMs' performance as they evolve over time.
Keywords: Artificial intelligence; Chronic rhinosinusitis; Computer-assisted diagnosis; Dental implant; Maxillary sinusitis; Oroantral fistula.
© 2024. The Author(s).
Conflict of interest statement
The authors have no potential conflict of interest or financial disclosures pertaining to this article.
Figures
Similar articles
-
Evaluation of the Performance of Generative AI Large Language Models ChatGPT, Google Bard, and Microsoft Bing Chat in Supporting Evidence-Based Dentistry: Comparative Mixed Methods Study.J Med Internet Res. 2023 Dec 28;25:e51580. doi: 10.2196/51580. J Med Internet Res. 2023. PMID: 38009003 Free PMC article.
-
Diagnosing odontogenic sinusitis: An international multidisciplinary consensus statement.Int Forum Allergy Rhinol. 2021 Aug;11(8):1235-1248. doi: 10.1002/alr.22777. Epub 2021 Feb 14. Int Forum Allergy Rhinol. 2021. PMID: 33583151 Review.
-
Management of odontogenic sinusitis: multidisciplinary consensus statement.Int Forum Allergy Rhinol. 2020 Jul;10(7):901-912. doi: 10.1002/alr.22598. Epub 2020 Jun 7. Int Forum Allergy Rhinol. 2020. PMID: 32506807 Review.
-
Evidence-based potential of generative artificial intelligence large language models in orthodontics: a comparative study of ChatGPT, Google Bard, and Microsoft Bing.Eur J Orthod. 2024 Apr 13:cjae017. doi: 10.1093/ejo/cjae017. Online ahead of print. Eur J Orthod. 2024. PMID: 38613510
-
Extrasinus Complications From Odontogenic Sinusitis: A Systematic Review.Otolaryngol Head Neck Surg. 2022 Apr;166(4):623-632. doi: 10.1177/01945998211026268. Epub 2021 Jul 13. Otolaryngol Head Neck Surg. 2022. PMID: 34253072
Cited by
-
Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced.Eur Arch Otorhinolaryngol. 2024 Sep;281(9):5001-5006. doi: 10.1007/s00405-024-08746-2. Epub 2024 May 25. Eur Arch Otorhinolaryngol. 2024. PMID: 38795148 Free PMC article.
-
Chat Generative Pre-Trained Transformer (ChatGPT) in Oral and Maxillofacial Surgery: A Narrative Review on Its Research Applications and Limitations.J Clin Med. 2025 Feb 18;14(4):1363. doi: 10.3390/jcm14041363. J Clin Med. 2025. PMID: 40004892 Free PMC article. Review.
-
Enhancing AI Chatbot Responses in Health Care: The SMART Prompt Structure in Head and Neck Surgery.OTO Open. 2025 Jan 16;9(1):e70075. doi: 10.1002/oto2.70075. eCollection 2025 Jan-Mar. OTO Open. 2025. PMID: 39822375 Free PMC article.
-
ChatGPT's role in alleviating anxiety in total knee arthroplasty consent process: a randomized controlled trial pilot study.Int J Surg. 2025 Mar 1;111(3):2546-2557. doi: 10.1097/JS9.0000000000002223. Int J Surg. 2025. PMID: 39903546 Free PMC article. Clinical Trial.
-
Application of machine learning in dentistry: insights, prospects and challenges.Acta Odontol Scand. 2025 Mar 27;84:145-154. doi: 10.2340/aos.v84.43345. Acta Odontol Scand. 2025. PMID: 40145687 Free PMC article. Review.