Large Language Models in peri-implant disease: How well do they perform?

Vasiliki P Koidou¹, Georgios S Chatzopoulos², Lazaros Tsalikis³, Eleutherios G Kaklamanos⁴

Affiliations

¹ Research Associate, Centre for Oral Immunobiology and Regenerative Medicine and Centre for Oral Clinical Research, Institute of Dentistry, Queen Mary University of London (QMUL), London, England, UK. Electronic address: v.koidou@qmul.ac.uk.
² PhD candidate, Department of Preventive Dentistry, Periodontology and Implant Biology, School of Dentistry, Aristotle University of Thessaloniki, Thessaloniki, Greece; and Visiting Research Assistant Professor, Division of Periodontology, Department of Developmental and Surgical Sciences, School of Dentistry, University of Minnesota, Minneapolis, Minn.
³ Professor, Department of Preventive Dentistry, Periodontology and Implant Biology, School of Dentistry, Aristotle University of Thessaloniki, Thessaloniki, Greece.
⁴ Associate Professor, Department of Preventive Dentistry, Periodontology and Implant Biology, School of Dentistry, Aristotle University of Thessaloniki, Greece; Associate Professor, School of Dentistry, European University Cyprus, Nicosia, Cyprus; and Adjunct Associate Professor, Hamdan bin Mohammed College of Dental Medicine, Mohammed bin Rashid University of Medicine and Health Sciences (MBRU), Dubai, United Arab Emirates.

PMID: 40055086
DOI: 10.1016/j.prosdent.2025.02.008

Large Language Models in peri-implant disease: How well do they perform?

Vasiliki P Koidou et al. J Prosthet Dent. 2025.

. 2025 Mar 6:S0022-3913(25)00102-7.

doi: 10.1016/j.prosdent.2025.02.008. Online ahead of print.

Authors

Vasiliki P Koidou¹, Georgios S Chatzopoulos², Lazaros Tsalikis³, Eleutherios G Kaklamanos⁴

Affiliations

¹ Research Associate, Centre for Oral Immunobiology and Regenerative Medicine and Centre for Oral Clinical Research, Institute of Dentistry, Queen Mary University of London (QMUL), London, England, UK. Electronic address: v.koidou@qmul.ac.uk.
² PhD candidate, Department of Preventive Dentistry, Periodontology and Implant Biology, School of Dentistry, Aristotle University of Thessaloniki, Thessaloniki, Greece; and Visiting Research Assistant Professor, Division of Periodontology, Department of Developmental and Surgical Sciences, School of Dentistry, University of Minnesota, Minneapolis, Minn.
³ Professor, Department of Preventive Dentistry, Periodontology and Implant Biology, School of Dentistry, Aristotle University of Thessaloniki, Thessaloniki, Greece.
⁴ Associate Professor, Department of Preventive Dentistry, Periodontology and Implant Biology, School of Dentistry, Aristotle University of Thessaloniki, Greece; Associate Professor, School of Dentistry, European University Cyprus, Nicosia, Cyprus; and Adjunct Associate Professor, Hamdan bin Mohammed College of Dental Medicine, Mohammed bin Rashid University of Medicine and Health Sciences (MBRU), Dubai, United Arab Emirates.

PMID: 40055086
DOI: 10.1016/j.prosdent.2025.02.008

Abstract

Statement of problem: Artificial intelligence (AI) has gained significant recent attention and several AI applications, such as the Large Language Models (LLMs) are promising for use in clinical medicine and dentistry. Nevertheless, assessing the performance of LLMs is essential to identify potential inaccuracies or even prevent harmful outcomes.

Purpose: The purpose of this study was to evaluate and compare the evidence-based potential of answers provided by 4 LLMs to clinical questions in the field of implant dentistry.

Material and methods: A total of 10 open-ended questions pertinent to prevention and treatment of peri-implant disease were posed to 4 distinct LLMs including ChatGPT 4.0, Google Gemini, Google Gemini Advanced, and Microsoft Copilot. The answers were evaluated independently by 2 periodontists against scientific evidence for comprehensiveness, scientific accuracy, clarity, and relevance. The LLMs responses received scores ranging from 0 (minimum) to 10 (maximum) points. To assess the intra-evaluator reliability, a re-evaluation of the LLM responses was performed after 2 weeks and Cronbach α and interclass correlation coefficient (ICC) was used (α=.05).

Results: The scores assigned by the examiners on the 2 occasions were not statistically different and each LLM received an average score. Google Gemini Advanced ranked higher than the rest of the LLMs, while Google Gemini scored worst. The difference between Google Gemini Advanced and Google Gemini was statistically significantly different (P=.005).

Conclusions: Dental professionals need to be cautious when using LLMs to access content related to peri-implant diseases. LLMs cannot currently replace dental professionals and caution should be exercised when used in patient care.

PubMed Disclaimer

LinkOut - more resources

Full Text Sources
- ClinicalKey
- Elsevier Science
Research Materials
- NCI CPTC Antibody Characterization Program

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Large Language Models in peri-implant disease: How well do they perform?

Affiliations

Large Language Models in peri-implant disease: How well do they perform?

Authors

Affiliations

Abstract

LinkOut - more resources

Full Text Sources

Research Materials