. 2020 Sep 11;22(9):e18091.

doi: 10.2196/18091.

Artificial Intelligence and Its Effect on Dermatologists' Accuracy in Dermoscopic Melanoma Image Classification: Web-Based Survey Study

Roman C Maron¹, Jochen S Utikal^{2

3}, Achim Hekler¹, Axel Hauschild⁴, Elke Sattler⁵, Wiebke Sondermann⁶, Sebastian Haferkamp⁷, Bastian Schilling⁸, Markus V Heppt⁹, Philipp Jansen⁶, Markus Reinholz⁵, Cindy Franklin¹⁰, Laurenz Schmitt¹¹, Daniela Hartmann⁵, Eva Krieghoff-Henning¹, Max Schmitt¹, Michael Weichenthal⁴, Christof von Kalle¹², Stefan Fröhling¹³, Titus J Brinker^#¹

Affiliations

¹ Digital Biomarkers for Oncology Group (DBO), National Center for Tumor Diseases (NCT), German Cancer Research Center (DKFZ), Heidelberg, Germany.
² Skin Cancer Unit, German Cancer Research Center (DKFZ), Heidelberg, Germany.
³ Department of Dermatology, Venereology and Allergology, University Medical Center Mannheim, University of Heidelberg, Heidelberg, Germany.
⁴ Department of Dermatology, University Hospital Kiel, University of Kiel, Kiel, Germany.
⁵ Department of Dermatology, University Hospital Munich (LMU), Munich, Germany.
⁶ Department of Dermatology, University Hospital Essen, University of Duisburg-Essen, Essen, Germany.
⁷ Department of Dermatology, University Hospital Regensburg, Regensburg, Germany.
⁸ Department of Dermatology, University Hospital Würzburg, University of Würzburg, Würzburg, Germany.
⁹ Department of Dermatology, University Hospital Erlangen, University of Erlangen, Erlangen, Germany.
¹⁰ Department of Dermatology, University Hospital Cologne, Cologne, Germany.
¹¹ Department of Dermatology, University Hospital Aachen, Aachen, Germany.
¹² Department of Clinical-Translational Sciences, Charité University Medicine and Berlin Institute of Health (BIH), Berlin, Germany.
¹³ National Center for Tumor Diseases, German Cancer Research Center, Heidelberg, Germany.

^# Contributed equally.

PMID: 32915161
PMCID: PMC7519424
DOI: 10.2196/18091

Artificial Intelligence and Its Effect on Dermatologists' Accuracy in Dermoscopic Melanoma Image Classification: Web-Based Survey Study

Roman C Maron et al. J Med Internet Res. 2020.

. 2020 Sep 11;22(9):e18091.

doi: 10.2196/18091.

Authors

Affiliations

¹ Digital Biomarkers for Oncology Group (DBO), National Center for Tumor Diseases (NCT), German Cancer Research Center (DKFZ), Heidelberg, Germany.
² Skin Cancer Unit, German Cancer Research Center (DKFZ), Heidelberg, Germany.
³ Department of Dermatology, Venereology and Allergology, University Medical Center Mannheim, University of Heidelberg, Heidelberg, Germany.
⁴ Department of Dermatology, University Hospital Kiel, University of Kiel, Kiel, Germany.
⁵ Department of Dermatology, University Hospital Munich (LMU), Munich, Germany.
⁶ Department of Dermatology, University Hospital Essen, University of Duisburg-Essen, Essen, Germany.
⁷ Department of Dermatology, University Hospital Regensburg, Regensburg, Germany.
⁸ Department of Dermatology, University Hospital Würzburg, University of Würzburg, Würzburg, Germany.
⁹ Department of Dermatology, University Hospital Erlangen, University of Erlangen, Erlangen, Germany.
¹⁰ Department of Dermatology, University Hospital Cologne, Cologne, Germany.
¹¹ Department of Dermatology, University Hospital Aachen, Aachen, Germany.
¹² Department of Clinical-Translational Sciences, Charité University Medicine and Berlin Institute of Health (BIH), Berlin, Germany.
¹³ National Center for Tumor Diseases, German Cancer Research Center, Heidelberg, Germany.

^# Contributed equally.

PMID: 32915161
PMCID: PMC7519424
DOI: 10.2196/18091

Abstract

Background: Early detection of melanoma can be lifesaving but this remains a challenge. Recent diagnostic studies have revealed the superiority of artificial intelligence (AI) in classifying dermoscopic images of melanoma and nevi, concluding that these algorithms should assist a dermatologist's diagnoses.

Objective: The aim of this study was to investigate whether AI support improves the accuracy and overall diagnostic performance of dermatologists in the dichotomous image-based discrimination between melanoma and nevus.

Methods: Twelve board-certified dermatologists were presented disjoint sets of 100 unique dermoscopic images of melanomas and nevi (total of 1200 unique images), and they had to classify the images based on personal experience alone (part I) and with the support of a trained convolutional neural network (CNN, part II). Additionally, dermatologists were asked to rate their confidence in their final decision for each image.

Results: While the mean specificity of the dermatologists based on personal experience alone remained almost unchanged (70.6% vs 72.4%; P=.54) with AI support, the mean sensitivity and mean accuracy increased significantly (59.4% vs 74.6%; P=.003 and 65.0% vs 73.6%; P=.002, respectively) with AI support. Out of the 10% (10/94; 95% CI 8.4%-11.8%) of cases where dermatologists were correct and AI was incorrect, dermatologists on average changed to the incorrect answer for 39% (4/10; 95% CI 23.2%-55.6%) of cases. When dermatologists were incorrect and AI was correct (25/94, 27%; 95% CI 24.0%-30.1%), dermatologists changed their answers to the correct answer for 46% (11/25; 95% CI 33.1%-58.4%) of cases. Additionally, the dermatologists' average confidence in their decisions increased when the CNN confirmed their decision and decreased when the CNN disagreed, even when the dermatologists were correct. Reported values are based on the mean of all participants. Whenever absolute values are shown, the denominator and numerator are approximations as every dermatologist ended up rating a varying number of images due to a quality control step.

Conclusions: The findings of our study show that AI support can improve the overall accuracy of the dermatologists in the dichotomous image-based discrimination between melanoma and nevus. This supports the argument for AI-based tools to aid clinicians in skin lesion classification and provides a rationale for studies of such classifiers in real-life settings, wherein clinicians can integrate additional information such as patient age and medical history into their decisions.

Keywords: artificial intelligence; deep learning; dermatology; diagnosis; machine learning; melanoma; neural network; nevi; skin neoplasm.

©Roman C Maron, Jochen S Utikal, Achim Hekler, Axel Hauschild, Elke Sattler, Wiebke Sondermann, Sebastian Haferkamp, Bastian Schilling, Markus V Heppt, Philipp Jansen, Markus Reinholz, Cindy Franklin, Laurenz Schmitt, Daniela Hartmann, Eva Krieghoff-Henning, Max Schmitt, Michael Weichenthal, Christof von Kalle, Stefan Fröhling, Titus J Brinker. Originally published in the Journal of Medical Internet Research (http://www.jmir.org), 11.09.2020.

PubMed Disclaimer

Conflict of interest statement

Conflicts of Interest: TJB reports owning a company that develops mobile apps (Smart Health Heidelberg GmbH, Handschuhsheimer Landstr. 9/1, 69120 Heidelberg). SH reports advisory roles for or has received honoraria from Pierre Fabre Pharmaceuticals, Novartis, Roche, BMS, Amgen and MSD outside the submitted work. AH reports clinical trial support, speaker’s honoraria, or consultancy fees from the following companies: Amgen, BMS, Merck Serono, MSD, Novartis, Oncosec, Philogen, Pierre Fabre, Provectus, Regeneron, Roche, OncoSec, Sanofi-Genzyme, and Sun Pharma. ES reports advisory roles for Heine Optotechnik GmbH and has received honoraria or travel support from Heine Optotechnik, LaRoche Posay, Naos and Abbott outside the submitted work. BS reports advisory roles for or has received honoraria from Pierre Fabre Pharmaceuticals, Incyte, Novartis, Roche, BMS and MSD, research funding from BMS, Pierre Fabre Pharmaceuticals and MSD, and travel support from Novartis, Roche, BMS, Pierre Fabre Pharmaceuticals and Amgen outside the submitted work. JSU is on the advisory board or has received honoraria and travel support from Amgen, Bristol Myers Squibb, GlaxoSmithKline, LeoPharma, Merck Sharp and Dohme, Novartis, Pierre Fabre, Roche, outside the submitted work.

Figures

**Figure 1**
Combined and individual dermatologists’ performance without and with artificial intelligence (AI) support. Every dot represents a single participant. A line between 2 dots connects the participants' metric without AI support to the corresponding metric with AI support. Highlighted dots represent the participants combined. Boxes indicate 25th and 75th percentile while the horizontal line within shows the median (50th percentile). Whiskers indicate the data range (1.5*IQR) where points beyond are considered as outliers.

**Figure 2**
Combined and individual dermatologists’ diagnostic accuracy without and with artificial intelligence (AI) support. Diagnostic accuracy is measured using sensitivity and specificity. Arrows represent the change in the diagnostic accuracy from without AI support to with AI support. Highlighted arrows represent the participants combined. In addition, the black curve denotes the mean receiver operating characteristic curve of the classifier. ROC: receiver operating characteristic; CNN: convolutional neural network; AUC: area under the curve.

See this image and copyright information in PMC

References

1. Rosendahl C, Tschandl P, Cameron A, Kittler H. Diagnostic accuracy of dermatoscopy for melanocytic and nonmelanocytic pigmented lesions. J Am Acad Dermatol. 2011 Jun;64(6):1068–73. doi: 10.1016/j.jaad.2010.03.039. - DOI - PubMed
1. Sinz C, Tschandl P, Rosendahl C, Akay BN, Argenziano G, Blum A, Braun RP, Cabo H, Gourhant J, Kreusch J, Lallas A, Lapins J, Marghoob AA, Menzies SW, Paoli J, Rabinovitz HS, Rinner C, Scope A, Soyer HP, Thomas L, Zalaudek I, Kittler H. Accuracy of dermatoscopy for the diagnosis of nonpigmented cancers of the skin. J Am Acad Dermatol. 2017 Dec;77(6):1100–1109. doi: 10.1016/j.jaad.2017.07.022. - DOI - PubMed
1. Vestergaard ME, Macaskill P, Holt PE, Menzies SW. Dermoscopy compared with naked eye examination for the diagnosis of primary melanoma: a meta-analysis of studies performed in a clinical setting. Br J Dermatol. 2008 Sep;159(3):669–76. doi: 10.1111/j.1365-2133.2008.08713.x. - DOI - PubMed
1. Tschandl P, Wiesner T. Advances in the diagnosis of pigmented skin lesions. Br J Dermatol. 2018 Jan;178(1):9–11. doi: 10.1111/bjd.16109. - DOI - PubMed
1. Brinker TJ, Hekler A, Enk AH, Klode J, Hauschild A, Berking C, Schilling B, Haferkamp S, Schadendorf D, Holland-Letz T, Utikal JS, von Kalle C, Collaborators Deep learning outperformed 136 of 157 dermatologists in a head-to-head dermoscopic melanoma image classification task. Eur J Cancer. 2019 May;113:47–54. doi: 10.1016/j.ejca.2019.04.001. https://linkinghub.elsevier.com/retrieve/pii/S0959-8049(19)30221-7 - DOI - PubMed

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Medical
- MedlinePlus Health Information

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Artificial Intelligence and Its Effect on Dermatologists' Accuracy in Dermoscopic Melanoma Image Classification: Web-Based Survey Study

Affiliations

Artificial Intelligence and Its Effect on Dermatologists' Accuracy in Dermoscopic Melanoma Image Classification: Web-Based Survey Study

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

MeSH terms

LinkOut - more resources

Full Text Sources

Medical