• Vol. 55 No. 7, 370–382
  • 14 July 2026
Accepted: 30 June 2026 | Published Online First: 14 July 2026

Artificial intelligence tools in dermatology education: A scoping review on their application, efficacy, and limitations

,

ABSTRACT

Introduction: Artificial Intelligence (AI) is increasingly explored for medical education, and dermatology, with its visual diagnostic focus, holds promise for AI-enhanced learning. However, evidence on its educational effectiveness remains limited and fragmented. This scoping review aimed to assess the evidence base for applications and limitations of AI-based educational tools in dermatology.

Methods: A systematic search of PubMed, Embase, Web of Science, Scopus, and PsycINFO up to July 2025 was conducted. Data were synthesised narratively, considering types of AI interventions and their assessed outcomes. These were analysed with the Cost, Usability, Credibility, Fairness, Accountability, Transparency, Explainability (CUC-FATE) framework. Study quality was assessed with ROBINS-I and a COSMIN-informed checklist. 

Results: A total of 827 records were screened, with 360 duplicates removed, yielding 467 studies. o full-length studies and 1 conference abstract met inclusion criteria, mostly from 2023 to 2025. These explored AI-generated clinical images, Large Language Model-generated vignettes, intelligent tutoring systems, and clinical decision support tools. Content validation studies generally reported favourable ratings for accuracy, clarity, and educational utility, while intervention studies suggested possible benefits for learning performance and diagnostic accuracy. Usability and credibility were commonly assessed, whereas cost, accountability, fairness, transparency, and explainability were rarely examined.

Conclusions: Most of the studies reviewed had high risk of bias, small sample sizes, and limited methodological rigour, with significant heterogeneity in examined educational outcomes limiting synthesis. While current initial studies on AI hold promise, this scoping review underscores the need for more robust studies with standardised evaluation frameworks, prioritising ethical principles such as fairness, explainability, and accountability for safe integration in training.


CLINICAL IMPACT

What is New

  • This scoping review focuses on the emerging evidence of the direct educational impact of AI-based educational tools in dermatology and identifies major gaps in methodological quality, standardised evaluation, and ethical oversight.
  • As the use of AI in healthcare and medical education expands, these findings provide a preview of opportunities and challenges in AI-based dermatology training.

Clinical Implications

  • The findings support the need for rigorous and ethically grounded frameworks to guide safe integration of AI into dermatology education, with emphasis on educational effectiveness, fairness, explainability, and accountability.


Artificial Intelligence (AI) technologies—particularly deep learning algorithms and Large Language Model (LLM) computer systems—have progressed at an unprecedented pace in recent years, driving innovation across a wide range of healthcare applications.1 These advancements have resulted in a corresponding surge in research on the application of AI in medical education,2 including the development of virtual reality simulators3 and intelligent tutoring systems.4 Such tools enable the provision of individualised, innovative and adaptive learning experiences, allowing students and physicians to practise medical reasoning.5,6 Moreover, integrating AI into medical education allows for distance learning, making medical education more accessible in resource-constrained areas with limited resources7 or time due to clinical demands.8

In the field of dermatology, its inherent visual diagnostic process makes it an attractive possibility for AI to be applied as an educational tool,9 with similar promise being noted in other visually intensive specialties, such as radiology, for the enhancement of clinical competencies and optimising learning outcomes.10 AI applications—ranging from machine learning-based diagnostic simulators to image-based learning platforms11—have been proposed to improve diagnostic accuracy, engagement, and learning efficiency.12 However, the evidence evaluating the effectiveness of these technologies in dermatology education—such as in assessing improvement in diagnostic accuracy, knowledge retention, clinical decision-making—remains fragmented and underexplored, with robust evidence for their use in dermatology education currently lacking.13

As such, this scoping review aims to address this gap, examining the current methods by which AI has been applied to dermatological education. By charting the existing research landscape and evaluating current evidence, this scoping review aims to highlight promising applications and identify limitations in existing studies. In doing so, this study aims to suggest future innovations for more efficacious AI usage, including areas of learner and faculty development necessary to integrate AI into dermatological education.

METHODS

The protocol detailing the search strategies of this review was formulated by the authors, who registered this on the International Prospective Register of Systematic Reviews (PROSPERO, registration number CRD420251058835) on 22 May 2025 before the study’s initiation. While initially designed as a systematic review, in view of the small number of highly heterogeneous studies identified, the study was redesigned as a scoping review after article identification to provide a broader map of existing evidence. With a narrow evidence base, assessment on efficacy would be limited, and hence research questions were refined to define the applications and limitations of the selected studies.

Search strategy

The search included the databases of MEDLINE (through PubMed), Embase, Web of Science, SCOPUS, and PsycINFO, with the search strategies described in Supplementary Table S1. All studies published from the inception of the respective database till 1 July 2025 were included. Based on titles and abstracts, duplicate studies were first excluded. The remaining studies were then evaluated based on titles and abstracts for their relevance to the use of artificial intelligence in dermatological education.

Selection of articles

Studies published in any language, from all countries, were considered for review. Randomised controlled trials, non-randomised controlled studies, pre-post intervention studies, and observational studies were included if they referred to direct application of AI in dermatological education. Efforts were made to obtain full-length studies for findings highlighted in conference abstracts. Studies were considered at various levels of healthcare education, including that of medical students, residents (especially dermatology and general medical trainees), and non-dermatology clinicians involved in dermatology learning. Tools used solely for clinical diagnosis without an educational aim, editorials, opinion pieces, protocols without results, and studies focused solely on AI tool development without educational evaluation were excluded. As title and abstract screening was done jointly, formal recording of reviewer disagreements and calculation of inter-rater reliability were not applicable.

Data collection

Data were extracted via a standardised collection form, with authors’ names, publication date, country of publication (as defined by the setting of data collection), type of study, and level of healthcare education of the intended learner audience. Following this, categorisation was performed of the types of artificial intelligence educational tools assessed, any specific learning outcomes examined and overall conclusions.  

Data analysis

In this review, the articles identified displayed significant heterogeneity, preventing meaningful pooling and synthesis of educational outcomes. As such, a scoping review was conducted according to the methodological framework proposed by Arksey and O’Malley14 and refined by Levac et al.,15 and was reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) checklist (Supplementary Table S2).16 With the extracted data, a narrative account of findings was produced, considering the types of AI interventions and the educational outcomes assessed. The limitations and gaps in these studies were also identified.

The studies were categorised based on various aspects of AI technologies specifically measured by each study. Utilising an approach developed by Quttainah et al.17 examining enablers for the use of LLMs in medical education, an assessment was made on whether the parameters of Cost, Usability, Credibility, Fairness, Accountability, Transparency, and Explainability (CUC-FATE) were considered by each included study. 

Quality assessment

Only non-randomised studies were obtained, and their methodological quality was assessed. Study quality of quasi-experimental studies was assessed using the Cochrane Collaboration’s Risk Of Bias In Non-randomized Studies of Interventions (ROBINS-I) tool,18 modified for usage in before-after studies. As no established appraisal tool exists for validation of AI-based educational content, content validation studies were assessed based on a framework devised by the authors, informed by the COnsensus-based Standards for the selection of health Measurement Instruments (COSMIN) checklist for content validity.19 This framework examined core tenets of the COSMIN checklist, including relevance, comprehensiveness and clarity, as well as recommendations regarding expert involvement and qualitative evaluation during instrument development.

There are also currently no suitable frameworks for assessing the quality of AI-powered interventions in medical education. The commonly used Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) criteria, while modified recently to account for machine learning methods,20 work to evaluate predictive models by their performance, and cannot be directly applied to effectiveness in medical education. As such, a framework previously utilised by Bates et al.,21 also referenced in other reviews on artificial intelligence in healthcare education,22 was utilised.

RESULTS

Search results

The comprehensive search of PubMed, Embase, Web of Science, Scopus, and PsycINFO yielded 827 studies, of which 360 duplicates were excluded (Fig, 1). From the remaining 467 studies, 406 were excluded based on titles and abstracts. In total, 61 studies were retrieved and assessed for eligibility, of which four full-length studies and one conference abstract were identified. Further search of citations yielded two other relevant articles,23-24 with a total of 6 full-length studies and 1 conference abstract23-29 selected for review.

Fig. 1. Database search.

A decision was made to include a conference abstract27 to allow comprehensive mapping of the dermatological educational landscape, in accordance with a proposal by Scherer et al.30 An attempt was made to contact the authors and examine study registers and published protocols to obtain further information on methods and results, but without success. However, due to its brevity, there was limited ability to appraise the abstract for its design, risk of bias, and assessed educational outcomes, which is noted in the subsequent discussion. The PRISMA flowchart for this study was created with the aid of an app designed by Haddaway et al.31

Description of included studies

The general characteristics of studies are summarised in Table 1, with specific outcomes examined summarised in Table 2 and further laid out in a frequency table for content validation studies (Supplementary Table S3). Out of the 6 identified studies and 1 conference abstract (henceforth considered with the rest of the studies), 6 of these were conducted over the past 3 years (2023–2025)23-28—with 1 study focusing on an intelligent tutoring system in 2008.29 Four of the studies were conducted in the US,24,27-29 and 1 each in Italy, Australia, and Germany.23,25,26

Table 1. Description of included studies.

CME: continuing medical education; LLM: large language model

Table 2. Results of included studies.
AI: artificial intelligence; BNS: best-next-step; VR: virtual reality
a Only 95 out of the 115 items were deemed eligible for assessment of consistency by the authors.
b This confidence is not symmetric about the point estimate with a much higher upper bound (4.89), which may be a possible error in the original manuscript.

Three of the studies (including the conference abstract) were assessed as quasi-experimental, defined here as studies analysing the effect of interventions without the assignment of experimental and control groups26-27,29; while the remaining 4 studies23-25,28 were designed as cross-sectional studies for content validation, defined here as studies which focused on rating the quality and accuracy of AI output. Three studies had a stated targeted audience who were residents,25,28,29 three studies had a stated targeted audience who were medical students,24,26,27 and 1 study general practitioners.23

AI technologies examined

Many studies focused on the utilisation of generative AI, including 2 for image generation25,26 and 3 for text generation.23.24,28 All studies examining text generation utilised ChatGPT in its various models, with 2 studies24,28 attempting to generate questions for use in clinical examinations. The last study23 attempted to assess the LLM’s ability to answer questions about acne vulgaris for education and was unique in assessing the LLM’s ability to identify appropriate reference literature for learners.

Out of the 2 remaining studies, the first evaluated an intelligent tutoring system in dermatopathology serving as a precursor to modern AI today.29 The second, a conference abstract, provided a different approach, utilising a clinical decision support system for medical students in an outpatient setting, creating a list of differential diagnoses during patient review.27

Educational outcomes examined

Both full-length quasi-experimental studies26,29 attempted to examine the improvement of clinical competence, with 1 image-generation study26 looking at the improvement in learning accuracy (albeit in a small sample of students, which cannot be interpreted as evidence of educational benefit), while another study on intelligent tutoring29 looked at the frequency of errors. The conference abstract27 evaluating a clinical decision support system appeared to only look at patient perception and utilisation time of the AI tool, although the absence of a full-text limits assessment. The significant heterogeneity of these educational outcomes prevented further synthesis and comparison.

All studies designed for content validation attempted to break down the characteristics of AI-generated output important for educational use. The 3 studies focused on LLMs23,24,28 looked at accuracy and clarity of material generated, though they diverged significantly in aspects they found important for education. The content validation study for AI-generated images25 looked similarly at accuracy and utility for clinical education.

Quality assessment

The 3 quasi-experimental trials (including the conference abstract) were analysed with the ROBINS-I tool with all studies assessed having a serious overall risk of bias (Table 3). None of the trials provided data on the participant recruitment or measures to control for confounders. In the conference abstract,27 while there was limited information to assess for the risk of bias, the authors only recruited two medical students to test the intervention.

Table 3. Cochrane risk-of-bias assessment tool for non-randomised studies.

NI: no information; PN: probably no; PY: probably yes
a Assessment for this study was limited as a conference abstract

In assessing content validation studies on AI-generated materials,23-25,28 a content validity checklist informed by the COSMIN approach was adopted19 (Table 4). All studies had clear descriptions of measured constructs and assessed relevant (but different) characteristics of AI-generated materials. They also all used clear Likert scales for assessment, involving experts in its development and reporting on inter-rater agreement. Three out of 4 studies were significantly lacking in having a small number of individuals assessing the suitability of teaching material, and only 1 out of 4 studies reported qualitative insights.

Table 4. COSMIN-inspired assessment for content validation studies.

CME: continuing medical education; JAAD: Journal of the American Academy of Dermatology; USMLE: United States Medical Licensing Examination

When analysed using the approach by Bates et al.21 (Supplementary Table S4), the accuracy of the AI-based educational technologies was often unvalidated, and information with regard to uncertainty and data-handling was not provided in all studies. 

Enablers of AI in medical education

Studies were assessed for examination of the key features of the CUC-FATE framework17 (Table 5) for effective and safe usage of AI in medical education, but many of these elements were not assessed in identified studies. While all studies examined the usability of AI technology, and 6 out of 7 studies examined their credibility, none of the studies assessed issues regarding cost.

Table 5. Assessment of studies using the CUC-FATE framework.

NR: not reported; R: reported
a Assessment for this study was limited as a conference abstract.

Two of the studies26,29 reported AI interventions created by the authors, allowing assessment of their transparency, but studies that utilised proprietary LLMs, such as the content validation studies, lacked the ability to do so. Author-created AI interventions in these studies also allowed assessment of explainability. Including a study23 which asked the LLM to generate reference literature for provided information, a total of 3 studies specifically assessed explainability.

This same study also highlighted the presence of inaccuracies and errors in the LLM’s responses, urging the necessity of expert review to prevent misinformation, thus being the only study assessing accountability.23 Only 1 study that looked at AI-generated questions considered fairness, specifically assessing the risk of demographic bias in AI-generated questions and concluding that the risk of this is low.24

DISCUSSION

This scoping review has identified a small but emerging body of literature examining the application of AI in dermatological education. Though the included studies examined a range of AI-based interventions, such as AI-generated images of dermatological conditions and LLM-generated educational content, the overall evidence base remains limited. Most studies were preliminary, small-scale and heterogeneous in design, and current evidence is insufficient in supporting conclusions regarding its effect on improving learner performance and clinical competence.

The review also highlights several questions that are still unanswered. It remains unclear which AI-based interventions are the most effective for different groups, such as medical students, residents, and general practitioners, and whether these interventions necessarily provide meaningful added value over existing traditional educational methods. Important concerns on implementation, such as transparency and explainability, are also inconsistently addressed in studies, and represent critical concerns for educators.

Since completion of the search by this review, a broader scoping review of AI and dermatological education has been published by Lau et al.,32 also highlighting educational applications, including the use of LLMs as well as generative and adaptive imaging tools. In contrast, this present review focused on studies examining direct educational applications, deliberately excluding the substantial number of studies solely assessing AI performance in answering dermatology examination questions. By adopting this narrower educational focus, this review provides a more targeted appraisal of the current evidence for the use of AI in dermatological education, concluding that despite growing interest in the field, current evidence is sparse and limits strong conclusions.

Nevertheless, this field continues to evolve rapidly. For example, a study published after the search period examining the usage of LLMs as a dermatology case narrator for medical students33 would have been eligible for this review. As this study was published after the prespecified search period, it was not included in the formal synthesis, but its relevance suggests that future updated reviews may identify a larger body of eligible evidence.  

The findings of this review are also consistent with emerging evidence from the broader medical education literature, with a rapidly increasing number of AI-related publications over the past few years.6 A scoping review of AI medical image generation across health professions education (including dermatology) by Gupta et al.34 also concluded that generative AI should be treated as an experimental adjunct requiring rigorous human review. Similarly, in radiology education (a similarly visually intensive specialty), a scoping review of AI-based educational approaches by Hui et al.35 identified a range of possible applications, from personalised curriculum generation and diagnostic support tools to automated evaluation systems.

While similar themes were observed in the present review, the evidence base identified for dermatology was considerably smaller, with fewer studies evaluating educational outcomes and learner performance. Together, these reviews suggest that while dermatology shares characteristics with more visual specialties—making it potentially well-suited for AI-enhanced learning, robust evidence demonstrating educational effectiveness is lacking, making further research with rigorous validation necessary to identify the most effective AI tools in education.36 Other potential interventions not included in this review, such as personalised learning platforms,37 integrating AI with virtual-reality based training for skin cancer screening38 as well as AI-assisted learning for knowledge and milestone assessment,39 have been proposed by other authors, and would benefit from implementation studies to assess educational outcomes.

Key limitations  

There was considerable heterogeneity noted in study design and evaluation metrics for educational efficacy in quasi-experimental studies, with a diversity of intervention types explored. These studies were also small-scale with low sample sizes (ranging from 2 to 44), with few participants and inadequate control for confounding variables causing serious risk of bias. None of the studies employed validated instruments to assess learning outcomes, often relying on self-developed Likert scales or informal expert ratings. Similarly, none of the content validation studies used standardised measures to assess content validity of AI-generated information. This lack of standardisation prevents meaningful synthesis.

None of the studies examined compared AI-based interventions with traditional teaching methods. This absence of a control group limited conclusions about added educational value and generalisability, and such comparative analysis is essential in evaluating AI-based interventions’ educational effectiveness.7

All studies were conducted in pilot or isolated research settings, outside the realm of formal medical school or residency training curricula, suggesting institutional adoption of AI in dermatology education is currently limited. This is consistent with existing research, which suggests that adoption of AI into medical education is still in its early developmental stages.40 A key challenge is the alignment of existing educational frameworks and curricula with rapidly evolving AI capacities,41 with educational centres often struggling with infrastructure and policy support required.42

A further limitation relates to the review process itself, where study screening was conducted through consensus between 2 reviewers rather than independent duplicate reviews. While this collaborative approach allowed for discussion and agreement at each stage, it may have increased susceptibility to shared reviewer bias.

CUC-FATE framework

This scoping review utilised the CUC-FATE framework to examine crucial enablers for AI in medical education, demonstrating that most dimensions were unassessed by included studies. Credibility and utility were frequently assessed, and while an earlier content validation study28 using an older LLM noted high levels of inaccuracy, this was not noted in newer LLM based content validation studies23,24 utilising newer models of the same LLM. This underscores how LLMs are rapidly evolving and improving, and older studies with less positive results may be less relevant with newer generative AI models.

Most studies lacked reporting of assessments of transparency and explainability. Content validity is limited by the inability to trace presented evidence to its origins, which is especially pertinent when proprietary technologies, such as the LLMs ChatGPT or Gemini, are used.43,44 These deficiencies may undermine clinicians’ trust in AI-assisted dermatological education, particularly when verification of accuracy in clinical contexts is essential.45

Encouragingly, an included image-generation study26 described the creation of a custom diffusion model, providing detailed methodological transparency. While other studies focusing on deep learning algorithms and clinical decision support systems for diagnosis of melanoma46,47 and other skin diseases48 were excluded from review because of the absence of educational focus, their transparent designs would allow effective adaptation for educational purposes.

Ethical concepts of accountability and fairness were rarely assessed in included studies. With the use of AI in dermatological education rapidly expanding, these omissions are concerning, given that the identification of ethical challenges is paramount for responsible development and implementation.49 The issue of accountability, establishing liability for harmful results sustained from incorrect educational content,50 was assessed in only 1 study23 which emphasised the need for human oversight to ensure clinical accuracy. For fairness, AI algorithmic bias can lead to error in diagnosis of under-represented populations such as individuals of colour,51 causing erroneous education of medical professionals. This was briefly assessed in one study on AI-generated dermatological examination questions,24 but should be more widely evaluated.

Implications for practice and future research

Despite promising innovation in the application of AI in dermatological education, this review has demonstrated that significant gaps remain, highlighting key areas for future development. Current quasi-experimental studies had a high rate of bias, emphasising the need for larger, robust, multicentre studies involving diverse participant cohorts in institutional contexts for both medical student and healthcare professional training in dermatology. 

Significant methodological heterogeneity across studies severely limited comparison, highlighting the need for standardised reporting frameworks and verified evaluation tools to ensure replicability. This review has demonstrated a framework for content validation studies devised by the authors and informed by the COSMIN approach, and the CUC-FATE framework, which incorporates not just practical but important ethical considerations. Adopting such standardised frameworks would improve the quality and transparency of AI-related educational research in dermatology, allowing appropriate and effective implementation. There is a need for studies to explicitly incorporate ethical principles in their design, ensuring AI-based educational tools are equitably and responsibly implemented.

Singapore has emerged as an early regional adopter of AI in healthcare training, resulting in an increased impetus to equip faculty and students with the necessary foundational knowledge in AI in a forward-looking medical education landscape.52 In this aspect, efforts have been made to design curricula locally to integrate AI into medical education,53 with local medical educators highlighting the benefits in equipping medical learners with skills in digital technology.54 By using dermatology as a targeted case study, this review provides timely evidence that can inform curriculum design for AI-enhanced medical education in Singapore and similar healthcare systems.

Given the diversity of AI interventions, we recommend that further research involve learners in their design, which can improve engagement and reduce barriers to adoption. Multiple studies have consistently demonstrated that dermatology specialists,55 trainees,56 and students,40,57 while agreeing that AI will play a key role in dermatological education, expressed that they lacked training to utilise it effectively. A mixed-methods approach, combining quantitative outcomes of efficacy with qualitative data of learner experience, would be useful in examining these perceptions to enhance usage of AI in dermatological education.

Beyond the design of learner-facing interventions, successful implementation of AI in dermatological education will also need investment in faculty development. Recent medical education literature has emphasised the importance of equipping educators with AI literacy through structured faculty development models, which can include competency-based certification and interdisciplinary training.58,59 As AI becomes increasingly integrated into educational practice, frameworks for faculty development have been devised in broader medical education contexts,60 although these have not been explored specifically in dermatology. Future work should therefore examine not only the effectiveness of AI interventions for learners, but also the competencies and support required for educators to implement these tools safely and effectively.

CONCLUSION

This review suggests that AI is an evolving tool for dermatology education, with potential benefits for learners. Existing literature is characterised by methodological heterogeneity and limited validation, and there is insufficient evidence to determine whether AI-based educational tools improve learning outcomes compared with standard educational approaches. Establishing standardised frameworks for the development, validation, and evaluation of AI-based educational technologies is crucial to generate higher-quality evidence, allowing for effective and ethical implementation in dermatology training.

Supplementary materials

Supplementary Table S1. Search strategy.
Supplementary Table S2. Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) checklist.
Supplementary Table S3. Frequency table of outcomes of included content validation studies.
Supplementary Table S4. Assessment of studies using the framework proposed by Bates et al.

Ethics statement

This research did not involve the collection of patient data, human participants, or biological samples. All analyses were conducted on publicly available, published literature.

Declaration

The authors declare they have no affiliations or financial involvement with any commercial organisation with a direct financial interest in the subject or materials discussed in the manuscript.

Correspondence

Dr Ziying Vanessa Lim, National Skin Centre, 1 Mandalay Road, Singapore 308205. Email: [email protected]