Predictive Capability of Both GPT-4 and GPT-5 for Histopathological Diagnosis in Oral Lichen Planus: A Retrospective study

Mustafa Mohammed Abdulhussain 1, *, Huda Elias Ali 2 and Ali Sami Mohsin 3

1 Department of Oral Pathology, College of Dentistry, Mustansiriyah University, Baghdad, Iraq.
2 Department of Pedodontic, orthodontic and prevention, College of Dentistry, Mustansiriyah University, Baghdad, Iraq.
3 Department of Oral Pathology, College of Dentistry, Mustansiriyah University, Baghdad, Iraq.
* Corresponding Author
ORCID Details
Mustafa Mohammed Abdulhussain:  https://orcid.org/ 0000-0003-3794-2299
 
Research Article
World Journal of Biological and Pharmaceutical Research, 2026, 10(02), 001–007.
Article DOI: 10.53346/wjbpr.2026.10.2.0016
Publication history: 
Received on 20 July 2026; revised on 29 August 2026; accepted on 31 August 2026
 
Abstract: 
Background: Oral Lichen Planus (OLP) refers to a persistent inflammatory condition that has a tendency for cancerous transformation; therefore, its identification is primarily dependent on histological examination. But symptoms that combine with different disorders, such as oral lesions or autoimmune conditions, may make a definitive diagnosis of Oral Lichen Planus difficult to achieve. In recent years, artificial intelligence systems like GPT-4 and GPT-5 have demonstrated potential skills in patient information processing.
Objective: The purpose of this research was to interpret and contrast the predictive ability of the GPT-4 test and GPT-5 test as well as in diagnosing oral lichen planus using histopathology findings.
Materials and Methods: The retrospective diagnostic procedure has been carried out on 50 histopathologically verified samples with oral lichen planus. All of them were assessed individually by two expert oral pathologists, so only instances with perfect agreement were used as the highest level of validation. Histopathological findings were anonymously classified and modified prior to being examined using GPT-4 and GPT-5, respectively, with a single query. The results of the two models have been compared to the guideline indication. The examination efficiency, specificity, sensitivity, and Cohen's kappa coefficients were computed. The McNemar exam was executed to analyze the performances of two distinct versions.
Results: The two tests, GPT-4 and GPT-5, were highly accurate in detecting oral lichen planus, with considerable concordance regarding the gold-based standard. GPT-5 had stronger analytic reliability and fewer variations than GPT-4. Furthermore, GPT-5 enabled more thorough recognition of essential histopathological characteristics.
Conclusion: The artificial intelligence (AI) may reach 100% concordance with experienced oral reviewers in recognizing oral lichen planus. These results show the promise of sophisticated artificial intelligence ("AI") models as guidance instruments in oral pathology, notably in improving diagnostic stability and suppressing observer heterogeneity.
 
Keywords: 
Artificial Intelligence, Oral Lichen Planus, Accuracy, Sensitivity, Histopathological study.
 
Full text article in PDF: