Next Article in Journal
Ophthalmic Artery Doppler Indices at 11–13 Weeks of Gestation in Relation to Early and Late Preeclampsia
Next Article in Special Issue
Artificial Intelligence in Ophthalmology: Acceptance, Clinical Integration, and Educational Needs in Switzerland
Previous Article in Journal
Can Adjunctive Lithium Therapy Influence Emotional Dysregulation in Adolescents? Findings from a Retrospective Study
Previous Article in Special Issue
Inter-Relationships Between the Deep Learning-Based Pachychoroid Index and Clinical Features Associated with Neovascular Age-Related Macular Degeneration
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Comparison of Validity and Reliability of Manual Consensus Grading vs. Automated AI Grading for Diabetic Retinopathy Screening in Oslo, Norway: A Cross-Sectional Pilot Study

by
Mia Karabeg
1,2,
Goran Petrovski
1,2,3,4,
Katrine Holen
2,
Ellen Steffenssen Sauesund
2,
Dag Sigurd Fosmark
2,
Greg Russell
5,
Maja Gran Erke
2,
Vallo Volke
6,
Vidas Raudonis
7,
Rasa Verkauskiene
8,
Jelizaveta Sokolovska
9,
Morten Carstens Moe
1,2,
Inga-Britt Kjellevold Haugen
10 and
Beata Eva Petrovski
1,2,8,*
1
Center for Eye Research and Innovative Diagnostics, Department of Ophthalmology, Institute for Clinical Medicine, University of Oslo, Kirkeveien 166, 0450 Oslo, Norway
2
Department of Ophthalmology, Oslo University Hospital, Kirkeveien 166, 0450 Oslo, Norway
3
Department of Ophthalmology, University Hospital Centre, University of Split School of Medicine, 21000 Split, Croatia
4
UKLONetwork, University St. Kliment Ohridski-Bitola, 7000 Bitola, North Macedonia
5
Clinical Development, Eyenuk Inc., Woodland Hills, CA 91367, USA
6
Faculty of Medicine, Tartu University, 50411 Tartu, Estonia
7
Automation Department, Kaunas University of Technology, 51368 Kaunas, Lithuania
8
Institute of Endocrinology, Lithuanian University of Health Sciences, 50161 Kaunas, Lithuania
9
Faculty of Medicine, University of Latvia, Jelgavas Street 3, LV1004 Riga, Latvia
10
Norwegian Association of the Blind and Partially Sighted, 0354 Oslo, Norway
*
Author to whom correspondence should be addressed.
J. Clin. Med. 2025, 14(13), 4810; https://doi.org/10.3390/jcm14134810
Submission received: 17 April 2025 / Revised: 25 June 2025 / Accepted: 26 June 2025 / Published: 7 July 2025
(This article belongs to the Special Issue Artificial Intelligence and Eye Disease)

Abstract

Background: Diabetic retinopathy (DR) is a leading cause of visual impairment worldwide. Manual grading of fundus images is the gold standard in DR screening, although it is time-consuming. Artificial intelligence (AI)-based algorithms offer a faster alternative, though concerns remain about their diagnostic reliability. Methods: A cross-sectional pilot study among patients (≥18 years) with diabetes was established for DR and diabetic macular edema (DME) screening at the Oslo University Hospital (OUH), Department of Ophthalmology, and the Norwegian Association of the Blind and Partially Sighted (NABP). The aim of the study was to evaluate the validity (accuracy, sensitivity, specificity) and reliability (inter-rater agreement) of automated AI-based compared to manual consensus (MC) grading of DR and DME, performed by a multidisciplinary team of healthcare professionals. Grading of DR and DME was performed manually and by EyeArt (Eyenuk) software version v2.1.0, based on the International Clinical Disease Severity Scale (ICDR) for DR. Agreement was measured by Quadratic Weighted Kappa (QWK) and Cohen’s Kappa (κ). Sensitivity, specificity, and diagnostic test accuracy (Area Under the Curve (AUC)) were also calculated. Results: A total of 128 individuals (247 eyes) (51 women, 77 men) were included, with a median age of 52.5 years. Prevalence of any vs. referable DR (RDR) was 20.2% vs. 11.7%, while sensitivity was 94.0% vs. 89.7%, specificity was 72.6% was 83.0%, and AUC was 83.5% vs. 86.3%, respectively. DME was detected only in one eye by both methods. Conclusions: AI-based grading offered high sensitivity and acceptable specificity for detecting DR, showing moderate agreement with manual assessments. Such grading may serve as an effective screening tool to support clinical evaluation, while ongoing training of human graders remains essential to ensure high-quality reference standards for accurate diagnostic accuracy and the development of AI algorithms.
Keywords: diabetic retinopathy; artificial intelligence (AI); automated grading; EyeArt; diabetic macular edema; fundus photography; screening program; manual consensus grading; diagnostic accuracy diabetic retinopathy; artificial intelligence (AI); automated grading; EyeArt; diabetic macular edema; fundus photography; screening program; manual consensus grading; diagnostic accuracy

Share and Cite

MDPI and ACS Style

Karabeg, M.; Petrovski, G.; Holen, K.; Steffenssen Sauesund, E.; Fosmark, D.S.; Russell, G.; Erke, M.G.; Volke, V.; Raudonis, V.; Verkauskiene, R.; et al. Comparison of Validity and Reliability of Manual Consensus Grading vs. Automated AI Grading for Diabetic Retinopathy Screening in Oslo, Norway: A Cross-Sectional Pilot Study. J. Clin. Med. 2025, 14, 4810. https://doi.org/10.3390/jcm14134810

AMA Style

Karabeg M, Petrovski G, Holen K, Steffenssen Sauesund E, Fosmark DS, Russell G, Erke MG, Volke V, Raudonis V, Verkauskiene R, et al. Comparison of Validity and Reliability of Manual Consensus Grading vs. Automated AI Grading for Diabetic Retinopathy Screening in Oslo, Norway: A Cross-Sectional Pilot Study. Journal of Clinical Medicine. 2025; 14(13):4810. https://doi.org/10.3390/jcm14134810

Chicago/Turabian Style

Karabeg, Mia, Goran Petrovski, Katrine Holen, Ellen Steffenssen Sauesund, Dag Sigurd Fosmark, Greg Russell, Maja Gran Erke, Vallo Volke, Vidas Raudonis, Rasa Verkauskiene, and et al. 2025. "Comparison of Validity and Reliability of Manual Consensus Grading vs. Automated AI Grading for Diabetic Retinopathy Screening in Oslo, Norway: A Cross-Sectional Pilot Study" Journal of Clinical Medicine 14, no. 13: 4810. https://doi.org/10.3390/jcm14134810

APA Style

Karabeg, M., Petrovski, G., Holen, K., Steffenssen Sauesund, E., Fosmark, D. S., Russell, G., Erke, M. G., Volke, V., Raudonis, V., Verkauskiene, R., Sokolovska, J., Moe, M. C., Kjellevold Haugen, I.-B., & Petrovski, B. E. (2025). Comparison of Validity and Reliability of Manual Consensus Grading vs. Automated AI Grading for Diabetic Retinopathy Screening in Oslo, Norway: A Cross-Sectional Pilot Study. Journal of Clinical Medicine, 14(13), 4810. https://doi.org/10.3390/jcm14134810

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop