Next Article in Journal
Assessment of Serum Neopterin as a Biomarker in Peripheral Artery Disease
Next Article in Special Issue
Improving Skin Cancer Classification Using Heavy-Tailed Student T-Distribution in Generative Adversarial Networks (TED-GAN)
Previous Article in Journal
Before and after Endovascular Aortic Repair in the Same Patients with Aortic Dissection: A Cohort Study of Four-Dimensional Phase-Contrast Magnetic Resonance Imaging
Previous Article in Special Issue
Using Transfer Learning Method to Develop an Artificial Intelligence Assisted Triaging for Endotracheal Tube Position on Chest X-ray
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Mining Primary Care Electronic Health Records for Automatic Disease Phenotyping: A Transparent Machine Learning Framework

1
Swansea University Medical School, Swansea University, Swansea SA2 8PP, UK
2
Arthritis Research UK CREATE Centre, Division Infection and Immunity, Cardiff University, Cardiff CF10 3NB, UK
3
Welsh Arthritis Research Network, School of Medicine, Cardiff University, Cardiff CF10 3NB, UK
4
China-ASEAN Research Institute, Guangxi University, Nanning 530004, China
5
Centre for Health Technology, Faculty of Health, University of Plymouth, Plymouth PL4 8AA, UK
*
Author to whom correspondence should be addressed.
Joint lead authors.
Diagnostics 2021, 11(10), 1908; https://doi.org/10.3390/diagnostics11101908
Submission received: 16 September 2021 / Revised: 6 October 2021 / Accepted: 13 October 2021 / Published: 15 October 2021
(This article belongs to the Special Issue Intelligent Data Analysis for Medical Diagnosis)

Abstract

(1) Background: We aimed to develop a transparent machine-learning (ML) framework to automatically identify patients with a condition from electronic health records (EHRs) via a parsimonious set of features. (2) Methods: We linked multiple sources of EHRs, including 917,496,869 primary care records and 40,656,805 secondary care records and 694,954 records from specialist surgeries between 2002 and 2012, to generate a unique dataset. Then, we treated patient identification as a problem of text classification and proposed a transparent disease-phenotyping framework. This framework comprises a generation of patient representation, feature selection, and optimal phenotyping algorithm development to tackle the imbalanced nature of the data. This framework was extensively evaluated by identifying rheumatoid arthritis (RA) and ankylosing spondylitis (AS). (3) Results: Being applied to the linked dataset of 9657 patients with 1484 cases of rheumatoid arthritis (RA) and 204 cases of ankylosing spondylitis (AS), this framework achieved accuracy and positive predictive values of 86.19% and 88.46%, respectively, for RA and 99.23% and 97.75% for AS, comparable with expert knowledge-driven methods. (4) Conclusions: This framework could potentially be used as an efficient tool for identifying patients with a condition of interest from EHRs, helping clinicians in clinical decision-support process.
Keywords: phenotyping; rheumatology; cohort identification; electronic health records; feature selection; transparent machine learning; text mining; big data; artificial intelligence phenotyping; rheumatology; cohort identification; electronic health records; feature selection; transparent machine learning; text mining; big data; artificial intelligence

Share and Cite

MDPI and ACS Style

Fernández-Gutiérrez, F.; Kennedy, J.I.; Cooksey, R.; Atkinson, M.; Choy, E.; Brophy, S.; Huo, L.; Zhou, S.-M. Mining Primary Care Electronic Health Records for Automatic Disease Phenotyping: A Transparent Machine Learning Framework. Diagnostics 2021, 11, 1908. https://doi.org/10.3390/diagnostics11101908

AMA Style

Fernández-Gutiérrez F, Kennedy JI, Cooksey R, Atkinson M, Choy E, Brophy S, Huo L, Zhou S-M. Mining Primary Care Electronic Health Records for Automatic Disease Phenotyping: A Transparent Machine Learning Framework. Diagnostics. 2021; 11(10):1908. https://doi.org/10.3390/diagnostics11101908

Chicago/Turabian Style

Fernández-Gutiérrez, Fabiola, Jonathan I. Kennedy, Roxanne Cooksey, Mark Atkinson, Ernest Choy, Sinead Brophy, Lin Huo, and Shang-Ming Zhou. 2021. "Mining Primary Care Electronic Health Records for Automatic Disease Phenotyping: A Transparent Machine Learning Framework" Diagnostics 11, no. 10: 1908. https://doi.org/10.3390/diagnostics11101908

APA Style

Fernández-Gutiérrez, F., Kennedy, J. I., Cooksey, R., Atkinson, M., Choy, E., Brophy, S., Huo, L., & Zhou, S.-M. (2021). Mining Primary Care Electronic Health Records for Automatic Disease Phenotyping: A Transparent Machine Learning Framework. Diagnostics, 11(10), 1908. https://doi.org/10.3390/diagnostics11101908

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop