Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (28)

Search Parameters:
Keywords = human phenotype ontology (HPO)

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
11 pages, 945 KB  
Article
Proposal of an Algorithm for the Clinical and Molecular Diagnosis of RASopathies Based on HPO Nomenclature
by Fernanda Meneses, Carlos Quintero, Juliana Lores, Eidith Gómez-Pineda, Diana Ramírez-Montaño, Estephania Candelo and Harry Pachajoa
Int. J. Mol. Sci. 2026, 27(16), 7348; https://doi.org/10.3390/ijms27167348 - 17 Aug 2026
Viewed by 192
Abstract
RASopathies are a group of genetic disorders caused by germline variants affecting the RAS/MAPK pathway. Their shared phenotypic features—craniofacial anomalies, cardiac defects, cutaneous findings, neurodevelopmental issues, and cancer predisposition—make diagnosis challenging, especially since most lack standardized clinical criteria. This study aimed to develop [...] Read more.
RASopathies are a group of genetic disorders caused by germline variants affecting the RAS/MAPK pathway. Their shared phenotypic features—craniofacial anomalies, cardiac defects, cutaneous findings, neurodevelopmental issues, and cancer predisposition—make diagnosis challenging, especially since most lack standardized clinical criteria. This study aimed to develop a practical diagnostic algorithm based on high-frequency Human Phenotype Ontology (HPO) features. Key clinical variables for each RASopathy were identified through HPO, PubMed, and GeneReviews. Only findings present in 80–99% of cases or supported by expert consensus were included. A decision-tree algorithm was constructed and preliminarily evaluated using a blinded cohort of 50 individuals with confirmed molecular diagnoses. Patients were eligible for inclusion if they met the following criteria: (1) molecularly confirmed diagnosis of a RASopathy by next-generation sequencing identifying a pathogenic or likely pathogenic variant; (2) availability of complete phenotypic records in the institutional clinical database; and (3) age at evaluation between 0 and 18 years. Patients were excluded if phenotypic data were incomplete or if molecular confirmation was absent. The algorithm integrates phenotypic patterns and genotype–phenotype correlations. Validation showed 78% accuracy (95% CI: 64.0–88.4%) for clinical diagnosis and 66% accuracy (95% CI: 51.2–78.8%) for molecular prediction. To our knowledge, this is the first HPO-based diagnostic algorithm for the clinical and molecular approach to RASopathies. It provides a structured, accessible tool to improve early recognition and guide molecular testing, particularly for the RASopathy subtypes represented in the validation cohort. Further external validation including underrepresented subtypes is required. Full article
Show Figures

Figure 1

13 pages, 745 KB  
Article
Integration of Machine Learning-Based Pathogenicity Prediction and Phenotype Matching Improves Variant Prioritization in Rare Clinical Testing
by Jiri Ruzicka, Jean-Marie Ravel, Jérôme Audoux, Alexandre Boulat, Julien Thévenon, Kévin Yauy, Marine Dancer, Laure Raymond, Yannis Lombardi, Nicolas Philippe, Michael GB Blum, Nicolas Duforet-Frebourg and Laurent Mesnard
Curr. Issues Mol. Biol. 2026, 48(7), 706; https://doi.org/10.3390/cimb48070706 - 11 Jul 2026
Viewed by 622
Abstract
Genome and exome sequencing have become central to diagnosing rare hereditary diseases, but each test returns thousands of variants that a clinical scientist must review by hand to find the one responsible for the patient’s condition. This manual interpretation is the main bottleneck [...] Read more.
Genome and exome sequencing have become central to diagnosing rare hereditary diseases, but each test returns thousands of variants that a clinical scientist must review by hand to find the one responsible for the patient’s condition. This manual interpretation is the main bottleneck in clinical genomics. To reduce it, we developed DiagAI, a machine-learning system that ranks the variants found in a patient and returns a short list of the most likely causal candidates. DiagAI combines three sources of evidence: a pathogenicity score from the Universal Pathogenicity Predictor (UP2), a model we trained to estimate how damaging a variant is on the five-tier scale of the American College of Medical Genetics and Genomics (ACMG); a phenotype-matching score from PhenoGenius, which weighs how well a gene’s known clinical features match the patient’s symptoms (encoded as Human Phenotype Ontology, or HPO, terms); and expert rules covering inheritance pattern and sequencing quality. We evaluated DiagAI on 966 exomes from adults investigated for kidney disease of unknown cause, of which 196 had a confirmed genetic diagnosis. We first tested UP2 on its own by ranking 62 confirmed disease-causing missense variants that were absent from its training data: UP2 placed the causal variant within the top 100 candidates in 87% of cases, compared with 61% for the widely used tool REVEL. Across the 196 diagnosed exomes, the full DiagAI shortlist contained the causal variant in 94.9% of cases when the patient’s symptoms were provided and in 90.8% when they were not, with a typical shortlist of about 10 variants. When symptoms were provided, the single top-ranked variant was the correct diagnosis in 74% of cases, versus 42% without symptoms, exceeding the performance of the established tools Exomiser and AI-MARRVEL on the same cohort. DiagAI produces compact, accurate shortlists that can reduce the manual interpretation workload as diagnostic sequencing volumes continue to grow. Full article
(This article belongs to the Special Issue Emerging Trends in Bioinformatics and Computational Biology)
Show Figures

Figure 1

19 pages, 14726 KB  
Article
MSeqDR PMD-VR: An Expert-Curated Virtual Registry of 11,000 Mitochondrial Disease Cases Established Through Literature Mining and Generative AI Augmentation
by Lishuang Shen, Marie T. Lott, Elizabeth M. Mccormick, Colleen C. Muraresku, Kierstin Keller, Douglas C. Wallace, Zarazuela Zolkipli-Cunningham, Shamima Rahman, Marni J. Falk and Xiaowu Gai
Genes 2026, 17(7), 757; https://doi.org/10.3390/genes17070757 - 30 Jun 2026
Viewed by 653
Abstract
Background/Objectives: Patient registries are essential for rare disease research, yet the extensive genetic and phenotypic heterogeneity of primary mitochondrial diseases (PMDs) makes traditional registry development slow and resource-intensive. We established the MSeqDR PMD virtual registry (PMD-VR) to address this gap through systematic literature [...] Read more.
Background/Objectives: Patient registries are essential for rare disease research, yet the extensive genetic and phenotypic heterogeneity of primary mitochondrial diseases (PMDs) makes traditional registry development slow and resource-intensive. We established the MSeqDR PMD virtual registry (PMD-VR) to address this gap through systematic literature mining and semi-automated data harmonization. Methods: The PMD-VR captures, standardizes, and harmonizes published case-level PMD data using a semi-automated curation pipeline. A data transformation framework maps heterogeneous raw data terms to standardized common data elements (CDEs). A generative AI (GenAI) platform leveraging large language models (LLMs), augmented by Human Phenotype Ontology (HPO) and external biomedical knowledge sources, accelerates data transformation and generates simulated clinical reports. Results: Currently, PMD-VR contains approximately 11,000 de-identified literature-derived cases, including over 2300 Leigh syndrome spectrum (LSS), 278 MELAS, and 300 CPEO cases. The pipeline mapped 872 heterogeneous terms to 102 standardized CDEs. Pathogenicity assessments were captured for variants in over 7900 cases, including 3800 with mtDNA pathogenic or likely pathogenic variants. Modes of inheritance were inferred for 5212 cases. PMD-VR has supported ClinGen Mitochondrial Diseases Gene Curation Expert Panel (Mito-GCEP) efforts, providing phenotyped evidence for 440 curated LSS cases across 113 PMD genes. Conclusions: PMD-VR is among the largest single PMD registries, offering a scalable, web-accessible platform for generating analysis-ready cohorts from the published literature. It represents a rich resource enabling comprehensive PMD characterization with unprecedented breadth of genetic and phenotypic knowledge. Full article
(This article belongs to the Special Issue Mitochondrial Genetics in Health and Disease)
Show Figures

Figure 1

20 pages, 769 KB  
Article
Note-Level Phenotyping of Multiple-Sclerosis Notes by a Large Language Model Achieves near Human-Level Agreement
by Daniel B. Hier, Pavankumar Y. Srinivasula and Michael D. Carrithers
J. Clin. Med. 2026, 15(11), 4092; https://doi.org/10.3390/jcm15114092 - 25 May 2026
Cited by 1 | Viewed by 414
Abstract
Background/Objectives: Clinical phenotyping from narrative electronic health records (EHRs) often relies on multi-stage pipelines involving span-level extraction, ontology mapping, and aggregation. Large language models (LLMs) may enable direct document-level abstraction of clinically meaningful phenotype features from complete notes. We evaluated whether GPT-5.2 [...] Read more.
Background/Objectives: Clinical phenotyping from narrative electronic health records (EHRs) often relies on multi-stage pipelines involving span-level extraction, ontology mapping, and aggregation. Large language models (LLMs) may enable direct document-level abstraction of clinically meaningful phenotype features from complete notes. We evaluated whether GPT-5.2 could approximate human annotation for note-level multiple sclerosis (MS) phenotyping and compared its performance with human annotators, a locally run open-source LLM, HPO-based extraction tools, and a supervised clinical transformer encoder. Methods: We analyzed 100 de-identified MS neurology progress notes from a single academic medical center. Each note was annotated for the presence or absence of 17 predefined neurological phenotype categories. Two human annotators independently labeled all notes using a multi-label note-level framework in Prodigy, and disagreements were adjudicated to create a reference annotation set. GPT-5.2 was evaluated in a zero-shot setting using structured JSON output. Comparator methods included Llama-3.1 8B, Doc2Hpo, ClinPhen, PhenoSnap, and BioClinical ModernBERT. Performance was assessed using agreement, precision, recall, F1, Matthews correlation coefficient, and false-positive and false-negative assignments per note. Results: Human–human agreement was generally high, although lower for rare or ambiguously documented features. GPT-5.2 achieved the strongest automated performance, with macro-precision 0.734, macro-recall 0.921, macro-F1 0.801, and macro-averaged MCC 0.777, approaching human annotator performance. GPT-5.2 showed the lowest false-negative count per note but more false-positive assignments than either human annotator, reflecting a sensitive but more inclusive annotation profile. Llama-3.1 8B performed competitively among automated methods, whereas HPO-based extraction tools and BioClinical ModernBERT showed lower performance on this low-resource note-level task. Secondary review of GPT-5.2 discordant assignments found no clear hallucinations and suggested that some apparent false positives reflected phenotype evidence missed in the human-derived reference set. Conclusions: GPT-5.2 achieved near-human performance for document-level recognition of MS phenotype categories from narrative neurology notes. Direct note-level abstraction may provide a scalable approach for research and population-health phenotyping of large EHR note corpora. Full article
Show Figures

Figure 1

27 pages, 3347 KB  
Article
Generative AI Accelerates Genotype–Phenotype Characterization of a 1600-Case Leigh Syndrome Virtual Cohort from Published Literature
by Lishuang Shen
Biology 2026, 15(4), 334; https://doi.org/10.3390/biology15040334 - 14 Feb 2026
Cited by 1 | Viewed by 1531
Abstract
Leigh Syndrome Spectrum (LSS) is a rare and heterogeneous disease continuum with most published cohorts in small sizes that limit the statistical power. Large-scale meta-analyses with published case-level clinical data extracted from the literature are essential for robust population analysis but are hindered [...] Read more.
Leigh Syndrome Spectrum (LSS) is a rare and heterogeneous disease continuum with most published cohorts in small sizes that limit the statistical power. Large-scale meta-analyses with published case-level clinical data extracted from the literature are essential for robust population analysis but are hindered by the burden of manually standardizing the unstructured, heterogeneous, and sparse case-level data from the literature. We developed a novel workflow which is among the first to combine Generative AI (GenAI) with human-in-the-loop curation to overcome this barrier. This pipeline utilized Google’s Gemini-2.5-pro and rapidly processed over 2300 cases from published case data tables in two weeks and achieved >90% accuracy in mapping raw clinical data to Human Phenotype Ontology (HPO) terms. This process rapidly yielded a harmonized LSS virtual cohort of 1679 data-rich cases, which is the largest LSS virtual cohort reported so far, and thus enables characterization of LSS phenotypic and genetic architectures, revealing that autosomal recessive (932 cases) and mitochondrial (752 cases) inheritance are the most common. The most frequently mutated genes were SURF1 (240 cases), MT-ATP6 (199), and MT-ND3 (183). HPO term consolidation identified common hallmark phenotypes, including lactic acidosis, hypotonia, bilateral basal ganglia lesions, and mitochondrial respiratory chain deficiency. The cohort’s scale enabled large-scale survival analysis, revealing that defects in mitochondrial translation are associated with the poorest prognosis (84% mortality in this group) and early onset (0.23 years). Among the deceased group, patients with Complex V mutations were linked to a significantly shorter mean survival time (1.77 years) than those with Complex I (3.70 years) or IV (3.57 years) mutations. This GenAI-driven methodology establishes a scalable framework for rapidly creating analysis-ready virtual cohorts from heterogeneous literature and accelerating population-level study for rare diseases including Leigh Syndrome and other mitochondrial diseases. Full article
(This article belongs to the Section Bioinformatics)
Show Figures

Figure 1

24 pages, 6717 KB  
Review
Dissecting the Genetic Contribution of Tooth Agenesis
by Antonio Fallea, Mirella Vinci, Simona L’Episcopo, Massimiliano Bartolone, Antonino Musumeci, Alda Ragalmuto, Simone Treccarichi and Francesco Calì
Int. J. Mol. Sci. 2025, 26(21), 10485; https://doi.org/10.3390/ijms262110485 - 28 Oct 2025
Cited by 5 | Viewed by 5564
Abstract
Tooth agenesis (TA), the congenital absence of one or more teeth, is the most common manifestation of defective dental morphogenesis in humans. TA can occur as an isolated (non-syndromic) condition or as part of a broader syndromic presentation. In this review, we analyzed [...] Read more.
Tooth agenesis (TA), the congenital absence of one or more teeth, is the most common manifestation of defective dental morphogenesis in humans. TA can occur as an isolated (non-syndromic) condition or as part of a broader syndromic presentation. In this review, we analyzed a total of 73 manuscripts to provide a comprehensive update on the genetic landscape of TA. To investigate the genes, variants, and associated phenotypes, we reviewed data from curated databases including Human Phenotype Ontology (HPO), OMIM, ClinVar and MalaCards. Based on the current evidence, the genes most frequently implicated in TA are MSX1, EDA, and PAX9. However, chromosomal abnormalities, such as those seen in Down syndrome and Williams syndrome, along with structural variations (e.g., deletions and duplications), also contribute significantly to TA etiology. The most involved pathways include TNF receptor binding, encompassing genes such as EDA, EDA2R, EDAR, and EDARADD, and the mTOR signaling pathway, which includes AXIN2, FGFR1, LRP6, WNT10A, and WNT10B. The aim of this review is to provide an critical synthesis of the genetic mechanisms underlying TA, highlighting the contribution of major signaling pathways, regulatory networks, and emerging molecular insights that may inform diagnostic and therapeutic advances. Full article
(This article belongs to the Section Molecular Genetics and Genomics)
Show Figures

Figure 1

19 pages, 1561 KB  
Article
Integrating Genomics and Deep Phenotyping for Diagnosing Rare Pediatric Neurological Diseases: Potential for Sustainable Healthcare in Resource-Limited Settings
by Nigara Yerkhojayeva, Nazira Zharkinbekova, Sovet Azhayev, Ainash Oshibayeva, Gulnaz Nuskabayeva and Rauan Kaiyrzhanov
Int. J. Transl. Med. 2025, 5(4), 47; https://doi.org/10.3390/ijtm5040047 - 4 Oct 2025
Cited by 1 | Viewed by 2780
Abstract
Background: Rare pediatric neurological diseases (RPND) often remain undiagnosed for years, creating prolonged and costly diagnostic odysseys. Combining Human Phenotype Ontology (HPO)-based deep phenotyping with exome sequencing (ES) and reverse phenotyping offers the potential to improve diagnostic yield, accelerate diagnosis, and support sustainable [...] Read more.
Background: Rare pediatric neurological diseases (RPND) often remain undiagnosed for years, creating prolonged and costly diagnostic odysseys. Combining Human Phenotype Ontology (HPO)-based deep phenotyping with exome sequencing (ES) and reverse phenotyping offers the potential to improve diagnostic yield, accelerate diagnosis, and support sustainable healthcare in resource-limited settings. Objectives: To evaluate the diagnostic yield and clinical impact of an integrated approach combining deep phenotyping, ES, and reverse phenotyping in children with suspected RPNDs. Methods: In this multicenter observational study, eighty-one children from eleven hospitals in South Kazakhstan were recruited via the Central Asian and Transcaucasian Rare Pediatric Neurological Diseases Consortium. All patients underwent standardized HPO-based phenotyping and ES, with variant interpretation following ACMG guidelines. Reverse phenotyping and interdisciplinary discussions were used to refine clinical interpretation. Results: A molecular diagnosis was established in 34 of 81 patients (42%) based on pathogenic or likely pathogenic variants. Variants of uncertain significance (VUS) were identified in an additional 9 patients (11%), but were reported separately and not included in the diagnostic yield. Reverse phenotyping clarified or expanded clinical features in one-third of genetically diagnosed cases and provided supportive evidence in most VUS cases, although their classification remained unchanged. Conclusions: Integrating deep phenotyping, ES, and reverse phenotyping substantially improved diagnostic outcomes and shortened the diagnostic odyssey. This model reduces unnecessary procedures, minimizes delays, and provides a scalable framework for advancing equitable access to genomic diagnostics in resource-constrained healthcare systems. Full article
Show Figures

Figure 1

15 pages, 3574 KB  
Article
Prior Knowledge Shapes Success When Large Language Models Are Fine-Tuned for Biomedical Term Normalization
by Daniel B. Hier, Steven K. Platt and Anh Nguyen
Information 2025, 16(9), 776; https://doi.org/10.3390/info16090776 - 7 Sep 2025
Cited by 2 | Viewed by 2227
Abstract
Large language models (LLMs) often fail to correctly associate biomedical terms with their standardized ontology identifiers, posing challenges for downstream applications that rely on accurate, machine-readable codes. These linking failures can compromise the integrity of data used in precision medicine, clinical decision support, [...] Read more.
Large language models (LLMs) often fail to correctly associate biomedical terms with their standardized ontology identifiers, posing challenges for downstream applications that rely on accurate, machine-readable codes. These linking failures can compromise the integrity of data used in precision medicine, clinical decision support, and population health. Fine-tuning can partially remedy these issues, but the degree of improvement varies across terms and terminologies. Focusing on the Human Phenotype Ontology (HPO), we show that a model’s prior knowledge of term–identifier pairs, acquired during pre-training, strongly predicts whether fine-tuning will enhance its linking accuracy. We evaluate prior knowledge in three complementary ways: (1) latent probabilistic knowledge, revealed through stochastic prompting, captures hidden associations not evident in deterministic output; (2) partial subtoken knowledge, reflected in incomplete but non-random generation of identifier components; and (3) term familiarity, inferred from annotation frequencies in the biomedical literature, which serve as a proxy for training exposure. We then assess how these forms of prior knowledge influence the accuracy of deterministic identifier linking. Fine-tuning performance varies most for terms in what we call the reactive middle zone of the ontology—terms with intermediate levels of prior knowledge that are neither absent nor fully consolidated. Fine-tuning was most successful when prior knowledge as measured by partial subtoken knowledge, was ‘weak’ or ‘medium’ or when prior knowledge as measured by latent probabilistic knowledge was ‘unknown’ or ‘weak’ (p<0.001). These terms from the ‘reactive middle’ exhibited the largest gains or losses in accuracy during fine-tuning, suggesting that the success of knowledge injection critically depends on the level of term–identifier pair knowledge in the LLM before fine-tuning. Full article
Show Figures

Figure 1

19 pages, 272 KB  
Review
Artificial Intelligence in the Diagnosis of Pediatric Rare Diseases: From Real-World Data Toward a Personalized Medicine Approach
by Nikola Ilić and Adrijan Sarajlija
J. Pers. Med. 2025, 15(9), 407; https://doi.org/10.3390/jpm15090407 - 1 Sep 2025
Cited by 9 | Viewed by 4489
Abstract
Background: Artificial intelligence (AI) is increasingly applied in the diagnosis of pediatric rare diseases, enhancing the speed, accuracy, and accessibility of genetic interpretation. These advances support the ongoing shift toward personalized medicine in clinical genetics. Objective: This review examines current applications of AI [...] Read more.
Background: Artificial intelligence (AI) is increasingly applied in the diagnosis of pediatric rare diseases, enhancing the speed, accuracy, and accessibility of genetic interpretation. These advances support the ongoing shift toward personalized medicine in clinical genetics. Objective: This review examines current applications of AI in pediatric rare disease diagnostics, with a particular focus on real-world data integration and implications for individualized care. Methods: A narrative review was conducted covering AI tools for variant prioritization, phenotype–genotype correlations, large language models (LLMs), and ethical considerations. The literature was identified through PubMed, Scopus, and Web of Science up to July 2025, with priority given to studies published in the last seven years. Results: AI platforms provide support for genomic interpretation, particularly within structured diagnostic workflows. Tools integrating Human Phenotype Ontology (HPO)-based inputs and LLMs facilitate phenotype matching and enable reverse phenotyping. The use of real-world data enhances the applicability of AI in complex and heterogeneous clinical scenarios. However, major challenges persist, including data standardization, model interpretability, workflow integration, and algorithmic bias. Conclusions: AI has the potential to advance earlier and more personalized diagnostics for children with rare diseases. Achieving this requires multidisciplinary collaboration and careful attention to clinical, technical, and ethical considerations. Full article
16 pages, 1534 KB  
Article
Clinician-Based Functional Scoring and Genomic Insights for Prognostic Stratification in Wolf–Hirschhorn Syndrome
by Julián Nevado, Raquel Blanco-Lago, Cristina Bel-Fenellós, Adolfo Hernández, María A. Mori-Álvarez, Chantal Biencinto-López, Ignacio Málaga, Harry Pachajoa, Elena Mansilla, Fe A. García-Santiago, Pilar Barrúz, Jair A. Tenorio-Castaño, Yolanda Muñoz-GªPorrero, Isabel Vallcorba and Pablo Lapunzina
Genes 2025, 16(7), 820; https://doi.org/10.3390/genes16070820 - 12 Jul 2025
Cited by 1 | Viewed by 1802
Abstract
Background/Objectives: Wolf–Hirschhorn syndrome (WHS; OMIM #194190) is a rare neurodevelopmental disorder, caused by deletions in the distal short arm of chromosome 4. It is characterized by developmental delay, epilepsy, intellectual disability, and distinctive facial dysmorphism. Clinical presentation varies widely, complicating prognosis and [...] Read more.
Background/Objectives: Wolf–Hirschhorn syndrome (WHS; OMIM #194190) is a rare neurodevelopmental disorder, caused by deletions in the distal short arm of chromosome 4. It is characterized by developmental delay, epilepsy, intellectual disability, and distinctive facial dysmorphism. Clinical presentation varies widely, complicating prognosis and individualized care. Methods: We assembled a cohort of 140 individuals with genetically confirmed WHS from Spain and Latin-America, and developed and validated a multidimensional, Clinician-Reported Outcome Assessment (ClinRO) based on the Global Functional Assessment of the Patient (GFAP), derived from standardized clinical questionnaires and weighted by HPO (Human Phenotype Ontology) term frequencies. The GFAP score quantitatively captures key functional domains in WHS, including neurodevelopment, epilepsy, comorbidities, and age-corrected developmental milestones (selected based on clinical experience and disease burden). Results: Higher GFAP scores are associated with worse clinical outcomes. GFAP showed strong correlations with deletion size, presence of additional genomic rearrangements, sex, and epilepsy severity. Ward’s clustering and discriminant analyses confirmed GFAP’s discriminative power, classifying over 90% of patients into clinically meaningful groups with different prognoses. Conclusions: Our findings support GFAP as a robust, WHS-specific ClinRO that may aid in stratification, prognosis, and clinical management. This tool may also serve future interventional studies as a standardized outcome measure. Beyond its clinical utility, GFAP also revealed substantial social implications. This underscores the broader socioeconomic burden of WHS and the potential value of GFAP in identifying high-support families that may benefit from targeted resources and services. Full article
(This article belongs to the Special Issue Molecular Basis of Rare Genetic Diseases)
Show Figures

Figure 1

14 pages, 1324 KB  
Article
Preprocessing of Physician Notes by LLMs Improves Clinical Concept Extraction Without Information Loss
by Daniel B. Hier, Michael A. Carrithers, Steven K. Platt, Anh Nguyen, Ioannis Giannopoulos and Tayo Obafemi-Ajayi
Information 2025, 16(6), 446; https://doi.org/10.3390/info16060446 - 27 May 2025
Cited by 7 | Viewed by 4609
Abstract
Clinician notes are a rich source of patient information, but often contain inconsistencies due to varied writing styles, abbreviations, medical jargon, grammatical errors, and non-standard formatting. These inconsistencies hinder their direct use in patient care and degrade the performance of downstream computational applications [...] Read more.
Clinician notes are a rich source of patient information, but often contain inconsistencies due to varied writing styles, abbreviations, medical jargon, grammatical errors, and non-standard formatting. These inconsistencies hinder their direct use in patient care and degrade the performance of downstream computational applications that rely on these notes as input, such as quality improvement, population health analytics, precision medicine, clinical decision support, and research. We present a large-language-model (LLM) approach to the preprocessing of 1618 neurology notes. The LLM corrected spelling and grammatical errors, expanded acronyms, and standardized terminology and formatting, without altering clinical content. Expert review of randomly sampled notes confirmed that no significant information was lost. To evaluate downstream impact, we applied an ontology-based NLP pipeline (Doc2Hpo) to extract biomedical concepts from the notes before and after editing. F1 scores for Human Phenotype Ontology extraction improved from 0.40 to 0.61, confirming our hypothesis that better inputs yielded better outputs. We conclude that LLM-based preprocessing is an effective error correction strategy that improves data quality at the level of free text in clinical notes. This approach may enhance the performance of a broad class of downstream applications that derive their input from unstructured clinical documentation. Full article
(This article belongs to the Special Issue Biomedical Natural Language Processing and Text Mining)
Show Figures

Figure 1

11 pages, 3831 KB  
Brief Report
Expanding the Clinical Spectrum Associated with the Recurrent Arg203Trp Variant in PACS1: An Italian Cohort Study
by Stefano Pagano, Diego Lopergolo, Alessandro De Falco, Camilla Meossi, Sara Satolli, Rosa Pasquariello, Rosanna Trovato, Alessandra Tessa, Claudia Casalini, Lucia Pfanner, Guja Astrea, Roberta Battini and Filippo M. Santorelli
Genes 2025, 16(2), 227; https://doi.org/10.3390/genes16020227 - 16 Feb 2025
Cited by 5 | Viewed by 1904
Abstract
Background/Objectives: Schuurs–Hoeijmakers syndrome (SHMS), also known as PACS1 neurodevelopmental disorder, is a rare condition characterized by intellectual disability, distinctive craniofacial abnormalities, and congenital malformations. SHMS has already been associated with variants in the PACS1 gene in 63 patients. In this study, we [...] Read more.
Background/Objectives: Schuurs–Hoeijmakers syndrome (SHMS), also known as PACS1 neurodevelopmental disorder, is a rare condition characterized by intellectual disability, distinctive craniofacial abnormalities, and congenital malformations. SHMS has already been associated with variants in the PACS1 gene in 63 patients. In this study, we describe 10 new Italian SHMS patients all harboring the common de novo p.(Arg203Trp) variant. Methods: The 10 patients we studied were evaluated by clinical geneticists and child neurologists and a detailed description of clinical features was recorded. Data were then coded using the Human Phenotype Ontology (HPO) terms. The recurrent p.(Arg203Trp) variant in PACS1 was identified by clinical exome sequencing or whole exome sequencing in trio using standard methodologies. To facilitate mutation identification, we designed a new PCR-RFLP strategy adopting the endonuclease DpnII. Results: We define a detailed clinical phenotyping in patients with intellectual disability and facial characteristics (thick eyebrows, down-slanting palpebral fissures, ocular hypertelorism, low-set ears, a thin upper lip, and a wide mouth) that can help clinicians form a more efficient diagnosis of SHMS even through neuroimaging and neuropsychological evaluation. Conclusions: Our series of 10 newly affected Italian children highlights specific clinical features that may help clinicians recognize and better manage this syndrome, contributing to precision medicine approaches in medical genetics. Full article
(This article belongs to the Section Genetic Diagnosis)
Show Figures

Figure 1

26 pages, 6531 KB  
Article
Analysis of Regions of Homozygosity: Revisited Through New Bioinformatic Approaches
by Susana Valente, Mariana Ribeiro, Jennifer Schnur, Filipe Alves, Nuno Moniz, Dominik Seelow, João Parente Freixo, Paulo Filipe Silva and Jorge Oliveira
BioMedInformatics 2024, 4(4), 2374-2399; https://doi.org/10.3390/biomedinformatics4040128 - 16 Dec 2024
Cited by 4 | Viewed by 6411
Abstract
Background: Runs of homozygosity (ROHs), continuous homozygous regions across the genome, are often linked to consanguinity, with their size and frequency reflecting shared parental ancestry. Homozygosity mapping (HM) leverages ROHs to identify genes associated with autosomal recessive diseases. Whole-exome sequencing (WES) improves [...] Read more.
Background: Runs of homozygosity (ROHs), continuous homozygous regions across the genome, are often linked to consanguinity, with their size and frequency reflecting shared parental ancestry. Homozygosity mapping (HM) leverages ROHs to identify genes associated with autosomal recessive diseases. Whole-exome sequencing (WES) improves HM by detecting ROHs and disease-causing variants. Methods: To streamline personalized multigene panel creation, using WES and ROHs, we developed a methodology integrating ROHMMCLI and HomozygosityMapper algorithms, and, optionally, Human Phenotype Ontology (HPO) terms, implemented in a Django Web application. Resorting to a dataset of 12,167 WES, we performed the first ROH profiling of the Portuguese population. Clustering models were applied to predict consanguinity from ROH features. Results: These resources were applied for the genetic characterization of two siblings with epilepsy, myoclonus and dystonia, pinpointing the CSTB gene as disease-causing. Using the 2021 Census population distribution, we created a representative sample (3941 WES) and measured genome-wide autozygosity (FROH). Portalegre, Viseu, Bragança, Madeira, and Vila Real districts presented the highest FROH scores. Multidimensional scaling showed that ROH count and sum were key predictors of consanguinity, achieving a test F1-score of 0.96 with additional features. Conclusions: This study contributes with new bioinformatics tools for ROH analysis in a clinical setting, providing unprecedented population-level ROH data for Portugal. Full article
Show Figures

Figure 1

14 pages, 2615 KB  
Perspective
Rare Genetic Developmental Disabilities: Mabry Syndrome (MIM 239300) Index Cases and Glycophosphatidylinositol (GPI) Disorders
by Miles D. Thompson and Alexej Knaus
Genes 2024, 15(5), 619; https://doi.org/10.3390/genes15050619 - 14 May 2024
Cited by 6 | Viewed by 3582
Abstract
The case report by Mabry et al. (1970) of a family with four children with elevated tissue non-specific alkaline phosphatase, seizures and profound developmental disability, became the basis for phenotyping children with the features that became known as Mabry syndrome. Aside from improvements [...] Read more.
The case report by Mabry et al. (1970) of a family with four children with elevated tissue non-specific alkaline phosphatase, seizures and profound developmental disability, became the basis for phenotyping children with the features that became known as Mabry syndrome. Aside from improvements in the services available to patients and families, however, the diagnosis and treatment of this, and many other developmental disabilities, did not change significantly until the advent of massively parallel sequencing. As more patients with features of the Mabry syndrome were identified, exome and genome sequencing were used to identify the glycophosphatidylinositol (GPI) biosynthesis disorders (GPIBDs) as a group of congenital disorders of glycosylation (CDG). Biallelic variants of the phosphatidylinositol glycan (PIG) biosynthesis, type V (PIGV) gene identified in Mabry syndrome became evidence of the first in a phenotypic series that is numbered HPMRS1-6 in the order of discovery. HPMRS1 [MIM: 239300] is the phenotype resulting from inheritance of biallelic PIGV variants. Similarly, HPMRS2 (MIM 614749), HPMRS5 (MIM 616025) and HPMRS6 (MIM 616809) result from disruption of the PIGO, PIGW and PIGY genes expressed in the endoplasmic reticulum. By contrast, HPMRS3 (MIM 614207) and HPMRS4 (MIM 615716) result from disruption of post attachment to proteins PGAP2 (HPMRS3) and PGAP3 (HPMRS4). The GPI biosynthesis disorders (GPIBDs) are currently numbered GPIBD1-21. Working with Dr. Mabry, in 2020, we were able to use improved laboratory diagnostics to complete the molecular diagnosis of patients he had originally described in 1970. We identified biallelic variants of the PGAP2 gene in the first reported HPMRS patients. We discuss the longevity of the Mabry syndrome index patients in the context of the utility of pyridoxine treatment of seizures and evidence for putative glycolipid storage in patients with HPMRS3. From the perspective of the laboratory innovations made that enabled the identification of the HPMRS phenotype in Dr. Mabry’s patients, the need for treatment innovations that will benefit patients and families affected by developmental disabilities is clear. Full article
Show Figures

Figure 1

21 pages, 1059 KB  
Review
Advancing the Management of Long COVID by Integrating into Health Informatics Domain: Current and Future Perspectives
by Radha Ambalavanan, R Sterling Snead, Julia Marczika, Karina Kozinsky and Edris Aman
Int. J. Environ. Res. Public Health 2023, 20(19), 6836; https://doi.org/10.3390/ijerph20196836 - 26 Sep 2023
Cited by 11 | Viewed by 4304
Abstract
The ongoing COVID-19 pandemic has profoundly affected millions of lives globally, with some individuals experiencing persistent symptoms even after recovering. Understanding and managing the long-term sequelae of COVID-19 is crucial for research, prevention, and control. To effectively monitor the health of those affected, [...] Read more.
The ongoing COVID-19 pandemic has profoundly affected millions of lives globally, with some individuals experiencing persistent symptoms even after recovering. Understanding and managing the long-term sequelae of COVID-19 is crucial for research, prevention, and control. To effectively monitor the health of those affected, maintaining up-to-date health records is essential, and digital health informatics apps for surveillance play a pivotal role. In this review, we overview the existing literature on identifying and characterizing long COVID manifestations through hierarchical classification based on Human Phenotype Ontology (HPO). We outline the aspects of the National COVID Cohort Collaborative (N3C) and Researching COVID to Enhance Recovery (RECOVER) initiative in artificial intelligence (AI) to identify long COVID. Through knowledge exploration, we present a concept map of clinical pathways for long COVID, which offers insights into the data required and explores innovative frameworks for health informatics apps for tackling the long-term effects of COVID-19. This study achieves two main objectives by comprehensively reviewing long COVID identification and characterization techniques, making it the first paper to explore incorporating long COVID as a variable risk factor within a digital health informatics application. By achieving these objectives, it provides valuable insights on long COVID’s challenges and impact on public health. Full article
(This article belongs to the Section Health Communication and Informatics)
Show Figures

Figure 1

Back to TopTop