Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,555)

Search Parameters:
Keywords = data repositories

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
17 pages, 594 KB  
Systematic Review
Quality of Outcome Reporting for Older Subgroups in FDA Registration Gastroesophageal Cancer Trials (2010–2024): A Systematic Review
by Ilse Trip, Martha Kormi, Elizabeth Smyth, Russell D. Petty and Mark A. Baxter
Cancers 2026, 18(17), 2747; https://doi.org/10.3390/cancers18172747 - 24 Aug 2026
Abstract
Background: Gastroesophageal adenocarcinoma is predominantly a disease of older adults. However, older adults are underrepresented in randomised controlled trials (RCTs), and outcomes specific to this subgroup are underreported. Clinicians therefore extrapolate data from younger, fitter trial populations, which can lead to over- [...] Read more.
Background: Gastroesophageal adenocarcinoma is predominantly a disease of older adults. However, older adults are underrepresented in randomised controlled trials (RCTs), and outcomes specific to this subgroup are underreported. Clinicians therefore extrapolate data from younger, fitter trial populations, which can lead to over- and undertreatment of older cancer patients. This study examines the inclusion and completeness of outcome reporting for older patients in practice-changing trials in gastroesophageal cancer. Methods: RCTs were identified from United States Food and Drug Administration (accelerated) approvals between 2010 and 2024. The primary study report, clinicaltrial.gov repository and all relevant secondary publications were assessed for the level of reporting of outcomes. Efficacy outcomes were reported as complete, partial, qualitative or quantitative based on availability of sample size, effect size and measures of precision. Baseline characteristics, toxicity outcomes and health-related quality of life (HRQOL) outcomes were assessed using minimum data thresholds. Results: In total, 13 trials, 53 publications and 13 repository webpages were assessed. A total of 9320 patients were included, of which 40.1% were classified as older adults. Among the patients, 42.2% had an Eastern Cooperative Oncology Group (ECOG) Performance Status (PS) of 0, with all but one trial excluding patients with an ECOG PS of 2 (n = 63). Baseline characteristics and toxicity outcomes for older adults were fully reported in only two trials (15.4%). Reporting of HRQOL outcomes (25.0%), secondary (20.0%) and primary endpoints (66.7%) was more complete. Overall survival (76.9%) and progression-free survival (67.0%) were the most commonly reported primary or secondary endpoints. Conclusions: This study highlights deficits in the outcome reporting of older cancer patients in GOA trials. Simultaneously, this subgroup also remains heavily underrepresented in GOA trials compared to the real-world patient population. These findings highlight the need for improvements in both trial design and outcome reporting standards for older adults. Full article
(This article belongs to the Special Issue Cancer and Aging: Challenges in Geriatric Cancer Survivorship)
23 pages, 3553 KB  
Article
An Offline Digital-Twin-Assisted Decision-Support Framework for Dynamic RO Under Kuwait Solar-Availability Conditions
by Fajer M. Alelaj, Mohammed A. Bou-Rabee, Mustafa Fadel, Shafqat Aziz, Adil Aslam Mir, Abdulrahman Alharbi and Hussain Al-Sairfi
Membranes 2026, 16(9), 281; https://doi.org/10.3390/membranes16090281 - 23 Aug 2026
Abstract
Reverse osmosis (RO) desalination is a major technology for freshwater production in arid regions, but its energy demand becomes more challenging when the system is supplied by variable renewable energy. This study presents an offline digital-twin-assisted decision-support framework for dynamic RO under Kuwait [...] Read more.
Reverse osmosis (RO) desalination is a major technology for freshwater production in arid regions, but its energy demand becomes more challenging when the system is supplied by variable renewable energy. This study presents an offline digital-twin-assisted decision-support framework for dynamic RO under Kuwait solar-availability conditions. Within this framework, the predictive models are driven primarily by the dynamic RO process variables, while NASA Prediction Of Worldwide Energy Resources (POWER) data provide the Kuwait solar-availability context, and the PV power margin serves as a scenario-level energy indicator. The purpose is to predict instantaneous permeate flow rate, estimate specific energy consumption, and identify energy-efficient operating conditions using machine learning. Kuwait City was used as the solar case-study location. Hourly solar and meteorological data were obtained from NASA POWER, while dynamic RO membrane data were obtained from the open experimental wave desalination dataset published by the National Renewable Energy Laboratory (NREL) through Data.gov and the Marine and Hydrokinetic Data Repository. The RO dataset includes steady-state, ramp, sinusoidal, and Wave Energy Converter SIMulator (WEC-Sim) pressure/flow experiments. The process-flow image used in the system description was also taken from the same NREL dataset and is cited in the figure caption. The raw RO files were cleaned, harmonized, and transformed into a process-informed modeling dataset. Derived features included pressure rate, recovery ratio, salt rejection, estimated pump power, specific energy consumption (SEC), PV power margin, and rolling pressure/flow features. Three supervised regression models were tested: Gradient Boosting, Random Forest, and XGBoost. A representative subset of 60,000 records was used to preserve the main experimental conditions while reducing redundancy in the densely sampled sequential data. Results show that permeate flow rate can be predicted with high accuracy using Gradient Boosting (R2 = 0.981; RMSE = 0.161 L/min). The moderate energy prediction performance yielded an R2 of 0.654 and RMSE of 7.570 kWh/m3 for Random Forest. The accuracy of permeate conductivity predictions was lower (R2 = 0.257; RMSE = 245.44 µS/cm) because membrane and feed characterizing parameters should be included for an adequate water quality control. The proposed approach is best suited as an offline decision-support framework for dynamic RO process analysis. Full article
Show Figures

Figure 1

21 pages, 4550 KB  
Article
Investigating Privacy-Preserving Federated Learning for Telecom Customer Churn Prediction Using Differential Privacy
by Alisha Sikri, Shalini Gambhir, Roshan Jameel, Sheikh Mohammad Idrees and Mariusz Nowostawski
Information 2026, 17(8), 811; https://doi.org/10.3390/info17080811 - 21 Aug 2026
Viewed by 74
Abstract
Predicting customer churn in the telecom sector is critical for retaining subscribers, maintaining brand reputation, and staying ahead of competitors. Losing customers not only reduces revenue but can also weaken long-term market position in a highly competitive industry. While machine learning has been [...] Read more.
Predicting customer churn in the telecom sector is critical for retaining subscribers, maintaining brand reputation, and staying ahead of competitors. Losing customers not only reduces revenue but can also weaken long-term market position in a highly competitive industry. While machine learning has been widely used to address this challenge, most traditional approaches depend on centralizing customer data. This raises major concerns about user privacy, data ownership, and compliance with strict regulations such as GDPR. These challenges make it difficult for businesses to fully utilize customer data while safeguarding sensitive information. In this paper, we investigate a privacy-preserving approach to churn prediction that combines federated learning (FL) with differential privacy (DP). Rather than collecting all customer data in a single repository, the investigated framework enables multiple clients to collaboratively train a deep neural network while maintaining data locality during the federated training process. To further enhance privacy protection, we employ Differentially Private Stochastic Gradient Descent (DP-SGD) and add controlled noise to model updates, reducing the possibility of inferring individual data contributions. This work systematically evaluates how different privacy levels, expressed through ε and δ, influence model performance under simulated non-IID client distributions. The experiments analyze the privacy–utility trade-off using multiple evaluation metrics and compare the results with centralized and non-private federated-learning approaches. The findings show that the investigated framework maintains competitive predictive performance across a range of privacy budgets while demonstrating a clear privacy–utility trade-off. Very strict privacy budgets result in substantial performance degradation, particularly for smaller and more imbalanced datasets, whereas moderate privacy budgets maintain competitive predictive performance with limited degradation. This study highlights the potential of privacy-preserving federated learning for practical distributed analytics applications where protecting sensitive data is essential. Full article
(This article belongs to the Section Information Security and Privacy)
Show Figures

Figure 1

18 pages, 382 KB  
Data Descriptor
A Georeferenced Dataset of Electromagnetic Field Exposure Measurements in Colombia
by David L. Ocampo-Rodríguez, Diógenes de Jesus Ramirez-Ramirez and Cristian David Correa-Álvarez
Data 2026, 11(8), 211; https://doi.org/10.3390/data11080211 - 21 Aug 2026
Viewed by 92
Abstract
This Data Descriptor presents a georeferenced dataset of electromagnetic-field exposure measurements collected in Colombia between 2 November 2023 and 30 June 2024. The records originate from the Sistema de Monitoreo de Campos of the Agencia Nacional del Espectro (ANE), which uses isotropic probes [...] Read more.
This Data Descriptor presents a georeferenced dataset of electromagnetic-field exposure measurements collected in Colombia between 2 November 2023 and 30 June 2024. The records originate from the Sistema de Monitoreo de Campos of the Agencia Nacional del Espectro (ANE), which uses isotropic probes to monitor broadband radiofrequency electromagnetic fields from 100 kHz to 8 GHz and reports six-minute averages of incident power density in W/m2. The comma-separated value file contains 1,286,346 timestamped records and seven source variables, including Well-Known Text (WKT) point geometry. The principal 2024 subset comprises 995,238 records from 24 monitoring stations. We provide a reproducible workflow for timestamp and coordinate parsing, structural and numeric validation, station-level aggregation, and spatial sensitivity analysis using inverse distance weighting with the mean, median, and 95th percentile. The dataset does not provide frequency-resolved measurements, calibration certificates, or measurement-uncertainty budgets; these omissions limit the use of the public file for formal regulatory-compliance assessment. The accompanying repository includes validated station-level summaries, descriptive tables, figures, and reproducible R code. The package supports environmental monitoring, geospatial analysis, methodological comparison, and reproducible reuse of electromagnetic-field exposure data. Full article
Show Figures

Figure 1

48 pages, 7388 KB  
Article
IPA-ANN: A Novel Framework for Optimizing Artificial Neural Network Weights and Biases Using Immune Plasma Algorithm
by Sercan Demirci, Durmuş Özkan Şahin, Gülcan Yıldız, Doğan Yıldız and Samad Hasanlı
Biomimetics 2026, 11(8), 597; https://doi.org/10.3390/biomimetics11080597 - 20 Aug 2026
Viewed by 207
Abstract
Classification is a fundamental technique in data mining that predicts categorical labels by analyzing input features. However, training Artificial Neural Networks (ANNs) using traditional methods often encounters challenges, such as getting stuck in local minima and slow convergence. To address these issues, this [...] Read more.
Classification is a fundamental technique in data mining that predicts categorical labels by analyzing input features. However, training Artificial Neural Networks (ANNs) using traditional methods often encounters challenges, such as getting stuck in local minima and slow convergence. To address these issues, this study proposes a novel hybrid model, IPA-ANN, which integrates the Immune Plasma Algorithm (IPA) to optimize the ANN’s connection weights and biases. The IPA, inspired by the immune plasma treatment process, utilizes a unique donor-receiver mechanism to balance exploration and exploitation in the search space. The proposed model was evaluated on nine benchmark datasets from the UCI repository and compared with 18 state-of-the-art metaheuristic algorithms, including Grey Wolf Optimization (GWO), Differential Evolution (DE), and Particle Swarm Optimization (PSO). Experimental results were analyzed using accuracy, F1-score, confusion matrices, and convergence graphs. The findings indicate that IPA-ANN achieves competitive and stable classification performance across different datasets while demonstrating favorable convergence characteristics in several cases. Furthermore, the study investigates the influence of donor–receiver parameters on the optimization process, highlighting the adaptability of the proposed framework. The reliability of these findings was further examined through repeated stratified 5-fold cross-validation and paired Wilcoxon signed-rank tests with Holm–Bonferroni correction on representative datasets, confirming that a subset of the observed performance differences are statistically significant, and through a computational cost analysis showing that IPA-ANN incurs no additional overhead relative to the majority of the compared algorithms. This study contributes to the literature by presenting the first documented application of IPA in ANN training and by providing a modular infrastructure for future metaheuristic-based ANN optimization studies. Full article
Show Figures

Graphical abstract

17 pages, 11487 KB  
Article
Integrated Analysis of Multiple Databases Identifies Tissue Inhibitor of Metalloproteinase 1 Expression and Its Association with the Immune Microenvironment in Colorectal Cancer
by Yun Xie, Jun Li, Zuwei Yan and Wenguang Zhang
Genes 2026, 17(8), 977; https://doi.org/10.3390/genes17080977 - 20 Aug 2026
Viewed by 196
Abstract
Background: In recent decades, the incidence of colorectal cancer (CRC) has been rising worldwide. CRC ranks second in cancer-related mortality. The identification of reliable biomarkers for early diagnosis and prognosis prediction, along with a deeper understanding of the underlying molecular events, holds substantial [...] Read more.
Background: In recent decades, the incidence of colorectal cancer (CRC) has been rising worldwide. CRC ranks second in cancer-related mortality. The identification of reliable biomarkers for early diagnosis and prognosis prediction, along with a deeper understanding of the underlying molecular events, holds substantial promise for improving patient outcomes. The tissue inhibitor of the metalloproteinase 1 (TIMP1) gene is overexpressed in various gastrointestinal malignancies and contributes to tumor progression. However, its role in regulating the CRC tumor immune microenvironment (TIME) and its potential as a clinically actionable prognostic biomarker remain unclear. Methods: To probe how TIMP1 acts as a prognosis-related candidate biomarker in colorectal carcinoma, TCGA-derived datasets were adopted to conduct Kaplan–Meier survival assessment. We also investigated the connection between the expression abundance of TIMP1 and the infiltration of immune populations and intratumoral lymphocytes; furthermore, immune checkpoint-related genes were systematically assessed across multiple tumor types via the TISIDB and TIMER2.0 platforms, with particular emphasis on CRC. We adopted the ESTIMATE scoring system to figure out how TIMP1 gene expression correlates with the phenotypic properties of the colorectal-cancer TIME. We relied on the limma toolkit for the screening of differential transcripts from high-TIMP1 and low-TIMP1 cohorts. Enrichment assessments covering Gene Ontology terms and Kyoto Encyclopedia of Genes and Genomes entries were then carried out to predict the potential biological pathways associated with TIMP1. We constructed the protein–protein interaction map for TIMP1-interacting partners via the STRING repository. To further explore TIMP1-correlated genes, we performed Venn diagram intersection analysis combined with Spearman’s correlation test. Finally, quantitative reverse-transcription PCR was then implemented to detect TIMP1 messenger-RNA abundance inside the RKO colorectal carcinoma cell line as well as normal colonic epithelial CCD-18Co cells, which offered in vitro experimental verification for our bioinformatic outcomes. Results: According to outcome data, TIMP1 transcripts were markedly up-regulated in CRC specimens and cell lines relative to normal samples. Elevated TIMP1 expression served as a poor-prognosis indicator for overall survival (hazard ratio [HR] = 0.43, 95% confidence interval [CI] = 0.29–0.64, p < 0.001) and disease-specific survival (HR = 0.39, 95% CI = 0.22–0.68, p = 0.001) among colorectal-carcinoma patients. TIMP1-high and TIMP1-low groups exhibited notable differences in immune cell infiltration (CD8+ T, macrophage, mast, neutrophil, B, monocyte, dendritic, and CD4+ T cells). TIMP1 expression was also significantly correlated with tumor-infiltrating lymphocytes, key immune checkpoint genes (e.g., CD274 [PD-L1] and CTLA4), and immunomodulatory chemokines (e.g., CCL3 and CCL5). Twelve TIMP1-interacting DEGs were selected: COL5A1, FN1, PRG4, and a cluster of nine MMPs (MMP1/2/3/7/8/9/11/13/14), all of which showed significant positive correlations with TIMP1 (r = 0.31–0.63, all p < 0.001). Conclusions: TIMP1 expression correlates with features of the tumor immune microenvironment and extracellular matrix remodeling in CRC, suggesting that TIMP1 shows potential as a candidate biomarker. However, its potential as a therapeutic target warrants further experimental investigation. Full article
(This article belongs to the Section Human Genomics and Genetic Diseases)
Show Figures

Figure 1

22 pages, 723 KB  
Review
Closing the Loop with Gates: A Scale-up-Gated Design–Build–Test–Learn Framework for Industrial Fermentation
by Xiang He, Yanling Hu, Yao Zhu, Xinli Li, Kenan Wang, Liqing Dong, Xiaolong He, Yueqin Liu, Jianzhao Qi and Pengfei Jin
Microorganisms 2026, 14(8), 1830; https://doi.org/10.3390/microorganisms14081830 - 19 Aug 2026
Viewed by 250
Abstract
The global fermentation industry faces persistent bottlenecks in scaling laboratory innovations to industrial production, and the integration of synthetic biology (SynBio) and artificial intelligence (AI) within the Design–Build–Test–Learn (DBTL) loop has yielded inconsistent industrial outcomes. This review proposes that transformative impact requires a [...] Read more.
The global fermentation industry faces persistent bottlenecks in scaling laboratory innovations to industrial production, and the integration of synthetic biology (SynBio) and artificial intelligence (AI) within the Design–Build–Test–Learn (DBTL) loop has yielded inconsistent industrial outcomes. This review proposes that transformative impact requires a “scale-up-gated DBTL” framework, in which explicit decision gates constrain every iteration. At the Design phase, scale-down simulation data must inform genetic design choices. At the Test phase, downstream processing compatibility and industrial robustness metrics are enforced as non-negotiable evaluation criteria. At the Learn phase, techno-economic analysis (TEA) and life-cycle assessment (LCA) serve as the convergence criteria, replacing traditional titer plateaus. Through a qualitative cross-sectoral analysis of food, pharmaceutical, agricultural, and energy fermentation, the analysis reveals that workflows incorporating such constraints consistently bridge the valley of death, whereas unconstrained DBTL systematically converges on laboratory optima that are industrially unviable. Five strategic priorities are outlined—embedding TEA/LCA into DBTL, adopting scale-down simulation, building open fermentation data repositories, harmonizing regulatory frameworks, and fostering cross-disciplinary training—as prerequisites for progressing toward fully autonomous, scale-up-aware biomanufacturing. Full article
(This article belongs to the Section Microbial Biotechnology)
Show Figures

Figure 1

11 pages, 1284 KB  
Article
Minimal Efficacy of Single-Agent Anti-PD-(L)1 Re-Exposure in Anti-PD-(L)1-Refractory Merkel Cell Carcinoma: A Retrospective Cohort Study of 16 Patients
by Peter Y. Ch’en, Yuzheng Zhang, Daniel S. Hippe, Rashmi Bhakuni, Tomoko Akaike, Evan T. Hall and Paul Nghiem
Cancers 2026, 18(16), 2673; https://doi.org/10.3390/cancers18162673 - 18 Aug 2026
Viewed by 254
Abstract
Background/Objectives: Anti-PD-(L)1 immune checkpoint inhibitors (ICIs) provide durable responses in nearly 50% of patients with advanced Merkel cell carcinoma (MCC). However, for those progressing on first-line ICI therapy, optimal subsequent therapy remains unclear. Although presumed to have limited benefit, the efficacy of re-exposure [...] Read more.
Background/Objectives: Anti-PD-(L)1 immune checkpoint inhibitors (ICIs) provide durable responses in nearly 50% of patients with advanced Merkel cell carcinoma (MCC). However, for those progressing on first-line ICI therapy, optimal subsequent therapy remains unclear. Although presumed to have limited benefit, the efficacy of re-exposure with anti-PD-(L)1 alone has not been formally reported. We assessed real-world outcomes of patients within a Seattle-based MCC repository who received this salvage approach. Methods: Among 106 patients who received salvage therapy after first-line ICI progression, 16 underwent single-agent anti-PD-(L)1 monotherapy re-exposure during salvage. Patients progressing >3 months after their last immunotherapy dose were excluded. Outcomes included progression-free survival (PFS), disease-specific survival (DSS), and objective response rate (ORR). Results: The median time from end of first-line ICI therapy to anti-PD-(L)1 re-exposure was 51 days (IQR 22–92). Most patients switched between PD-1 and PD-L1 inhibitors (n = 9), while others were re-exposed with the same agent (n = 5) or a different PD-1 inhibitor (n = 2). One of 16 patients experienced a partial response with the same PD-1 inhibitor (ORR 6%; 95% CI: 0.2–30%) at 3 months after re-exposure, followed by progression 10 months after re-exposure. Median PFS was 2.2 months (95% CI: 1.3–5.1 months), and median DSS was 14.7 months (95% CI: 10.4–NR). Conclusions: These data suggest that re-exposure with anti-PD-(L)1 monotherapy confers minimal and short-lived benefit in ICI-refractory MCC, reinforcing the need to develop alternative salvage strategies. Future trials for ICI-refractory MCC mandating an ICI monotherapy arm are unlikely to be appealing to patients or physicians based on a low chance of clinical benefit for this approach. Full article
(This article belongs to the Special Issue Recent Advances in Diagnosis and Therapy of Skin Cancers)
Show Figures

Figure 1

37 pages, 6338 KB  
Article
Valorizing Residue Biomass into Bioenergy: An Explainable Hybrid Machine Learning Model for Predicting Higher Heating Value (HHV) from Elemental Composition
by Yıldırım Özüpak, Emrah Aslan, Mehmet Burukanli and Davut Ari
Sustainability 2026, 18(16), 8412; https://doi.org/10.3390/su18168412 - 17 Aug 2026
Viewed by 139
Abstract
Transforming waste and agricultural-residue biomass into bioenergy is central to the circular bioeconomy, yet routing such heterogeneous residues to the right thermochemical pathway depends on the higher heating value (HHV), which is conventionally measured by slow, resource-intensive bomb calorimetry. Here, we present an [...] Read more.
Transforming waste and agricultural-residue biomass into bioenergy is central to the circular bioeconomy, yet routing such heterogeneous residues to the right thermochemical pathway depends on the higher heating value (HHV), which is conventionally measured by slow, resource-intensive bomb calorimetry. Here, we present an explainable alternative that predicts HHV from inexpensive elemental inputs. We used a publicly archived compilation of 344 literature-reported biomass samples retrieved from an open data repository rather than assembled by the authors, including carbon (C), hydrogen (H), oxygen (O), nitrogen (N) and sulfur (S). Measured HHV was the target. The samples spanned woody, herbaceous and agricultural-residue biomass, and they were standardized through duplicate removal, consistency verification and outlier assessment. On these features, we developed a stacked hybrid model combining Random Forest, eXtreme Gradient Boosting and Artificial Neural Networks, which estimated the HHV with R2 = 0.99, RMSE = 0.45 MJ/kg and MAE = 0.30 MJ/kg. SHAP and LIME analyses showed that carbon exerts the strongest positive influence on HHV, whereas oxygen contributes negatively, which is consistent with established thermochemical principles. Within the compositional range covered by the training data, and subject to the absence of external validation, the framework offers a fast and interpretable complement to bomb calorimetry for screening residue biomass. Full article
Show Figures

Figure 1

23 pages, 1332 KB  
Article
From Word Embeddings to Semantic Projections: Interpretability and Context in Web-Scale Semantic Analysis
by Mabel López-Bordao, Antonia Ferrer-Sapena, Pablo Lara-Navarra, Carlos A. Reyes-Pérez and Claudia Sánchez-Arnau
Information 2026, 17(8), 787; https://doi.org/10.3390/info17080787 - 17 Aug 2026
Viewed by 160
Abstract
Recent developments in NLP and web-scale document analysis have increasingly emphasized the importance of interpretability and contextual dependence in semantic representations. Although modern word embeddings achieve remarkable empirical performance, their semantic structure is often difficult to interpret, since meaning is encoded through latent [...] Read more.
Recent developments in NLP and web-scale document analysis have increasingly emphasized the importance of interpretability and contextual dependence in semantic representations. Although modern word embeddings achieve remarkable empirical performance, their semantic structure is often difficult to interpret, since meaning is encoded through latent geometric relations in high-dimensional spaces. This paper discusses an alternative conceptual framework based on explicit contextual semantic relations. Building on ideas from distributional semantics, co-occurrence analysis, and fuzzy set theory, the study revisits semantic projections and related count-based representations as interpretable directional semantic structures for semantic analysis in document corpora and web-based information environments. In this setting, several classical association measures, including PMI and related transformations, may be understood as derived from simpler conditional semantic projections. The methodology is illustrated through a comparative analysis of semantic associations related to “ChatGPT” across general web-scale data and specialized scientific repositories. Our results demonstrate that semantic projections effectively capture persistent contextual structures while remaining sensitive to corpus-specific discourse communities. The resulting perspective emphasizes interpretability, asymmetry, contextual dependence, and direct empirical meaning as central principles for semantic representation. Full article
(This article belongs to the Special Issue Recent Developments and Implications in Web Analysis, 2nd Edition)
Show Figures

Graphical abstract

23 pages, 1517 KB  
Article
Development and Validation of an Interpretable Machine Learning Model for Inpatient Fall Risk Using Electronic Health Record Data
by Siti Zubaidah Mordiffi, Xiujuan Guo, Mien Li Goh, Kee Yuan Ngiam, Neng Wei Wong, Jenny Chua, Mohammad Shaheryar Furqan and Han Shi Jocelyn Chew
Nurs. Rep. 2026, 16(8), 283; https://doi.org/10.3390/nursrep16080283 - 13 Aug 2026
Viewed by 251
Abstract
Background: Falls are the most common hospital-acquired adverse event, leading to extended hospitalization, loss of independence, disability, and premature death. Routine fall risk assessments are time-consuming, even with limited factors. An AI-derived fall prediction model can provide more comprehensive and comparably accurate [...] Read more.
Background: Falls are the most common hospital-acquired adverse event, leading to extended hospitalization, loss of independence, disability, and premature death. Routine fall risk assessments are time-consuming, even with limited factors. An AI-derived fall prediction model can provide more comprehensive and comparably accurate risk predictions quickly and as often as needed. Objective: To develop and validate a fall prediction model for fall risk in adult inpatients. Methods: Patient records from 2016 were extracted from the adult inpatient database, including information from the Electronic Inpatient Medication Records, SAP, and Hospital Incident Reporting System. The sample consisted of 1506 cases (1:5 faller to non-faller). The fall prediction model was trained using the following four variables: demographics, diagnosis, medications, and surgery. Data sources included the hospital’s data repository, integrating admission/discharge, pharmacy, laboratory, and incident reports. Results: The support vector machine model performed best among all tested models, achieving an AUC of 0.803, recall of 0.816, and precision of 0.440. In the validation cohort (978 patients: 163 fallers and 815 non-fallers), the fall prediction model demonstrated moderate-to-good discrimination (AUC 0.79), with accuracy of 0.67, sensitivity of 0.46, and specificity of 0.86. Compared with the nursing four-item fall risk assessment, which showed lower discrimination (AUC 0.65, accuracy 0.65, sensitivity 0.58, specificity 0.72), the fall prediction model had better specificity and overall discrimination, though the nursing tool was more sensitive in identifying fallers. Conclusions: The fall prediction model using demographics, diagnoses, medication, and surgery data predicts falls risk effectively. It enables timely, accurate risk assessments and supports preventive interventions, saving nurses’ time for direct patient care. Full article
(This article belongs to the Special Issue AI in Nursing: Promoting Patient Safety and Care Quality)
Show Figures

Figure 1

29 pages, 998 KB  
Article
Antibiotic-Specific Genotype–Phenotype Concordance and Cross-Database Interoperability in Public Escherichia coli/Shigella WGS-AMR Metadata
by Abdullateef Abdullah Alshehri
Int. J. Mol. Sci. 2026, 27(16), 7157; https://doi.org/10.3390/ijms27167157 - 10 Aug 2026
Viewed by 248
Abstract
Public pathogen-genomics repositories are increasingly used for antimicrobial resistance (AMR) surveillance, yet the reliability and interoperability of database-derived genotype–phenotype inference remain incompletely characterized. This study evaluated antibiotic-specific concordance between exported AMR genotype annotations and antimicrobial susceptibility testing (AST) phenotypes in the NCBI Pathogen [...] Read more.
Public pathogen-genomics repositories are increasingly used for antimicrobial resistance (AMR) surveillance, yet the reliability and interoperability of database-derived genotype–phenotype inference remain incompletely characterized. This study evaluated antibiotic-specific concordance between exported AMR genotype annotations and antimicrobial susceptibility testing (AST) phenotypes in the NCBI Pathogen Detection metadata for the Escherichia coli/Shigella organism group and used BV-BRC to assess complementary phenotype-related coverage and cross-resource linkage. Frozen NCBI Pathogen Detection and BV-BRC exports were analyzed retrospectively. NCBI records were filtered for assembly accessions, AMR genotype annotations, and interpretable resistant or susceptible AST results. Prespecified antibiotic-specific mapping rules were applied. Performance metrics, including sensitivity, specificity, accuracy, balanced accuracy, positive and negative predictive values and F1-score, were calculated, and their 95% confidence intervals were estimated using 10,000 isolate-level bootstrap replicates. BV-BRC was evaluated descriptively for phenotype-related coverage, evidence type, and accession overlap. Among 539,918 NCBI records, 10,201 met prespecified criteria for assembly accession, AMR genotype annotation, and interpretable resistant/susceptible AST data. These records generated 56,141 genotype–phenotype comparisons across eight priority antibiotics. Concordance was high for tetracycline (accuracy, 98.4%; F1-score, 98.1%) and ceftriaxone (accuracy, 97.9%; F1-score, 94.6%), but lower for amoxicillin–clavulanic acid (accuracy, 73.6%; F1-score, 44.5%), indicating antibiotic-specific limits of genotype-field inference. Discordance was explicitly separated into 2524 genotype-positive/phenotype-susceptible and 923 phenotype-resistant observations with no mapped genomic evidence of resistance. BV-BRC contributed 21,298 taxonomy-strict E. coli/Shigella genomes and 4745 phenotype-linked identifiers. Although 15,214 BioSample and 10,121 assembly accessions were shared across exports, no analysis-ready overlap remained between the final NCBI genotype–AST and BV-BRC phenotype-linked subsets after eligibility filtering. Public WGS-AMR databases can support large-scale surveillance-oriented concordance analyses, but performance is antibiotic specific and depends on mapping rules, phenotype definitions, evidence provenance, and accession linkage. These estimates do not constitute clinical diagnostic validation and should not replace phenotypic AST. Because the analysis was conducted at the combined organism-group level, these estimates may mask species-, pathotype-, or lineage-specific resistance dynamics. Full article
Show Figures

Figure 1

35 pages, 2341 KB  
Article
Evaluating Open Government Data as a Tool for Planning and Sustainability in U.S. Cities: Portals, Policies, and Plans
by Gulnara N. Nabiyeva and Stephen M. Wheeler
Sustainability 2026, 18(16), 8177; https://doi.org/10.3390/su18168177 - 10 Aug 2026
Viewed by 343
Abstract
Open Government Data (OGD) has the potential to support urban planning by increasing transparency, improving evidence-based decision-making, enhancing public participation, and monitoring progress toward sustainability goals. Despite this potential, the extent to which OGD supports planning and sustainability in practice remains unclear. This [...] Read more.
Open Government Data (OGD) has the potential to support urban planning by increasing transparency, improving evidence-based decision-making, enhancing public participation, and monitoring progress toward sustainability goals. Despite this potential, the extent to which OGD supports planning and sustainability in practice remains unclear. This study evaluates the user-friendliness of OGD portals in 19 U.S. cities using four criteria identified in the literature: breadth of portal content, structure of governance, user-friendly support, and equitable user access. It also examines how cities incorporate OGD into OGD policies, general plans, and sustainability/climate action plans. This study was designed as an exploratory comparative case study of OGD practices, and the methods included portal assessment using four criteria, content analysis of OGD policies and plans, and semi-structured interviews with OGD and planning experts. The findings indicate that most OGD portals function primarily as data repositories, rather than as planning support tools, and vary considerably in portal content, governance, and user support, affecting their usefulness for planning and sustainability efforts. OGD policies are often absent or underdeveloped and typically emphasize transparency rather than planning applications or public engagement. Only nine cities identified a role for OGD in at least one planning-related document. Interviewees identified low public awareness, limited user capacity, and insufficient integration of data into decision-making as key barriers. The findings identify practical strategies that cities can adopt to better integrate OGD into urban planning and sustainability, including proactive identification of datasets and indicators that will be of interest to various stakeholders, improved stakeholder engagement, data standardization, use of graphics to highlight policy-relevant data trends, improved OGD management, and closer alignment between OGD initiatives and action-oriented planning processes. Full article
(This article belongs to the Section Sustainable Urban and Rural Development)
Show Figures

Figure 1

18 pages, 5026 KB  
Article
Flype: Integrating Molecular and Pharmacogenomic Results to Enhance Oncology Patient Care in a Community-Based Academic Cancer Center
by Donald L. Helseth, Nicholas Miller, Mathew Yang, Henry Wittich, Qin Zhao, Tom Werth, Linda M. Sabatini, Mir Alikhan, Megan Parilla, Amandeep Kaur, Xiaoyan Yang, Kathy A. Mangold, Michael Bouma, Henry M. Dunnenberger, Dyson T. Wake, Annette Sereika, Gayathri Moorthy, Peter J. Hulick, Karen L. Kaul and Janaradan D. Khandekar
Cancers 2026, 18(16), 2560; https://doi.org/10.3390/cancers18162560 - 10 Aug 2026
Viewed by 258
Abstract
Background/Objectives: We describe how our in-house bioinformatics platform, Flype, has evolved from being a variant repository to an enterprise role as an electronic medical record (EMR) content provider, powering molecular pathology reporting, pharmacogenomics reporting, sending discrete data to our EMR, aggregating internal and [...] Read more.
Background/Objectives: We describe how our in-house bioinformatics platform, Flype, has evolved from being a variant repository to an enterprise role as an electronic medical record (EMR) content provider, powering molecular pathology reporting, pharmacogenomics reporting, sending discrete data to our EMR, aggregating internal and external molecular test results and powering our molecular tumor board (MTB). Methods: In response to critical pain points, we developed Docket, a sample tracking system in Flype, which manages multiple individual in-house molecular tests for NGS assays, pharmacogenomic (PGX) assays and additional molecular testing. To help with interpretation and integration of all internal and external assays, we developed a clinical outcomes view in Flype. To improve the efficiency of our molecular pathologists reporting results, we developed Convo 2.0, which integrates OncoKB interpretations and other information. Flype can also be used by our pathologists to submit patient molecular results to NCI’s ComboMATCH and retrieve clinical trial recommendations. Results: Flype was used during our Kellogg Cancer Genomic Initiative for reporting PGX integration, MTB review and integration of EMR prescription information with internal and external molecular test results. Integrating PGX results led to several recommendations against the use of drugs metabolized by, for example, CYP2D6 or TPMT, along with warnings about altered pain relief. Enhancements in report sign-out and the use of file transfer scripts have led to reduced turnaround time. Conclusions: Flype supports hundreds of users performing different roles in molecular diagnostics. We discuss lessons learned adapting our software to support continuously changing test requirements. Full article
(This article belongs to the Special Issue Pharmacogenetics and Pharmacogenomics in Oncology)
Show Figures

Figure 1

24 pages, 4303 KB  
Review
Agent-Based Modeling and System Dynamics Integrated with AI and Analytical Methods: A Structured Review of Hybrid Approaches, Applications, and Future Directions for Decision Making in Complex Systems (2021–2026)
by Ionela Samuil, Andreea Ionica and Monica Leba
Systems 2026, 14(8), 956; https://doi.org/10.3390/systems14080956 - 7 Aug 2026
Viewed by 332
Abstract
The complexity of socio-economic, ecological, and public-health systems demands simulation frameworks capturing both macro dynamics and micro agent heterogeneity. Agent-based modeling (ABM) and system dynamics (SD), increasingly coupled with AI and analytical methods, support decision making in complex systems, yet no systematic review [...] Read more.
The complexity of socio-economic, ecological, and public-health systems demands simulation frameworks capturing both macro dynamics and micro agent heterogeneity. Agent-based modeling (ABM) and system dynamics (SD), increasingly coupled with AI and analytical methods, support decision making in complex systems, yet no systematic review covers 2021–2026. A two-round PRISMA 2020 search (September 2025; May 2026) in IEEE Xplore, Web of Science, and ProQuest identified 70 eligible papers across 14 thematic clusters; because the second search round closed in May 2026, papers published later in 2026 are necessarily under-represented relative to complete prior years. Data were extracted along eleven dimensions; inter-rater reliability was κ = 0.88–0.90. Output grew 133% between 2023 and 2025. C14 (Multi-Method Simulation) and C7 (Healthcare & Epidemiology) are the largest clusters; AnyLogic dominates as the only natively tri-paradigm platform (22.9%). Hybrid models consistently identify critical intervention thresholds invisible to mono-paradigm approaches. Only 44.3% of papers report formal structural validation and 88.6% withhold code. Nine papers (12.9%) integrate ML/AI, but none apply interpretability techniques (SHAP, LIME, ICE). ABM–SD hybridization is a maturing paradigm whose epistemic gain depends on closing validation, reproducibility, and interpretability gaps. Six gaps and five priority directions for 2026–2030 are identified, including a minimum validation protocol, a standardized repository, and a dedicated interpretability framework for ML+ABM–SD hybrids that preserves causal transparency for decision making in complex systems. Full article
Show Figures

Figure 1

Back to TopTop