Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (133)

Search Parameters:
Keywords = tabular-to-image

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
54 pages, 743 KB  
Article
Utilizing Concept Ontologies for Designing Neural Network-Based Classifiers
by Kamil Szwed, Jan G. Bazan, Stanislawa Bazan-Socha, Krzysztof Wójcik and Pawel Milan
Appl. Sci. 2026, 16(16), 8152; https://doi.org/10.3390/app16168152 - 15 Aug 2026
Viewed by 160
Abstract
Deep neural networks have achieved substantial success in image, text, and signal analysis, but their advantage is less consistent for heterogeneous tabular data, where tree-based ensemble methods often remain strong baselines. This study proposes CANON (Cross-Attention Neuro-symbolic Ontology Network), a neuro-symbolic architecture that [...] Read more.
Deep neural networks have achieved substantial success in image, text, and signal analysis, but their advantage is less consistent for heterogeneous tabular data, where tree-based ensemble methods often remain strong baselines. This study proposes CANON (Cross-Attention Neuro-symbolic Ontology Network), a neuro-symbolic architecture that integrates a hierarchical a priori concept ontology with a modular mixture-of-experts mechanism. CANON is designed to combine data-driven representation learning with explicit domain structure and to reduce the influence of irrelevant or weakly informative features. The architecture was evaluated on two clinical tabular cohorts—848 patients with ANCA-associated vasculitis described by 142 features, and 200 patients assessed for coronary artery stenosis described by 593 features—using stratified 8-fold cross-validation, and was compared with tree-based ensembles, dedicated tabular deep learning models, and classical neural architectures. CANON achieved the highest mean AUC on both cohorts (0.9426 and 0.9484). Two findings are reported. First, CANON significantly outperformed in AUC all four non-tree baselines included in the paired statistical analysis: all eight paired comparisons against GaussianNB, FT-Transformer, TabNet and LSTM were significant on both cohorts and favored CANON in 8 of 8 folds, with FT-Transformer and TabNet tuned separately for each cohort under an identical budget of 20 Optuna trials each. Second, CANON performed comparably to the tree-based ensembles, significantly outperforming Random Forest on the coronary artery stenosis cohort (Δ=+0.124, p=0.004), while the remaining comparisons against Random Forest and XGBoost did not reach significance. An ablation study comprising 40 cross-validation folds per variant shows that the semantic organization of features into concepts and the non-linear feature tokenizer contribute measurably to the harder, higher-dimensional cohort, whereas the directional fusion mechanism does not improve predictive performance and is retained for its interpretability role. These findings indicate that ontology-guided neural architectures can provide a competitive and interpretable basis for clinical decision support and may also be useful in other regulated domains in which predictive models must remain consistent with domain-specific knowledge. Full article
Show Figures

Figure 1

51 pages, 7279 KB  
Technical Note
From Medical Records to AI-Ready Datasets: A Practical Guide for Clinical Researchers
by Catalin Anghel, Andreea Alexandra Anghel, Marian Viorel Craciun, Simona Moldovanu, Adina Cocu, Diana-Elena Vulpe, Calina Maier, Vasile Potop, Christiana Diana Maria Dragosloveanu, Constantin Adrian Andrei, Serban Dragosloveanu and Cristian Scheau
J. Clin. Med. 2026, 15(16), 6297; https://doi.org/10.3390/jcm15166297 - 14 Aug 2026
Viewed by 202
Abstract
Background: Medical artificial intelligence (AI), machine learning (ML), and deep learning (DL) studies frequently begin with datasets collected for routine care rather than for computational modeling. Such datasets may contain inconsistent variables, heterogeneous measurement time points, unexplained NaN values, poorly defined outcomes, [...] Read more.
Background: Medical artificial intelligence (AI), machine learning (ML), and deep learning (DL) studies frequently begin with datasets collected for routine care rather than for computational modeling. Such datasets may contain inconsistent variables, heterogeneous measurement time points, unexplained NaN values, poorly defined outcomes, missing metadata, and insufficient documentation, which can compromise model development before any algorithm is selected. Methods: This Technical Note proposes a physician-facing Clinical AI-Readiness Guide for preparing medical datasets before AI-based analysis. The guide was developed as a practical framework organized around pre-modeling decisions, including the clinical task, cohort, minimum common dataset, outcome definition, predictor variables, measurement timing, missing-data logic, standardization, non-tabular data linkage, data dictionary, and validation readiness. Results: The proposed guide translates AI-readiness principles into concrete data-collection rules for clinical, laboratory, imaging, physiological-signal, textual, follow-up, and multimodal data. It emphasizes clinically consistent data acquisition, reliable target labeling, explicit missing-data logic, patient-level linkage, structured metadata, and validation feasibility. A structured checklist and scoring approach are also proposed as practical pre-modeling assessment tools to classify datasets as not ready, exploratory only, ML-ready with limitations, or AI-ready for model development. Conclusions: Medical AI-readiness should be established before model development begins. By helping physicians collect, structure, and document data more consistently, the proposed guide may improve collaboration between clinical and technical teams and reduce preventable dataset-related failures in medical AI research. Full article
(This article belongs to the Section Clinical Guidelines)
Show Figures

Figure 1

17 pages, 6118 KB  
Article
Time-Aware Late Fusion for Multimodal Rice Growth-Rate Prediction from UAV Imagery and Weather Context
by Alaa O. Elhadi, Saad M. Darwish and Mahmoud A. Mahdi
Computers 2026, 15(8), 523; https://doi.org/10.3390/computers15080523 - 12 Aug 2026
Viewed by 264
Abstract
Accurate crop-growth estimation from unmanned aerial vehicle (UAV) imagery is important for precision agriculture, but image-only models can struggle to represent seasonal context. This study evaluates whether combining UAV imagery with weather variables and elapsed time improves continuous rice growth-rate prediction in a [...] Read more.
Accurate crop-growth estimation from unmanned aerial vehicle (UAV) imagery is important for precision agriculture, but image-only models can struggle to represent seasonal context. This study evaluates whether combining UAV imagery with weather variables and elapsed time improves continuous rice growth-rate prediction in a single-site, publicly available rice-seedling dataset spanning multiple growing seasons. A Time-Aware Late Fusion (TALF) model is introduced in which a convolutional branch encodes image features, a multilayer perceptron encodes contextual features, and the two streams are merged only at the regression head. Relative humidity, wind speed, and elapsed time are used as contextual inputs after season-aware preprocessing. Evaluation is reported on a chronological multi-season split using internal ablations rather than external generalization claims. TALF achieved a mean absolute error (MAE) of 0.031, compared with 0.1455 for the optimized image-only baseline and 0.0890 for an early-fusion image-and-weather baseline. A secondary tolerance-based metric reached 94.9% under the reported threshold. The results indicate that weather and elapsed-time context improve prediction on this dataset and that separating image and tabular encoders until the final layers is a competitive multimodal learning design under the reported protocol. Full article
Show Figures

Graphical abstract

21 pages, 1901 KB  
Article
Computational Pathology Reveals Extracellular-Matrix Imaging Biomarkers of Therapy Response in Preclinical Breast Cancer
by Stelios Lamprou, Styliana Georgiou, Triantafyllos Stylianopoulos and Chrysovalantis Voutouri
Cancers 2026, 18(16), 2603; https://doi.org/10.3390/cancers18162603 - 12 Aug 2026
Viewed by 278
Abstract
Background/Objectives: Histological biomarkers of therapy response remain incompletely defined in preclinical breast cancer models. We evaluated whether quantitative immunofluorescence imaging and multimodal machine learning (ML) could identify image-derived tissue biomarkers associated with response in a 4T1 murine breast cancer therapy-response setting. Methods: We [...] Read more.
Background/Objectives: Histological biomarkers of therapy response remain incompletely defined in preclinical breast cancer models. We evaluated whether quantitative immunofluorescence imaging and multimodal machine learning (ML) could identify image-derived tissue biomarkers associated with response in a 4T1 murine breast cancer therapy-response setting. Methods: We analysed 1696 immunofluorescence images across nine immunofluorescence staining panels: aPDL1, alpha-smooth muscle actin-CD31, alpha-smooth muscle actin-Ki67, CD3-CD31, CD8-Ki67, collagen–hyaluronic acid (HA), granzyme B-CD8, HMGB1, and pimonidazole/hypoxia. A Python image-analysis algorithm extracted 117 RGB-channel, intensity, morphometric, and cross-channel spatial features per image. The modelling framework included image-only deep learning (DL), tabular ML on extracted histological features, and image-plus-tabular gated fusion. Models were evaluated using stratified cross-validation, train-fold-only class balancing, bootstrap confidence intervals, permutation testing, nested feature-selection sensitivity analysis, and grouped leakage controls. Results: Histological staining panels classified treatment-response status across the nine-panel benchmark. Collagen–HA was the strongest staining (AUC 0.954), followed by alpha-smooth muscle actin-Ki67 (AUC 0.851) and hypoxia (AUC 0.837). DL achieved the best performance in eight of nine stainings. In 346 collagen–HA images, combined collagen and HA features achieved AUC 0.848, collagen-only features AUC 0.833, and hyaluronic-acid-only features AUC 0.805. In 104 animals with matched hypoxia and collagen–hyaluronic-acid images, extracellular-matrix features achieved AUC 0.985, hypoxia-only features achieved AUC 0.929, and adding hypoxia did not improve extracellular-matrix-only prediction. Conclusions: Quantitative extracellular-matrix imaging provides a strong preclinical signal for therapy-response stratification in 4T1 breast cancer. The findings are hypothesis-generating and require independent preclinical and human validation before clinical translation. Full article
(This article belongs to the Section Tumor Microenvironment)
Show Figures

Figure 1

21 pages, 2777 KB  
Article
A Rationale-Conditioned Image-Contrast Audit of OCT Dependence in Vision-Language Models for Anti-VEGF Treatment-Response Prediction
by Wonbong Jang, Gwon Yul Jo, Siyun Lee, Shin Jeong Yoon, Gi Young Lee, Hee-Eun Lee, Tae Hyung Kim, Jong Won Baek and Joonhyung Kim
Bioengineering 2026, 13(8), 912; https://doi.org/10.3390/bioengineering13080912 - 12 Aug 2026
Viewed by 445
Abstract
Vision-language models (VLMs) are increasingly evaluated for retinal image interpretation, but end-task discrimination does not establish whether a prediction depends on the optical coherence tomography (OCT) image. We evaluated RetinaVLM, LLaVA-Med, and Qwen3.6-27B under zero-shot and parameter-efficient adaptation using an APTOS-2021 cohort (128 [...] Read more.
Vision-language models (VLMs) are increasingly evaluated for retinal image interpretation, but end-task discrimination does not establish whether a prediction depends on the optical coherence tomography (OCT) image. We evaluated RetinaVLM, LLaVA-Med, and Qwen3.6-27B under zero-shot and parameter-efficient adaptation using an APTOS-2021 cohort (128 training, 21 validation, and 69 test eyes) and a 100-eye cross-site stress-test cohort. The primary readout compared forced-choice continue/stop scores obtained with a real OCT B-scan and a uniform-grey image while holding the generated rationale fixed; it therefore estimates a rationale-conditioned direct image contrast rather than total image dependence. Across 22 internal cells, 21 confidence intervals included an AUC of 0.5, while one Qwen zero-shot cell was inversely aligned (AUC 0.354, 95% CI 0.224–0.493). No positively aligned cells survived the Benjamini–Hochberg adjustment. Selected backbones produced image-responsive biomarker outputs, including pigment epithelial detachment balanced accuracy up to 0.93, whereas a five-field tabular reference model achieved decision AUC 0.731 (95% CI 0.603–0.846). Similar direct-contrast findings occurred in the second-site stress test. Because the generated rationales were not verified as faithful representations of latent computation, the grey image was out of distribution, and the cohorts were small, null direct contrasts cannot show that the models ignored OCT or establish equivalence to no discrimination. Full article
(This article belongs to the Special Issue Advances in Ocular Diagnosis and Therapy)
Show Figures

Figure 1

22 pages, 11883 KB  
Article
Deep Learning-Based Prediction of Epithelial Cytokine Responses for the Selection of Functionally Consistent Airway Organoids
by Hyeokjin Kweon, Mi Hyun Lim, David W. Jang, Keonhyeok Park, Seungchul Lee and Do Hyun Kim
Biomimetics 2026, 11(8), 547; https://doi.org/10.3390/biomimetics11080547 - 3 Aug 2026
Viewed by 209
Abstract
Although airway organoids provide a physiologically relevant platform for modeling human airway inflammation, their utility is often limited by substantial heterogeneity in epithelial differentiation and functional responsiveness across Matrigel domes. Here, we present a non-destructive, imaging-guided framework to predict epithelial cytokine responses and [...] Read more.
Although airway organoids provide a physiologically relevant platform for modeling human airway inflammation, their utility is often limited by substantial heterogeneity in epithelial differentiation and functional responsiveness across Matrigel domes. Here, we present a non-destructive, imaging-guided framework to predict epithelial cytokine responses and enable the selection of functionally consistent airway organoid domes. Mature human airway organoids were stimulated with house dust mite (HDM) extract and dome-level inflammatory responsiveness was quantified by RT-qPCR for thymic stromal lymphopoietin (TSLP) and interleukin-33 (IL-33). Both cytokines exhibited wide dome-to-dome variability and showed a significant positive correlation, indicating coordinated allergic inflammatory regulation. Meanwhile, bright-field dome images were analyzed to segment individual organoids, define robust regions of interest, and extract quantitative morphological and texture descriptors based on gray-level co-occurrence matrix features. Organoid-level descriptors were aggregated into a single dome-level feature vector using distributional statistics, thereby capturing both central tendency and heterogeneity within each dome. Using these engineered dome-level features, we trained a deep tabular learning model (TabNet) to classify qPCR-defined inflammatory responsiveness. The resulting model achieved strong and consistent cross-validated performance for both targets, reaching balanced accuracies of 0.910 for TSLP and 0.833 for IL-33, demonstrating that bright-field phenotypes contain predictive signatures of cytokine activation. This approach provides a scalable enrichment strategy for robustly responsive organoid–Matrigel domes without destructive assay. It improves reproducibility in organoid-based airway inflammation studies and supports standardized dome selection for downstream mechanistic and translational applications. Full article
Show Figures

Graphical abstract

42 pages, 6187 KB  
Article
TL-RL-FusionNet: Reinforcement Learning-Guided Residual MLP with Fused CNN Embeddings for Efficient and Adaptive Ransomware Detection
by Jannatul Ferdous, Rafiqul Islam, Arash Mahboubi and Md Zahidul Islam
Sensors 2026, 26(15), 4775; https://doi.org/10.3390/s26154775 - 27 Jul 2026
Viewed by 338
Abstract
Ransomware detection remains challenging because modern variants exhibit diverse, elusive, and partly benign behaviors and can propagate rapidly across interconnected enterprises and sensor-enabled cyber-physical systems, causing cascading operational failures. These characteristics undermine signature-based and static-detection methods. Although machine learning has improved detection, many [...] Read more.
Ransomware detection remains challenging because modern variants exhibit diverse, elusive, and partly benign behaviors and can propagate rapidly across interconnected enterprises and sensor-enabled cyber-physical systems, causing cascading operational failures. These characteristics undermine signature-based and static-detection methods. Although machine learning has improved detection, many approaches still rely on fixed objectives that weight samples uniformly, limiting their adaptation to heterogeneity and overlaps between ransomware and benign activities. To address this challenge, we introduce TL-RL-FusionNet, a reinforcement learning (RL)-guided hybrid framework that combines dual transfer learning (TL) backbones, EfficientNetB0 and InceptionV3, with a lightweight residual multi-Layer perceptron (MLP) classifier. The framework converts sandbox reports into RGB grids, extracts features using frozen CNN backbone networks, and fuses embeddings for classification. Training is guided by a tabular Q-learning sample-weighting agent, formulated as a per-sample bandit over discrete weight actions. To prevent cross-fold information leakage, the Q-table is freshly initialized in each cross-validation fold and updated only using the fold-local training partition, whereas the held-out fold is used for the final evaluation. The framework was evaluated using two datasets. On our dataset, TL-RL-FusionNet achieved the best overall performance on Dataset 1, with 99.20% accuracy, 99.40% recall, and 99.84% AUC. On the public EldeRan benchmark, it achieved 90.36% accuracy using the full dynamic feature space and 92.08% using a Mutual Information-selected compact subset. Paired Wilcoxon tests across five folds were used to assess the RL contribution, while additional grid-order sensitivity analysis showed that the image-based representation remained robust under five random 10 × 10 feature-grid permutations. Interpretability analysis using t-distributed stochastic neighbor embedding (t-SNE) and gradient-weighted class activation mapping feature-grid mapping further showed that the model captured discriminative behavioral patterns. Overall, these results demonstrate that RL-guided sample reweighting improves adaptive ransomware detection while maintaining efficiency and interpretability. The dataset and supporting code are publicly available on GitHub. Full article
(This article belongs to the Special Issue Intelligent Sensors for Security and Attack Detection)
Show Figures

Figure 1

29 pages, 52303 KB  
Article
Landslide Susceptibility Mapping Using an Image–Tabular Joint Deep Learning Framework: A Case Study of the Tacheng Region, Xinjiang, China
by Qianjie Deng, Dingfan Xing, Xiong Wu, Lirui Song, Zhuoer Teng, Rui Wang, Shichen Gao, Zhiwu Zhang and Kun-Feng Qiu
Remote Sens. 2026, 18(15), 2436; https://doi.org/10.3390/rs18152436 - 23 Jul 2026
Viewed by 574
Abstract
Accurate landslide susceptibility mapping (LSM) is important for hazard prevention and land use planning in mountainous regions. Existing machine learning and deep learning methods mainly use raster-based conditioning factors. They often ignore landslide-related attribute information and spatial context. To address this issue, this [...] Read more.
Accurate landslide susceptibility mapping (LSM) is important for hazard prevention and land use planning in mountainous regions. Existing machine learning and deep learning methods mainly use raster-based conditioning factors. They often ignore landslide-related attribute information and spatial context. To address this issue, this study proposes an image–tabular joint deep learning framework for regional-scale LSM. The framework is based on a FiLM-conditioned U-Net. The model combines raster patches with an estimated soft attribute-prior vector and uses FiLM to guide condition-aware spatial feature learning. The proposed framework was tested in the Tacheng region, Xinjiang, China. The dataset includes a landslide inventory and conditioning factors related to terrain, hydrology, vegetation, geology, land cover, and human activities. Model performance was evaluated using stratified five-fold cross-validation, an independent test set, buffer-radius sensitivity tests, and spatial hold-out validation. FiLM-U-Net achieved the best performance among the tested models. It obtained an accuracy of 89.73%, an F1-score of 89.41%, and an AUC of 0.953 on the independent test set. In the spatial hold-out validation area, the model achieved an AUC of 0.921. Feature importance analysis showed that distance to roads, rainfall, NDVI, and terrain factors provided important predictive information. These results suggest that the proposed image–tabular joint framework can improve condition-aware feature learning and support regional landslide susceptibility assessment. Full article
Show Figures

Figure 1

65 pages, 3965 KB  
Systematic Review
Alzheimer’s Disease Detection Based on Machine Learning and Deep Learning Frameworks: A Cross-Dataset Comparative Performance Analysis and Assessment of Clinical Readiness
by Keenan Ramnarain, Rito Clifford Maswanganyi and Philani Khumalo
Mach. Learn. Knowl. Extr. 2026, 8(7), 217; https://doi.org/10.3390/make8070217 - 22 Jul 2026
Viewed by 1295
Abstract
Alzheimer’s disease (AD) is the most prevalent neurodegenerative disorder worldwide, affecting approximately 56.9 million people in 2021 and projected to reach 152 million by 2050. Its defining pathological features, amyloid-beta plaques and neurofibrillary tangles, accumulate for up to two decades before cognitive symptoms [...] Read more.
Alzheimer’s disease (AD) is the most prevalent neurodegenerative disorder worldwide, affecting approximately 56.9 million people in 2021 and projected to reach 152 million by 2050. Its defining pathological features, amyloid-beta plaques and neurofibrillary tangles, accumulate for up to two decades before cognitive symptoms emerge, placing the preclinical and mild cognitive impairment (MCI) stages at the centre of the early detection problem. Despite this, current diagnostic practice in routine clinical settings remains unreliable, with post-mortem studies placing the specificity of clinical AD diagnosis between 44.3 and 70.8% even in specialist memory clinics. Machine learning (ML) and deep learning (DL) applied to neuroimaging and electrophysiological data have emerged as candidate tools for closing this diagnostic gap, yet whether the accuracy figures reported in published studies translate into clinically useful performance on independent data remains unresolved. This study presents a structured comparative review of machine learning and deep learning methods reported across four publicly available Alzheimer’s disease datasets, namely the Alzheimer’s Disease Neuroimaging Initiative (ADNI), the Open Access Series of Imaging Studies (OASIS), the OpenNeuro ds004504 electroencephalography (EEG) dataset, and the Kaggle Alzheimer’s magnetic resonance imaging (MRI) dataset. Thirteen model families are examined through the published literature rather than through new experiments, and for each model and dataset combination, the best accuracy reported in the source study is recorded alongside the model’s mathematical formulation. All performance figures reported in this abstract and throughout the paper are taken from the published studies reviewed, not from new experiments conducted by the authors. Across the reviewed studies, deep learning architectures pre-trained on ImageNet and fine-tuned on neuroimaging data are reported to produce the highest accuracy on MRI classification tasks. Residual Network (ResNet)-101 is reported at 98.21 percent on ADNI and 97.45 percent on OASIS, while the IncepRes fusion architecture reaches 98.35% on OASIS by combining multi-scale feature extraction from InceptionV3 with residual connectivity from ResNet152V2. Traditional machine learning classifiers remain competitive on tabular clinical and biomarker data, with Extreme Gradient Boosting (XGBoost) reaching 91% on ADNI multiclass features. Logistic Regression achieves 82 to 85% on binary ADNI classification and is the only classifier in this review that provides explicit per-feature prediction contributions without post hoc tooling. Gaussian Naïve Bayes achieves 80 to 83% on the same task. On the OpenNeuro EEG dataset, K-nearest neighbours (KNN) with singular value decomposition (SVD) entropy features achieves 91% binary accuracy, with feature engineering quality determining performance more reliably than classifier architecture. Eight principal findings emerge from the cross-dataset analysis. Binary classification consistently outperforms multiclass by 10 to 30% across all datasets, reflecting the genuine biological ambiguity of the mild cognitive impairment category. Dataset size and augmentation predict reported accuracy more reliably than model architecture. Ensemble methods outperform individual classifiers by 5 to 8% in nearly every imaging study. Deeper architectures can overfit small clinical cohorts. EEG models trail MRI models by approximately 10 to 15% on comparable binary tasks. Cross-dataset generalisation has not been systematically evaluated in most studies, and the few that have tested it report accuracy drops of 5 to 10% or more when models encounter data from different scanners or cohorts. Eight recurring limitations constrain the clinical utility of these findings. Small sample sizes and limited demographic diversity, severe class imbalance inflating raw accuracy metrics, poor cross-dataset generalisation driven by scanner heterogeneity, limited deep learning interpretability, the dominance of binary over multiclass tasks, the absence of longitudinal modelling despite available datasets, inadequate standardisation of preprocessing and evaluation protocols, and the signal-to-noise ratio constraints specific to EEG recordings of elderly patients collectively define the gap between benchmark performance and clinical readiness. Future work must prioritise multi-centre training cohorts, multimodal fusion architectures, longitudinal progression modelling, and standardised interpretability evaluation as non-optional requirements for any system intended for clinical deployment. Full article
(This article belongs to the Section Thematic Reviews)
Show Figures

Figure 1

16 pages, 5055 KB  
Article
Application of Machine Learning for Mean Glandular Dose Prediction Utilizing DICOM Mammography Images
by Ali A. A. Alghamdi
J. Imaging 2026, 12(7), 330; https://doi.org/10.3390/jimaging12070330 - 21 Jul 2026
Viewed by 462
Abstract
The growing demand for raw and processed scientific data has encouraged many researchers and research institutions to adopt an open-source data policy. At present, data accessibility is of paramount importance due to the growing demand for artificial intelligence (AI) and machine learning (ML) [...] Read more.
The growing demand for raw and processed scientific data has encouraged many researchers and research institutions to adopt an open-source data policy. At present, data accessibility is of paramount importance due to the growing demand for artificial intelligence (AI) and machine learning (ML) applications in various scientific fields, particularly medicine. Medium- to large-scale mammography datasets are widely used in breast cancer research to develop and evaluate computer-aided detection methods. However, there are only a few studies on using mammogram datasets for the prediction of the breast mean glandular dose (MGD) with AI or ML models. The aim of this study was to investigate the feasibility of using ML and deep ML for MGD prediction based on DICOM images and retrieved dosimetric data from DICOM mammogram images. A total of 26,988 mammography images in DICOM format were obtained from the Federated Research Data Repository (FRDR). Eleven regression algorithms and three neural network-based models were evaluated using five-fold cross-validation. In addition, a deep ML fusion model based on Vision Transformer (ViT) and tabular data was developed for the prediction of the MGD normalized conversion factor CF(DgN). A mean breast thickness of 61.37 mm and a mean MGD of 1.53 mGy (0.55–6.33 mGy) were calculated using this dataset. Regarding tabular data, the artificial neural network (ANN) sequential models outperformed other linear and tree-based models. The ViT deep ML fusion model was tested with three configuration versions differing on the number of features included. A comparison of the three versions revealed that the version with six features achieved the best overall predictor performance. This study demonstrates that ML and deep ML can effectively predict the MGD using dosimetric tabular data and mammography DICOM images. The use of ML with tabular data extracted from DICOM images can be further strengthened by incorporating larger and more diverse datasets. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

22 pages, 700 KB  
Article
Cross-Layer Resource Optimization for Ultra-Low-Power TinyML Inference on ARM Cortex-M Microcontrollers
by Abdulaziz G. Alanazi, Haifa A. Alanazi and Nasser S. Albalawi
Electronics 2026, 15(13), 2918; https://doi.org/10.3390/electronics15132918 - 3 Jul 2026
Viewed by 502
Abstract
Running neural networks on battery-powered Internet of Things (IoT) sensor nodes is difficult because flash memory, SRAM, latency, and energy per inference are limited at the same time. Existing TinyML co-design methods usually improve model size or memory use, but runtime voltage–frequency control [...] Read more.
Running neural networks on battery-powered Internet of Things (IoT) sensor nodes is difficult because flash memory, SRAM, latency, and energy per inference are limited at the same time. Existing TinyML co-design methods usually improve model size or memory use, but runtime voltage–frequency control is often handled as a separate step. This separation limits energy saving because the power policy does not use the layer-wise compute profile of the final compressed model. We propose the Cross-Layer Resource Optimizer (CLRO), a three-stage resource optimization pipeline for TinyML inference on an ARM Cortex-M7 target. The first stage, Mixed-Precision Aware Pruning and Distillation (MPAD), assigns per-layer bit widths and pruning ratios using calibration-set sensitivity scores. The second stage, consisting of the Activation Lifetime-Aware Tensor Scheduler (ALTS), uses the compressed graph to find an execution order that reduces peak live static random-access memory (SRAM). The third stage, Reinforcement Learning-Based Dynamic Voltage and Frequency Scaling (DVFS-RL), trains a tabular Q-learning policy from the multiply–accumulate (MAC) utilization profile of the compressed and scheduled model. The learned voltage–frequency policy is stored as a small flash lookup table, so it adds no runtime decision cost during inference. We evaluate the CLRO on all four MLPerf Tiny tasks using an STM32H743ZI microcontroller with 512 kB SRAM and 2 MB flash. The CLRO reaches 91.7% image classification accuracy, 95.4% keyword-spotting accuracy, 89.6% visual wake words accuracy, and 0.913 anomaly detection AUC. The final deployment uses 198 kB flash and 174 kB peak SRAM, with 387 μJ energy per inference and 38 ms latency. Compared with the MCUNet baseline, the CLRO reduces energy by 58.1% and peak SRAM by 39% while keeping the same accuracy level. Full article
Show Figures

Figure 1

17 pages, 807 KB  
Article
Adaptive A-Semilogarithmic Gradient Quantization for Efficient Deep Neural Network Training
by Stefan Panić, Milan Dubljanin, Milan Savić and Marko Smilić
Algorithms 2026, 19(7), 521; https://doi.org/10.3390/a19070521 - 29 Jun 2026
Viewed by 318
Abstract
This paper introduces an adaptive A-semilogarithmic gradient quantization framework aimed at reducing memory overhead and computational complexity during the training of deep neural networks. The approach employs a semilogarithmic companding function parameterized by a dynamically adjusted scaling factor A, which evolves [...] Read more.
This paper introduces an adaptive A-semilogarithmic gradient quantization framework aimed at reducing memory overhead and computational complexity during the training of deep neural networks. The approach employs a semilogarithmic companding function parameterized by a dynamically adjusted scaling factor A, which evolves in response to the statistical properties of gradients throughout the training process. Two distinct quantization strategies are proposed and evaluated: The switching piecewise A-quantizer, which adaptively toggles between low-bit uniform and high-bit semilogarithmic quantization according to an exponentially weighted moving-average (EMA) estimate of gradient variance; and the hybrid A-quantizer, which statically partitions the gradient domain, applying uniform quantization in low-magnitude regions and semilogarithmic companding in high-magnitude regions. The proposed methods are empirically evaluated on both multilayer perceptron (MLP) and convolutional neural network (CNN) architectures using tabular and image-classification benchmarks, including DCCC, CIFAR-10, CIFAR-100, and ImageNet. Quantitative results demonstrate that both models achieve comparable classification accuracy to full-precision (FP32) baselines while significantly reducing gradient reconstruction error. Notably, the hybrid A-quantizer consistently yields better validation accuracy, reduced RMSE, and improved convergence behavior relative to its switching counterpart. These findings underscore the effectiveness of hybrid semilogarithmic quantization as a robust and efficient solution for training deep models in resource-constrained or bandwidth-limited environments, with strong potential for scalable deployment across diverse hardware platforms. Full article
(This article belongs to the Special Issue Deep Neural Networks and Optimization Algorithms (2nd Edition))
Show Figures

Figure 1

25 pages, 4672 KB  
Article
Data-Efficient and Explainable Multimodal Survival Prediction in NSCLC Using Deep Image Embeddings, Clinical Variables, and Gradient-Boosted Trees
by Sevim Sahin and Adil Gursel Karacor
Diagnostics 2026, 16(12), 1941; https://doi.org/10.3390/diagnostics16121941 - 22 Jun 2026
Viewed by 482
Abstract
Background/Objectives: Survival prediction in non-small cell lung cancer (NSCLC) remains challenging, particularly in limited-sample settings where end-to-end deep learning models may suffer from limited generalization. This study aimed to develop a data-efficient, multimodal, and explainable framework integrating computed tomography (CT)-derived imaging information with [...] Read more.
Background/Objectives: Survival prediction in non-small cell lung cancer (NSCLC) remains challenging, particularly in limited-sample settings where end-to-end deep learning models may suffer from limited generalization. This study aimed to develop a data-efficient, multimodal, and explainable framework integrating computed tomography (CT)-derived imaging information with clinical variables for NSCLC survival prediction. Methods: CT images, tumor segmentations, and clinical data from the publicly available NSCLC Radiomics (LUNG1) dataset (377 patients) were used. Tumor-focused regions were extracted using segmentation masks, and pretrained RadImageNet-InceptionV3 embeddings were obtained from the largest tumor-containing slice and neighboring-slice summaries. Deep imaging embeddings, engineered imaging features, and clinical variables were fused into a unified tabular representation. To improve robustness under limited-sample conditions, feature blocks were compressed using principal component analysis. CatBoost, XGBoost, and LightGBM models were trained on a development set and evaluated on a strictly held-out final validation set. Results: In three-class survival stratification, assigning censored/non-event patients to the upper survival group produced the strongest ordinal prognostic performance. Under the EX_PLUS_NON_EX_TOP setting, CatBoost achieved the best holdout score-based class C-index of 0.655. In continuous survival regression, LightGBM achieved the best holdout event-patient C-index of 0.576. Clinical variables provided the dominant prognostic signal, while compact deep image embeddings contributed complementary information, particularly in separating short- and long-survival groups. SHAP analysis confirmed contributions from both clinical and image-derived features. Conclusions: The proposed framework provides a proof-of-concept demonstration of a data-efficient and explainable image-to-tabular approach for NSCLC survival prediction under strict internal holdout validation. The results suggest that pretrained CT embeddings, clinical variables, gradient-boosted trees, and SHAP-based interpretation can be combined in a feasible, limited-sample survival modeling pipeline, while external validation remains necessary before clinical translation. Full article
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)
Show Figures

Figure 1

25 pages, 1601 KB  
Article
A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction
by Ronish Shrestha, Md Masud Rana, Bo Sun, Frank Sun, Helen Lou and Alek Hutson
Sensors 2026, 26(12), 3804; https://doi.org/10.3390/s26123804 - 15 Jun 2026
Viewed by 393
Abstract
In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module [...] Read more.
In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R2. The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests. Full article
(This article belongs to the Special Issue AI-Empowered Internet of Things)
Show Figures

Figure 1

23 pages, 5972 KB  
Article
AI-Based Prediction of Post-ERCP Pancreatitis: A Comparative Study Using Tabular, Image, and Multimodal Data
by Anum Jamil, Waseemullah Nazir, Abeer Altaf and Saad Khalid Niaz
Diagnostics 2026, 16(12), 1824; https://doi.org/10.3390/diagnostics16121824 - 12 Jun 2026
Viewed by 469
Abstract
Background/Objectives: Post-Endoscopic Retrograde Cholangiopancreatography Pancreatitis (PEP) is a clinically significant complication of ERCP, occurring in approximately 2–10% of general cases and at higher rates in high-risk patients. Early prediction of PEP risk may support timely intervention and improved patient management. This retrospective [...] Read more.
Background/Objectives: Post-Endoscopic Retrograde Cholangiopancreatography Pancreatitis (PEP) is a clinically significant complication of ERCP, occurring in approximately 2–10% of general cases and at higher rates in high-risk patients. Early prediction of PEP risk may support timely intervention and improved patient management. This retrospective single-center study comparatively evaluated tabular clinical data, endoscopic image data, and multimodal fusion approaches for PEP prediction. Methods: Retrospective data collected from the Sindh Institute of Advanced Endoscopy and Gastroenterology were analyzed using machine learning and deep learning techniques. XGBoost(version 3.2.0) was applied to tabular clinical data, while EfficientNet-B0, ResNet50, and DenseNet201 were used for endoscopic image analysis. A multimodal contrastive learning (MMCL)-based framework combining ResNet50 image features with multilayer perceptron (MLP)-based tabular features was additionally implemented for binary PEP prediction. Class imbalance mitigation techniques, including data augmentation and balancing strategies, were applied during training. Model performance was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC), sensitivity, F1-score, and precision. SHAP analysis was performed to identify important predictive features. Results: The tabular XGBoost model achieved the best predictive performance with an AUC of 0.95 and a sensitivity of 0.50, while five-fold cross-validation yielded an AUC of 0.79 and a sensitivity of 0.48. Among image-based models, ResNet50 achieved the highest performance, with an AUC of 0.76 and a sensitivity of 0.40. The multimodal model achieved an AUC of 0.57 and a sensitivity of 0.20. SHAP analysis identified cannulation time, ampulla type, and age as prominent features associated with PEP prediction. Conclusions: This exploratory study suggests that structured clinical data currently provide stronger predictive signals for PEP prediction than the available image and multimodal data within this limited cohort. The relatively low occurrence of PEP contributed to class imbalance despite mitigation strategies. Future multicenter studies with larger datasets, improved image availability, synthetic data generation, and advanced multimodal fusion techniques may improve predictive performance and clinical applicability. Full article
(This article belongs to the Section Medical Imaging and Theranostics)
Show Figures

Graphical abstract

Back to TopTop