Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (887)

Search Parameters:
Keywords = Vision Transformers (ViT)

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
17 pages, 6389 KB  
Article
An Improved Deepfake Detection Approach Using Hybrid Architecture Based on EfficientNet and Vision Transformer
by Omar Banimelhem and Abeer O. Alsharu
J. Imaging 2026, 12(9), 427; https://doi.org/10.3390/jimaging12090427 - 9 Sep 2026
Abstract
The rapid spread of deepfake content poses a significant threat to the credibility of digital multimedia, creating an urgent need for accurate and robust detection methods. This paper proposes a hybrid deep learning framework that combines EfficientNet-B0 and Vision Transformer (ViT-B/16) through feature [...] Read more.
The rapid spread of deepfake content poses a significant threat to the credibility of digital multimedia, creating an urgent need for accurate and robust detection methods. This paper proposes a hybrid deep learning framework that combines EfficientNet-B0 and Vision Transformer (ViT-B/16) through feature concatenation to exploit both local spatial representations and global contextual dependencies for image-level deepfake detection. The proposed model was trained and evaluated on the deepfake and real images dataset. Experimental results demonstrate that the proposed framework achieves an accuracy of 0.9870, an F1-score of 0.9871, and an AUC of 0.9990. Additional cross-dataset evaluation on the CelebDF-v2 image dataset and robustness experiments under common image degradations further demonstrates the strong generalization capability and practical applicability of the proposed approach. These results confirm that integrating CNN-based and Transformer-based feature extraction provides an effective and reliable solution for deepfake image detection. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

17 pages, 1752 KB  
Article
Classification of Oral Squamous Cell Carcinoma from Histopathological Images Using a Hybrid Deep Learning Model
by Furkan Talo and Ahmet Bedri Ozer
Diagnostics 2026, 16(18), 2885; https://doi.org/10.3390/diagnostics16182885 - 8 Sep 2026
Viewed by 123
Abstract
Background/Objectives: Oral squamous cell carcinoma (OSCC) has high mortality rates and leads to serious health problems when diagnosed late. This situation is considered a public health problem. Histopathological examination, which is an important point in the diagnosis of the disease, is a [...] Read more.
Background/Objectives: Oral squamous cell carcinoma (OSCC) has high mortality rates and leads to serious health problems when diagnosed late. This situation is considered a public health problem. Histopathological examination, which is an important point in the diagnosis of the disease, is a time-consuming and manual process that requires expertise. Methods: This study presents a deep learning-based approach for the automated classification of normal oral epithelium and oral squamous cell carcinoma (OSCC) from histopathological images. The performance of state-of-the-art architectures such as Vision Transformer (ViT), CLIP, ConvNextV2, Deit, and Dinov2 was comparatively analyzed. Based on the results, a hybrid architecture combining the strengths of the models is proposed. Results: In experiments conducted on a dataset of 696 histopathological images, the ConvNextV2 and ViTL16 architectures stood out among the basic models with accuracy rates around 84%. However, the most significant contribution of this study is the proposed method, which combines the global context capability of Transformer-based models with the local feature extraction power of CNN-based models using the Efficient Channel Attention (ECA)—Gated Features model. This hybrid model, created by integrating the ViTL16, Dinov2, and ConvNextV2 architectures, achieved 89.95% accuracy, 89.82% F1 score, and 89.95% sensitivity with a KNN classifier, outperforming the baseline models in the literature. Conclusions: The results obtained demonstrate that the fusion of multiple architectures increases diagnostic reliability in medical image analysis and can assist pathologists as a decision support mechanism. Full article
Show Figures

Figure 1

25 pages, 2308 KB  
Article
Comparative Analysis of CNN and Transformer Models for Multi-Class Diabetic Retinopathy Grading Using Fundus Images
by Maha A. Thafar
Diagnostics 2026, 16(17), 2882; https://doi.org/10.3390/diagnostics16172882 - 7 Sep 2026
Viewed by 216
Abstract
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a [...] Read more.
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a unified experimental framework remain limited. This study systematically compares representative standard and lightweight CNN- and transformer-based architectures for multi-class DR grading. Methods: Six ImageNet-pretrained deep-learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer (ViT), Swin-Tiny, and Swin Transformer, were evaluated on the APTOS 2019 retinal fundus image dataset under a unified experimental configuration with consistent preprocessing, data augmentation, training, and evaluation settings. All models were fine-tuned and evaluated independently over five runs with different random seeds. Their performance was assessed using accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), Quadratic Weighted Kappa (QWK), per-class analysis, computational efficiency, and statistical analysis. Results: Transformer-based models generally achieved higher mean classification performance than the evaluated CNN-based models. Swin-Tiny achieved the highest mean accuracy (82.3%), macro F1-score (64.4%), weighted F1-score (82.1%), and QWK (89.8%) across the five runs. Among the CNN-based models, EfficientNet-B0 achieved the strongest overall classification performance, whereas MobileNetV2 provided the lowest computational complexity. The results also highlighted differences in learning behavior and computational requirements across the evaluated architectures. Repeated experiments demonstrated stable performance across different random seeds, supporting the reliability of the proposed evaluation. Conclusions: Overall, this study provides a comprehensive comparison of representative CNN- and transformer-based architectures under consistent experimental settings and offers practical guidance for selecting suitable deep learning models for automated diabetic retinopathy screening. Full article
Show Figures

Figure 1

32 pages, 4867 KB  
Article
Coupled Fourier Neural Operator and Vision Transformer Bottleneck for Parameter-Efficient Brain Tumor Segmentation
by Abel Alejandro Rubín Alvarado, Juan Humberto Sossa Azuela, Humberto de Jesús Ochoa Domínguez, Osslan Osiris Vergara Villegas and Vianey Guadalupe Cruz Sánchez
Mathematics 2026, 14(17), 3242; https://doi.org/10.3390/math14173242 - 7 Sep 2026
Viewed by 213
Abstract
Segmenting brain tumor subregions in multimodal MRI is difficult due to severe class imbalance and scarce annotated data, and current state-of-the-art models require 21.3–31 million parameters to reach a whole-tumor Dice of 0.906–0.921. We propose a parameter-efficient 2D encoder–decoder coupling a Fourier Neural [...] Read more.
Segmenting brain tumor subregions in multimodal MRI is difficult due to severe class imbalance and scarce annotated data, and current state-of-the-art models require 21.3–31 million parameters to reach a whole-tumor Dice of 0.906–0.921. We propose a parameter-efficient 2D encoder–decoder coupling a Fourier Neural Operator (FNO) and a Vision Transformer (ViT), trained with BraTSPipeline, which raises throughput from 257 to 16,040 slices per epoch, and two composite loss functions penalizing false negatives 2.3× more than false positives. On BraTS 2020, the 2D variant (5.6 M parameters) achieves whole-tumor (WT), tumor-core (TC), and enhancing-tumor (ET) Dice of 0.8985, 0.8263, and 0.7568 on the 56-case held-out test set, using 3.8–5.6× fewer parameters than published architectures. Encoding three consecutive axial slices as 12 channels (2.5D, 20.5 M parameters) raises WT to 0.9117, within 0.010 of the best published 2D result on BraTS 2020, Mod-R2AU-Net (WT = 0.921); a 4.6 M-parameter ablation without the ViT reaches WT of 0.9106 and TC of 0.8428, while the ViT adds 0.043 ET Dice. Full article
Show Figures

Figure 1

27 pages, 9393 KB  
Article
Intelligent Monitoring of Shear Damage Evolution at Bonded Sandstone Interfaces Based on ViT and Piezoelectric Ultrasonic Testing
by Jiancheng Liu, Chong Wang, Hongbo Zhang, Zhongshan Zhang, Dong Xu, Hongyu Zou and Zhenbin Xie
Sensors 2026, 26(17), 5586; https://doi.org/10.3390/s26175586 - 2 Sep 2026
Viewed by 232
Abstract
This paper presents a method for monitoring damage evolution at sandstone-binding material interfaces by combining a Vision Transformer (ViT) deep learning model with piezoelectric ultrasonic monitoring. Direct shear tests were conducted on bonded weak sandstone specimens. The results indicate that the damage evolution [...] Read more.
This paper presents a method for monitoring damage evolution at sandstone-binding material interfaces by combining a Vision Transformer (ViT) deep learning model with piezoelectric ultrasonic monitoring. Direct shear tests were conducted on bonded weak sandstone specimens. The results indicate that the damage evolution process can be divided into four stages: initial elastic, compaction and stabilization, crack propagation and coalescence, and frictional sliding and interlocking. Ultrasonic signals acquired during loading reveal that interface damage evolution is in good agreement with the time-domain waveforms, Continuous Wavelet Transform (CWT) time–frequency spectra, and wavelet packet energy. Based on the ViT-Small/16 backbone, a ViT model with adaptive frequency-feature extraction was developed for small-sample and cross-specimen interfacial damage-stage identification, using the Stage I health observations of each specimen prior to loading as the reference. Results from five repeated runs with different random seeds show that the method achieved an accuracy of 93.23% ± 0.72% and an F1-score of 92.32% ± 0.90%, demonstrating favorable recognition performance and stability under the current condition. This study offers insights into the damage monitoring and subsequent warning of similar binary interfaces in tunnel engineering, geotechnical engineering, and stone cultural heritage conservation. Full article
(This article belongs to the Special Issue Sensing Techniques for Intelligent Tunnel Construction)
Show Figures

Figure 1

35 pages, 15908 KB  
Article
HIFU Tissue Degeneration Classification Based on Multifractal Detrending Fluctuation Analysis and Vision Transformer
by Hu Dong, Xin Tong and Gang Liu
Fractal Fract. 2026, 10(9), 609; https://doi.org/10.3390/fractalfract10090609 - 1 Sep 2026
Viewed by 180
Abstract
For high-intensity focused ultrasound (HIFU) thermal ablation to be safe and effective, real-time, high-precision monitoring of tissue coagulative necrosis is essential. However, decoding the ultra-long, non-stationary radio frequency (RF) echoes produced during tissue phase transitions usually results in severe feature aliasing and high [...] Read more.
For high-intensity focused ultrasound (HIFU) thermal ablation to be safe and effective, real-time, high-precision monitoring of tissue coagulative necrosis is essential. However, decoding the ultra-long, non-stationary radio frequency (RF) echoes produced during tissue phase transitions usually results in severe feature aliasing and high computational costs for conventional deep learning models. This paper suggests a highly interpretable, asymmetric classification framework that combines a lightweight Vision Transformer (ViT) with multifractal detrended fluctuation analysis (MFDFA) in order to overcome this obstacle. In terms of methodology, ViT may independently capture cross-scale thermodynamic dependencies without local inductive biases by using MFDFA as a physical prior to compress 1D RF sequences into dense 2D fractal tensors. This method greatly improved the algorithmic recognition of the extremely elusive “partially degenerated” transient state, achieving a strong 96.5% classification accuracy when validated on an ex vivo pig liver dataset. Furthermore, by firmly attaching its classifications to the macroscopic statistical correlates of the acoustic scattering process, the model achieves great decision transparency instead of functioning as an opaque black box. Importantly, this MFDFA-ViT architecture provides an ideal accuracy–latency trade-off with only 3.45M parameters and an end-to-end inference latency of 20.6 ms. This offers a real-time, intelligent monitoring paradigm that is highly deployable and specifically designed for upcoming clinical HIFU applications. Full article
Show Figures

Figure 1

26 pages, 14705 KB  
Article
Federated Convolutional Transformer Network for Privacy-Preserving Photovoltaic Fault Detection in Distributed Solar Power Systems
by Priyanka Vyas, Sheetal U. Bhandari and Pramod R. Sonawane
AI 2026, 7(9), 338; https://doi.org/10.3390/ai7090338 - 31 Aug 2026
Viewed by 247
Abstract
Solar energy makes a considerable contribution to global power generation, necessitating photovoltaic (PV) fault detection to ensure high yields in solar power systems. In real-world solar installations, operational data are geographically dispersed, heterogeneous, and sensitive, which imposes privacy restrictions. Existing PV fault-detection techniques [...] Read more.
Solar energy makes a considerable contribution to global power generation, necessitating photovoltaic (PV) fault detection to ensure high yields in solar power systems. In real-world solar installations, operational data are geographically dispersed, heterogeneous, and sensitive, which imposes privacy restrictions. Existing PV fault-detection techniques have relied on centralized training, requiring raw image data from multiple plants to be collected at a single server, which has led to privacy risks, bias from heterogeneous datasets, and scalability issues. To overcome these challenges, this research provides a novel FL-based PV fault-detection model called Federated Convolutional Transformer Network (Fed-CVTNet), which combines Convolutional Neural Network (CNN) and Vision Transformer (ViT) architecture within the FL framework. Initially, the raw images are pre-processed to enhance the input quality, and the Region of Interest (ROI) is identified via a pretrained YOLO model. Then, the proposed Fed-CVTNet facilitates networked learning among many geographically dispersed clients by only sharing model updates using federated averaging (FedAvg). The experimental findings illustrate that the proposed FL model shows substantial quality improvements compared to the customized CNN-ViT models trained on a dataset of 5600 images and the CNN-ViT models that lack federated aggregation. The highest accuracy achieved by the proposed technique is 98.99%; the proposed Fed-CVTNet has better sensitivity (98.15%) and specificity (99.21%), and much lower false positive and false negative rates than its centralized counterparts. The comparative analysis establishes that federated weight aggregation outperforms centralized baseline models by 2.3% and is effective in reducing the data privacy risk. Full article
Show Figures

Figure 1

24 pages, 17179 KB  
Article
SmartFire Vision: An Attention-Pruned Hybrid Vision Transformer and Detection Transformer Framework for Accurate, Efficient, and Real-Time Fire and Smoke Detection in Smart City Video Surveillance
by Muhammad Azhar, Muhammad Arman, Asma Iqbal, Adeen Amjad and Deshinta Arrova Dewi
Information 2026, 17(9), 845; https://doi.org/10.3390/info17090845 - 31 Aug 2026
Viewed by 195
Abstract
Fire incidents can lead to significant destruction of lives and property, especially in urban and smart cities, and pose a great risk worldwide. Existing fire and smoke detection systems are often inadequate for detecting the location of a fire, assessing the speed of [...] Read more.
Fire incidents can lead to significant destruction of lives and property, especially in urban and smart cities, and pose a great risk worldwide. Existing fire and smoke detection systems are often inadequate for detecting the location of a fire, assessing the speed of its spread, and providing real-time alerts that can be acted upon quickly. This study proposes a method termed SmartFire Vision, which uses a hybrid deep learning framework consisting of an Efficient Vision Transformer (E-ViT) and a Detection Transformer (DETR) for real-time fire and smoke detection from video sequences. A major contribution of this study is the integration of a new Removing Inefficient Attention Heads (RIAH) pruning strategy to reduce the computational overhead and maintain a global context in the ViT encoder. The E-ViT and DETR feature representations were fused and passed to a fully connected classification head enhanced with a probabilistic thresholding function and an integrated alarm system. The proposed model was trained and evaluated using the FURG fire benchmark dataset, which comprises 28,022 annotated frames. The proposed model achieved an overall accuracy of 85.40%, precision of 85.33%, recall of 85.43%, and F1-score of 85.35%, surpassing the current state-of-the-art methods. The SmartFire Vision framework provides a highly capable and computationally efficient means of fire detection, is particularly beneficial for CCTV-based smart city surveillance, and shows promising computational efficiency on desktop-class GPUs, though dedicated edge-hardware validation remains a direction for future work. Full article
Show Figures

Graphical abstract

26 pages, 8441 KB  
Article
Explainable Superpixel-Guided Graph Vision Transformer for Hyperspectral Image Analysis
by Jieli Chen, Kah Phooi Seng, Chee Shen Lim, Li-Minn Ang and Jeremy Smith
Sensors 2026, 26(17), 5513; https://doi.org/10.3390/s26175513 - 31 Aug 2026
Viewed by 188
Abstract
Hyperspectral imaging provides rich spectral–spatial information for fine-grained material discrimination, but effective and interpretable modeling remains challenging because land-cover regions often have irregular spatial structures and class-specific spectral responses. Conventional methods typically rely on fixed grid patches or local neighborhoods, which may not [...] Read more.
Hyperspectral imaging provides rich spectral–spatial information for fine-grained material discrimination, but effective and interpretable modeling remains challenging because land-cover regions often have irregular spatial structures and class-specific spectral responses. Conventional methods typically rely on fixed grid patches or local neighborhoods, which may not align with natural object boundaries, whereas pure superpixel or graph models may lose fine pixel-level details. This paper proposes an explainable superpixel-guided graph vision transformer (ESG-ViT) for hyperspectral image classification and analysis. The proposed framework contains two complementary branches: a graph superpixel vision transformer (GS-ViT) that represents hyperspectral scenes as adaptive superpixel tokens and injects graph topology into self-attention, and a windowed pixel vision transformer (WP-ViT) that preserves dense local spectral–spatial details through efficient local attention. The two representations are adaptively fused for pixel-wise classification. To support interpretability, the model further derives class-wise superpixel relevance maps and spectral channel importance from gradient responses, revealing both the spatial regions and wavelength channels that contribute to each category. Experiments on multiple benchmark hyperspectral datasets demonstrate that the proposed method improves classification accuracy while producing clearer, more human-aligned explanations. Visualization results show that the model highlights meaningful class-related superpixel regions and assigns distinct spectral-channel importance patterns to different classes. These results indicate that the proposed framework provides an accurate and explainable alternative to conventional patch-based transformers for hyperspectral image analysis. Full article
Show Figures

Figure 1

10 pages, 2178 KB  
Proceeding Paper
Real-Time Bacterial Colony Count Detection and Classification Using Computer Vision
by Vasugi Ramdass, Dhilip Kumar Venkatesan, Oana Geman and Roxana Toderean
Eng. Proc. 2026, 148(1), 48; https://doi.org/10.3390/engproc2026148048 - 28 Aug 2026
Viewed by 167
Abstract
Counting bacterial colonies by hand is one of the most common tasks in microbiology, but it is slow and often yields different results depending on who does the counting. This paper presents an automated system for detecting and classifying bacterial colonies from Petri [...] Read more.
Counting bacterial colonies by hand is one of the most common tasks in microbiology, but it is slow and often yields different results depending on who does the counting. This paper presents an automated system for detecting and classifying bacterial colonies from Petri plate images. We tested four models: Support Vector Machine (SVM), Artificial Neural Network (ANN), Convolutional Neural Network (CNN), and a hybrid CNN+Vision Transformer (CNN+ViT). The system uses OpenCV for image preprocessing and extracts features such as area, perimeter, and circularity. Colonies are then classified by size and health condition. Confusion matrix results show that SVM and ANN achieved the highest overall accuracy (95.6%), while CNN+ViT eliminated false positives entirely. A Streamlit-based web interface allows real-time colony analysis directly in laboratory environments. Full article
Show Figures

Figure 1

23 pages, 3265 KB  
Article
A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison
by Binyu Guo, Laibao Yu, Yiming Yang and Chunzhi Wang
Appl. Sci. 2026, 16(17), 8496; https://doi.org/10.3390/app16178496 - 26 Aug 2026
Viewed by 184
Abstract
As a key pixel-level analysis technology, semantic segmentation is widely deployed in autonomous driving and medical imaging. Fully supervised segmentation relies on labour-intensive pixel-wise annotations, so weakly supervised semantic segmentation (WSSS) with only image-level labels has attracted wide attention. Existing Vision Transformer (ViT)-based [...] Read more.
As a key pixel-level analysis technology, semantic segmentation is widely deployed in autonomous driving and medical imaging. Fully supervised segmentation relies on labour-intensive pixel-wise annotations, so weakly supervised semantic segmentation (WSSS) with only image-level labels has attracted wide attention. Existing Vision Transformer (ViT)-based WSSS methods suffer from two critical limitations: ViT’s global self-attention mechanism leads to insensitivity to local small target features and incomplete foreground activation; its class-agnostic attention maps frequently misactivate background regions as foreground objects, introducing heavy noise. To tackle these two issues, this paper proposes a single-stage weakly supervised segmentation algorithm based on local–global class labelling comparison. First, we design a local–global class labelling comparison (LTG) module. By feeding both original images and randomly cropped local patches into ViT, we adopt InfoNCE contrastive loss to align local class tokens with global class tokens, enhancing the feature integrity of local target regions and suppressing background false activation. Second, a class-aware stimulus module (CSM) is embedded into ViT’s multi-head attention branch. It injects category semantic constraints into self-attention to generate class-aware attention maps, guiding the model to focus on real foreground targets and reduce background interference. Finally, we construct a feature fusion class-aware activation map (FFCAM) by fusing ViT global output features and CSM class-aware attention features to generate high-quality pseudo-labels for segmentation training. Extensive experiments are conducted on PASCAL VOC 2012 and MS COCO 2014 datasets. Our method achieves 78.5% mIoU on the validation set of PASCAL VOC 2012 and 50.9% mIoU on MS COCO 2014, showing competitive numerical performance among the compared single-stage ViT-based WSSS approaches. Ablation experiments verify the independent and joint effectiveness of LTG, CSM and FFCAM. The proposed method effectively improves the completeness of target activation regions and suppresses background noise, and it maintains strong generalization for slender, small and texture-sparse objects. In future work, we will further lightweight the ViT backbone to reduce computational overhead for embedded deployment. Full article
Show Figures

Figure 1

37 pages, 3015 KB  
Article
Deepfake Detection via Frequency-Aware Vision Transformer and Bidirectional Cross-Attention Fusion with Post-Processing Robustness
by Wasin Alkishri, Shahid Kamal and Jabar Yousif
Information 2026, 17(9), 819; https://doi.org/10.3390/info17090819 - 26 Aug 2026
Viewed by 1038
Abstract
Today, the use of increasingly ubiquitous synthetic media, or ‘deepfakes’, has become a risk to online trust, information integrity and individual security and is being created by artificial intelligence (AI). The current approaches are mainly based on either spatial features of CNNs or [...] Read more.
Today, the use of increasingly ubiquitous synthetic media, or ‘deepfakes’, has become a risk to online trust, information integrity and individual security and is being created by artificial intelligence (AI). The current approaches are mainly based on either spatial features of CNNs or high-level semantic representations of Vision Transformer; both have major drawbacks in effectively leveraging multi-domain forensic cues. This paper presents FAViT (Frequency-Aware Vision Transformer), a hybrid architecture capable of jointly utilizing spatial- and frequency-domain forensic information by the means of a bidirectional cross-attention fusion scheme. We use an 11-channel forensic tensor in each face image (including per-channel Fast Fourier Transform (FFT) magnitude maps, Discrete Wavelet Transform (DWT) sub-bands, channel noise residual maps, Sobel gradient magnitude and channels of Error Level Analysis (ELA)). A Frequency Branch CNN processes this multi-domain tensor and the original RGB image is encoded with a pretrained ViT-B/16 spatial branch. The two streams are combined through the bidirectional cross-attention which allows the model to localize both spatial and spectral manipulation artifacts. We also present an adversarial cleaning simulation pipeline which partitions the training process with five post-processing attack methods, namely GFPGAN neural face restoration, learned autoencoder cleaning, etc., to increase resistance to real-world forensic defenses. Tests of FaceForensics++ C23 (7926 images, consisting of four manipulation types) show that FAViT attains F1-score of 86.22, AUC-ROC of 94.26 and accuracy of 85.55 on the held-out test set. The strength analysis of 21 attack conditions shows that the max degradation in AUC is 30.3, with specific strengths in GFPGAN restoration (AUC = 98.51). Robustness is evaluated based on 21 post-processing attack cases that include JPEG compression, Gaussian blurring, down-sampling, and GFDGAN neural-based restoration; it should be noted that robustness against gradient-based adaptive attacks requires additional attention. Testing on the CIFAKE and Celeb-DF v2 datasets reveals some limitations of domain generalization. Full article
(This article belongs to the Special Issue Artificial Intelligence for Signal, Image and Video Processing)
Show Figures

Graphical abstract

14 pages, 2349 KB  
Article
Dynamic LoRA Fine-Tuning of DINOv3 for Multi-Component Pasture Biomass Estimation
by Shikha Sen, Nischay Dhankhar and Akram Bayat
Sensors 2026, 26(16), 5285; https://doi.org/10.3390/s26165285 - 20 Aug 2026
Viewed by 409
Abstract
Accurate estimation of pasture biomass components from imagery is essential for sustainable grazing management and precision agriculture. Conventional methods such as destructive harvesting, rising plate meters, and remote sensing are limited by scalability, reliability, or the ability to disaggregate biomass by species. We [...] Read more.
Accurate estimation of pasture biomass components from imagery is essential for sustainable grazing management and precision agriculture. Conventional methods such as destructive harvesting, rising plate meters, and remote sensing are limited by scalability, reliability, or the ability to disaggregate biomass by species. We propose a parameter-efficient multi-output regression framework predicting five biomass components (dry green, dry dead, dry clover, green dry matter, and total dry biomass) from high-resolution top-view pasture images. It employs a pretrained DINOv3 Vision Transformer backbone adapted via a dynamic, depth-aware Low-Rank Adaptation (LoRA) strategy, in which the adaptation rank and scaling factor increase exponentially with layer depth: early layers encoding generic visual primitives are minimally perturbed, while deeper layers receive stronger task-specific adaptation. This schedule is effective in low-data regimes, where uniform adaptation or full fine-tuning overfits. To handle rectangular image geometry, each image is split into two square halves processed as a dual-view stream with a contrastive alignment loss. The system ensembles ViT-Large and ViT-Huge backbones with test-time augmentation across five-fold cross-validation. On the CSIRO Image2Biomass benchmark, the full pipeline attains a cross-validated weighted R-squared of 0.81, indicating that depth-aware, parameter-efficient adaptation of large vision models is effective for non-invasive biomass estimation under data scarcity. Full article
Show Figures

Figure 1

21 pages, 1161 KB  
Article
Uncertainty-Aware AI-Assisted Diabetic Retinopathy Grading from Fundus Images with Ordinal Conformal Prediction
by Umar Hasan, Muhammad Ali Nayeem and Turki G. Alghamdi
Diagnostics 2026, 16(16), 2645; https://doi.org/10.3390/diagnostics16162645 - 19 Aug 2026
Viewed by 308
Abstract
Background: Artificial intelligence (AI) systems for diabetic retinopathy (DR) grading require reliable uncertainty estimates when applied to fundus images outside the development dataset. We evaluated whether conformal prediction can provide structured set-valued outputs and whether internal uncertainty calibration remains reliable during external evaluation. [...] Read more.
Background: Artificial intelligence (AI) systems for diabetic retinopathy (DR) grading require reliable uncertainty estimates when applied to fundus images outside the development dataset. We evaluated whether conformal prediction can provide structured set-valued outputs and whether internal uncertainty calibration remains reliable during external evaluation. Methods: EfficientNet-B0, ResNet-50, and Vision Transformer (ViT-Base) classifiers were trained on APTOS 2019. Split-conformal predictors were calibrated exclusively on held-out APTOS images using three categorical scores, LAC, APS, and RAPS, and an ordinal score restricted to adjacent severity grades. Performance was assessed internally on APTOS and externally on IDRiD, with an independent five-seed ViT replication extending evaluation to Messidor-2, at target coverages of 90% and 95%. Results: Coverage was approximately nominal internally but decreased on both external datasets. On IDRiD, the largest deficit occurred for severe DR (grade 3). The ordinal method produced contiguous intervals in 100% of cases and achieved the highest grade-3 coverage in every tested backbone–risk configuration. For ResNet-50 at 95% target coverage, grade-3 coverage increased from 0.750 with APS to 0.945 with the ordinal method, while average set size increased from 2.89 to 3.03. In the independent ViT replication on Messidor-2, ordinal marginal coverage exceeded APS at both targets (0.676 versus 0.632 and 0.741 versus 0.714). Conclusions: Internal calibration did not ensure reliable class-specific uncertainty during external evaluation. Ordinal prediction sets improved structural coherence and mitigated severe-grade undercoverage, but did not restore formal coverage guarantees after dataset shift. Full article
(This article belongs to the Special Issue Artificial Intelligence in Eye Disease, Fifth Edition)
Show Figures

Figure 1

37 pages, 32962 KB  
Article
FedSwin-LHTP: Structure-Aware Hessian-Inspired Token Pruning for Efficient Federated Skin Lesion Classification
by Muhammad Awais and Riaz Hussain Junejo
Diagnostics 2026, 16(16), 2637; https://doi.org/10.3390/diagnostics16162637 - 19 Aug 2026
Viewed by 299
Abstract
Background: Skin cancer encompasses a diverse range of malignancies and remains a significant global health challenge. Accurate machine-learning-assisted diagnosis can substantially improve patient outcomes through early detection and timely clinical intervention. Federated Learning (FL) enables privacy-preserving collaborative model training across multiple healthcare institutions [...] Read more.
Background: Skin cancer encompasses a diverse range of malignancies and remains a significant global health challenge. Accurate machine-learning-assisted diagnosis can substantially improve patient outcomes through early detection and timely clinical intervention. Federated Learning (FL) enables privacy-preserving collaborative model training across multiple healthcare institutions while ensuring that sensitive patient data remain decentralized. However, deploying advanced architectures such as Vision Transformers (ViTs) in clinical environments is challenging due to the high computational demands of self-attention mechanisms. Methods: This work proposes FedSwin-LHTP, an efficient federated learning framework for skin lesion classification that integrates a Swin Transformer backbone with a Lightweight Hessian-Inspired Token Pruning (LHTP) mechanism. LHTP estimates token importance using a second-order Taylor approximation around converged local model parameters to identify less informative patch tokens, enabling the early pruning of redundant representations without explicitly computing the Hessian matrix. Furthermore, the framework incorporates the FedProx optimization objective to mitigate client drift under heterogeneous non-IID data distributions. The proposed framework is evaluated on the HAM10000 and ISIC datasets under realistic non-IID federated settings. Results: Experimental results demonstrate stable convergence, effective knowledge aggregation, and robust diagnostic discrimination across distributed clients. By adaptively pruning approximately 60% of Stage-1 tokens, the proposed framework substantially reduces the computational burden of local transformer processing while maintaining high multiclass classification performance, achieving an accuracy of up to 96.1% on the evaluated datasets. Conclusions: These results highlight the potential of FedSwin-LHTP as a practical, privacy-preserving, and resource-efficient solution for collaborative healthcare intelligence. Full article
Show Figures

Figure 1

Back to TopTop