Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (3,870)

Search Parameters:
Keywords = ResNet50 network

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
25 pages, 33250 KB  
Article
DGCR-Net: Dynamic Graph Contextual Reasoning Network for Semantic Segmentation of Remote Sensing Imagery
by Hong Wang, Kun Gao, Xiaodian Zhang, Zhijia Yang, He Zhang, Zefeng Zhang and Jingyi Wang
Remote Sens. 2026, 18(18), 3069; https://doi.org/10.3390/rs18183069 - 8 Sep 2026
Abstract
Semantic segmentation of remote sensing images is challenging because multi-scale irregular objects in complex scenes often exhibit large intra-class variability, high inter-class similarity, and sparse spatial distributions. These factors hinder accurate boundary delineation and reliable contextual modeling among spatially distant but semantically related [...] Read more.
Semantic segmentation of remote sensing images is challenging because multi-scale irregular objects in complex scenes often exhibit large intra-class variability, high inter-class similarity, and sparse spatial distributions. These factors hinder accurate boundary delineation and reliable contextual modeling among spatially distant but semantically related regions. Considering the capability of graph neural networks in modeling irregular relationships, we propose DGCR-Net, a dynamic graph contextual reasoning network for semantic segmentation of remote sensing imagery. Specifically, DGCR-Net integrates a ResNet18 encoder with a multi-stage decoder composed of cascaded dynamic graph reasoning blocks (DGRBs), which adaptively infer complex contextual dependencies among irregular objects and progressively refine multi-scale semantic representations. A semantic graph adapter (SGA) is incorporated at each skip connection to enhance encoder features and project them into graph-compatible representations, ensuring robust contextual reasoning. Extensive experiments on the Vaihingen, Potsdam, LoveDA, and UAVid datasets demonstrate that DGCR-Net achieves competitive performance, with mIoU scores of 83.4%, 86.5%, 53.9%, and 69.4%, respectively. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

25 pages, 2308 KB  
Article
Comparative Analysis of CNN and Transformer Models for Multi-Class Diabetic Retinopathy Grading Using Fundus Images
by Maha A. Thafar
Diagnostics 2026, 16(17), 2882; https://doi.org/10.3390/diagnostics16172882 - 7 Sep 2026
Abstract
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a [...] Read more.
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a unified experimental framework remain limited. This study systematically compares representative standard and lightweight CNN- and transformer-based architectures for multi-class DR grading. Methods: Six ImageNet-pretrained deep-learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer (ViT), Swin-Tiny, and Swin Transformer, were evaluated on the APTOS 2019 retinal fundus image dataset under a unified experimental configuration with consistent preprocessing, data augmentation, training, and evaluation settings. All models were fine-tuned and evaluated independently over five runs with different random seeds. Their performance was assessed using accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), Quadratic Weighted Kappa (QWK), per-class analysis, computational efficiency, and statistical analysis. Results: Transformer-based models generally achieved higher mean classification performance than the evaluated CNN-based models. Swin-Tiny achieved the highest mean accuracy (82.3%), macro F1-score (64.4%), weighted F1-score (82.1%), and QWK (89.8%) across the five runs. Among the CNN-based models, EfficientNet-B0 achieved the strongest overall classification performance, whereas MobileNetV2 provided the lowest computational complexity. The results also highlighted differences in learning behavior and computational requirements across the evaluated architectures. Repeated experiments demonstrated stable performance across different random seeds, supporting the reliability of the proposed evaluation. Conclusions: Overall, this study provides a comprehensive comparison of representative CNN- and transformer-based architectures under consistent experimental settings and offers practical guidance for selecting suitable deep learning models for automated diabetic retinopathy screening. Full article
Show Figures

Figure 1

42 pages, 6230 KB  
Article
A Study on Multi-Tier Categorical Soil Classification Based on Decoupled Parallel Deep Learning: A Case Study in the Southern Foothills of Qilian Mountains
by Yueyong Pang, Heng Xu, Sen Zou, Liming Zhu, Lizhi Miao and Jieying Zheng
Land 2026, 15(9), 1657; https://doi.org/10.3390/land15091657 - 7 Sep 2026
Abstract
High-precision, multi-tier categorical soil classification faces critical bottlenecks, including the neglect of spatial context by conventional pixel-based models, the error cascade propagation phenomenon in multi-level classification networks, and the disruption of geophysical directional anisotropy by traditional geometric data augmentation. To address these challenges, [...] Read more.
High-precision, multi-tier categorical soil classification faces critical bottlenecks, including the neglect of spatial context by conventional pixel-based models, the error cascade propagation phenomenon in multi-level classification networks, and the disruption of geophysical directional anisotropy by traditional geometric data augmentation. To address these challenges, in this study, we propose a multi-level soil classification model based on MTSC-ResNet-Trans, which organically couples residual convolutional blocks with a 3-layer Transformer encoder to model long-range spatial dependencies. The framework integrates a geospatial-safe data augmentation pipeline to preserve the topological fidelity of absolute geographic coordinates alongside four decoupled parallel multi-task classification heads to substantially suppress inter-level error propagation. Evaluated in the Southern Foothills of Qilian Mountains using 18 environmental covariates, the framework achieves an Overall Accuracy of 0.8931 at the Great Group level under conventional random splitting. Under a distance-stratified spatial evaluation—which isolates the contribution of spatial autocorrelation to accuracy estimates—the framework maintains robust performance, with MTSC-ResNet-Trans consistently outperforming pixel-based baselines (Random Forest) by approximately 3.7 percentage points even at spatial separation distances exceeding 400 m. This protocol transparently decomposes predictive accuracy into a component attributable to spatial proximity and a component reflecting reduced spatial proximity performance. Across the four taxonomic levels, accuracy decay is suppressed to 5.11%. Although spatial-block cross-validation indicates a lower regional extrapolation accuracy (OA = 0.7389 ± 0.0671), the decoupled parallel framework provides an effective and robust baseline for high-resolution regional digital soil mapping. Full article
23 pages, 766 KB  
Article
Lightweight Shoulder Physiotherapy Exercise Recognition via Efficient Channel Attention and Depthwise Separable Residual Networks on Wrist-Worn IMU
by Sakorn Mekruksavanich and Anuchit Jitpattanakul
Computers 2026, 15(9), 595; https://doi.org/10.3390/computers15090595 - 7 Sep 2026
Abstract
Accurate and subject-independent recognition of shoulder physiotherapy exercises from wrist-worn inertial measurement unit (IMU) signals is essential for automated home-based rehabilitation monitoring, yet existing deep learning models are too parameter-intensive to deploy on resource-constrained smartwatch hardware. This paper presents ECA-ResNet1D-Lite, a lightweight one-dimensional [...] Read more.
Accurate and subject-independent recognition of shoulder physiotherapy exercises from wrist-worn inertial measurement unit (IMU) signals is essential for automated home-based rehabilitation monitoring, yet existing deep learning models are too parameter-intensive to deploy on resource-constrained smartwatch hardware. This paper presents ECA-ResNet1D-Lite, a lightweight one-dimensional depthwise separable residual network augmented with efficient channel attention (ECA), trained and evaluated on the SPARS9x dataset comprising six shoulder exercises recorded from 20 subjects using a commercial wrist-worn smartwatch at 50 Hz. Because 50% window overlap allows adjacent windows to share samples, we report three protocols—window-level five-fold, recording-level grouped five-fold, and leave-one-subject-out (LOSO) cross-validation—with model selection performed throughout on validation data disjoint from the test partition. Under LOSO, the primary protocol, the model attains 99.1 ± 1.0% accuracy while requiring only 13,612 parameters and 0.46 M multiply–accumulate operations per window—the highest accuracy and the lowest between-subject dispersion of the seven architectures trained under an identical protocol, ahead of the strongest unconstrained baseline (InceptionTime, 98.8 ± 1.4%) at 36.3× fewer parameters and 213× fewer operations, and ahead of the parameter-efficient designs TinyHAR (97.4 ± 2.5%) and TinierHAR (97.4 ± 2.1%). An ablation over four attention variants (SE, ECA, CBAM, and multi-head self-attention) shows ECA to be the cheapest, adding six parameters (0.04% overhead), while delivering the largest LOSO gain over the same backbone without attention (+0.2 percentage points) and reducing the cross-subject standard deviation from 1.4% to 1.0%. Per-class analysis further reveals that shoulder girdle stabilization is the most challenging exercise under LOSO (F1-score: 96.4%), despite being the most represented class, attributable to its quasi-static, low-amplitude IMU signature. Deployed to an Apple Watch Ultra 2, the model classifies a 4-s window in 0.24 ms (duty cycle 0.012%), establishing that subject-independent shoulder physiotherapy monitoring is computationally feasible on current smartwatch hardware. Full article
Show Figures

Figure 1

15 pages, 5812 KB  
Article
Microfluidic Light-Scattering Imaging Coupled with Deep Learning for Label-Free Single-Cell Classification of Lymphoma Cells
by Linyan Xie, Mengfei Wang, Xijia Luo, Shuoxian Xia, Qiongqiong Ren and Xuezhi Zhou
Biosensors 2026, 16(9), 500; https://doi.org/10.3390/bios16090500 - 7 Sep 2026
Abstract
Accurate classification of lymphoma cell subtypes is essential for disease diagnosis and therapeutic decision-making, yet conventional approaches often rely on fluorescence labeling, labor-intensive sample preparation, and specialized instrumentation, limiting their applicability for rapid, label-free single-cell analysis. Here, we present an AI-assisted microfluidic light-scattering [...] Read more.
Accurate classification of lymphoma cell subtypes is essential for disease diagnosis and therapeutic decision-making, yet conventional approaches often rely on fluorescence labeling, labor-intensive sample preparation, and specialized instrumentation, limiting their applicability for rapid, label-free single-cell analysis. Here, we present an AI-assisted microfluidic light-scattering imaging platform for label-free classification of lymphoma cells. The platform integrates hydrodynamic focusing within a microfluidic chip, continuous acquisition of two-dimensional (2D) light-scattering patterns, automated image preprocessing, and transfer learning based on a pretrained ResNet50 network for intelligent optical feature extraction and classification. Human B lymphoma (Daudi) and T lymphoblastic lymphoma (SUP-T1) cells were used to evaluate the proposed framework. The optical imaging system was first validated using standard microspheres, demonstrating reliable acquisition of light-scattering patterns under continuous-flow conditions. A dataset comprising 800 single-cell scattering patterns was subsequently established and evaluated using stratified five-fold cross-validation. The proposed framework achieved an average classification accuracy of 94.75% with an average area under the receiver operating characteristic (ROC) curve of 0.986. By integrating microfluidic optical biosensing with deep learning, this work enables automated interpretation of intrinsic optical scattering signatures and provides a promising AI-enabled strategy for rapid, label-free lymphoma screening and intelligent healthcare applications. Full article
Show Figures

Figure 1

25 pages, 5957 KB  
Article
A Multi-Scale Fractal Feature Extraction Method for CNN-Based Plant Disease Classification
by Egor Savchenko and Anna Maslovskaya
Mach. Learn. Knowl. Extr. 2026, 8(9), 273; https://doi.org/10.3390/make8090273 - 7 Sep 2026
Abstract
Plant diseases, being a subject of interdisciplinary research, significantly reduce crop yield, quality, and economic returns, while the misidentification of pathogens often leads to ineffective treatments and may harm beneficial organisms and ecosystems. This work develops an approach for robust visual classification of [...] Read more.
Plant diseases, being a subject of interdisciplinary research, significantly reduce crop yield, quality, and economic returns, while the misidentification of pathogens often leads to ineffective treatments and may harm beneficial organisms and ecosystems. This work develops an approach for robust visual classification of plant diseases under limited and heterogeneous data based on multi-scale fractal texture descriptors integrated into a convolutional neural network. The proposed method employs wavelet transform modulus maxima to extract two complementary fractal characteristics, local fractal dimension and singularity spectrum width, from leaf images at several spatial scales. These descriptors form multi-channel fractal maps fed into a fractal attention module (FAM) inserted after the third stage of a ResNet-50 architecture. The FAM learns to emphasize spatial regions where fractal properties are most discriminative, while a parallel branch encodes global fractal statistics into an auxiliary vector combined with backbone features at the final classification layer. Experiments are conducted on a large heterogeneous collection of 11 public plant disease datasets under 5-shot, 50-shot, and full-scale training regimes. The fractal-augmented model raises classification accuracy from 57.06% to 67.73% on 5 shots and from 80.81% to 86.11% on 50 shots, red outperforming the plain ResNet-50 in these settings, converges within 1–2 epochs versus 25–40, and shows markedly better resilience to color distortions, random occlusions, and grayscale conversion in most cases. The generated attention maps provide spatially explicit explanations of the model’s decisions, increasing transparency for practical use. The proposed approach demonstrates that fractal analysis, embedded as a modulating signal inside a deep network, can serve as an efficient and interpretable inductive bias, which is particularly valuable under data scarcity and noisy agricultural imagery. Full article
Show Figures

Figure 1

34 pages, 4038 KB  
Article
Template-Based Digital Surface Reconstruction of Shoe Lasts from Point Clouds
by Philip Azariadis
Algorithms 2026, 19(9), 764; https://doi.org/10.3390/a19090764 - 6 Sep 2026
Abstract
The shoe last is central to footwear design. Modern footwear CAD operates on parametric digital lasts, yet much last geometry—legacy collections and lasts that skilled last makers still sculpt by hand and copy by pantograph turning—exists only as physical models or as point-cloud [...] Read more.
The shoe last is central to footwear design. Modern footwear CAD operates on parametric digital lasts, yet much last geometry—legacy collections and lasts that skilled last makers still sculpt by hand and copy by pantograph turning—exists only as physical models or as point-cloud scans lacking the structured parametric form that footwear CAD requires. This paper presents a complete template-based method for reconstructing a watertight parametric last from a segmented point cloud without intermediate triangulation. The only manual input is three landmark points—for which the system proposes standard positions—and the interactive confirmation of two boundary lines on the digitized last. From these, the method defines four feature points, a median plane, and a four-curve boundary network; all subsequent stages run without user interaction. A curvature-adaptive quadrilateral grid is constructed on the cloud by geodesic tracing and monitor-weighted area-orthogonality relaxation. A periodic Coons tube interpolates the grid and initializes the parameterization for a periodic tensor-product cubic B-spline surface fitted by penalized least squares with cyclic/open difference penalties, exact boundary interpolation, and toe-aware weighting. Cap surfaces close both collar and sole openings, and the model is exported as a watertight B-rep solid. Tests on sixteen industrial lasts using one fixed parameter set produced a mean one-sided deviation of 0.034 mm (RMS 0.058 mm) from the withheld industrial reference meshes in approximately 12 s per last. With synthetic noise at 50 dB SNR, the mean deviation increased by only 0.011 mm. A sampling-density study indicated near-second-order convergence before the control-net reaches an upper limit. The resulting solids import directly into CAD systems and support re-lasting, footwear design, and customization. Full article
(This article belongs to the Collection Algorithms for Computer Vision Applications)
26 pages, 10042 KB  
Article
Unstructured Data Parsing Method Based on Asymmetric Convolution and 3D Attention Residual Networks
by Liping Wang, Pingwen Zheng, Changchun Liu, Dunbing Tang and Zehui Jin
Electronics 2026, 15(17), 4025; https://doi.org/10.3390/electronics15174025 - 6 Sep 2026
Abstract
Efficient parsing of manufacturing process data is a key enabler for the informatization of intelligent manufacturing systems. However, network isolation in aerospace-specific job shops makes large volumes of unstructured shop-floor data, such as handwritten production reports and equipment logs, inaccessible to existing information [...] Read more.
Efficient parsing of manufacturing process data is a key enabler for the informatization of intelligent manufacturing systems. However, network isolation in aerospace-specific job shops makes large volumes of unstructured shop-floor data, such as handwritten production reports and equipment logs, inaccessible to existing information systems. To tackle this, we propose a comprehensive parsing framework that covers data acquisition, parsing, and structured output. For handwritten report parsing, we devise a collaborative pipeline comprising text detection via the Differentiable Binarization Network (DBNet); text recognition using an enhanced Convolutional Recurrent Neural Network (CRNN) that incorporates Asymmetric Convolution (AC) and a Simple Attention Module (SimAM)-based residual module (SimRes, short for SimAM ResNet), referred to as AC-SimRes-CRNN; and table structure extraction via TableMaster. The predicted table cell coordinates, detected text-region coordinates, and recognized text contents are subsequently aggregated to reconstruct complete tables, which are then exported as Excel files. Experiments on the CASIA-HWDB2x and IAM datasets show that AC-SimRes-CRNN achieves an Accurate Rate (AR) of 91.36% and a Correct Rate (CR) of 93.17% on Chinese handwritten text recognition and a Character Error Rate (CER) of 7.85% and a Word Error Rate (WER) of 26.83% on English handwritten text recognition, demonstrating competitive performance against representative methods. Ablation studies validate the contributions of both AC and SimRes. A case study on an aerospace equipment maintenance report further illustrates the component-level feasibility of the proposed workflow. Full article
(This article belongs to the Section Computer Science & Engineering)
Show Figures

Figure 1

23 pages, 5994 KB  
Article
A Transfer Learning and Data Augmentation Approach for Classifying Field Images of Granite Residual Slope Soils
by Zuohui Qin, Can Wang, Xin Zhou, Tengfei Yao, Wei Yin, Huimin Liang and Jian Ou
Algorithms 2026, 19(9), 755; https://doi.org/10.3390/a19090755 - 4 Sep 2026
Viewed by 135
Abstract
Granite residual and slope-wash soils are important disaster-prone geological bodies in the hilly and mountainous areas of Hunan Province, China. Their engineering classification has long relied on manual visual inspection and laboratory testing, which is inefficient and subjective. In this study, an automatic [...] Read more.
Granite residual and slope-wash soils are important disaster-prone geological bodies in the hilly and mountainous areas of Hunan Province, China. Their engineering classification has long relied on manual visual inspection and laboratory testing, which is inefficient and subjective. In this study, an automatic classification method based on deep learning image recognition is proposed for granite residual and slope-wash soils in the mountainous areas of Hunan Province. First, a three-class primary classification scheme was established, comprising residual clay (RNC), residual sandy clay (RNSC), and residual gravelly clay (RNGC), based primarily on the gravel content of particles larger than 2 mm (RNC < 5%, RNSC 5–20%, RNGC > 20%). Second, 7678 geotechnical test records from 21 counties in Hunan Province were collected, and classification labels were assigned through a strategy combining manual verification and automatic inference using Random Forest (5-fold cross-validation macro F1 = 0.913). From approximately 10,096 original field images, 3380 pure soil image patches were retained after segmentation and screening. A dataset of 43,940 samples was then generated through two-stage preprocessing (including denoising and illumination correction) and 13-fold data augmentation. A CNN image classification model was constructed based on a ResNet18 backbone network pre-trained on ImageNet. On the independent test set (6591 images), the primary classification accuracy reached 91.46%, with a macro F1-score of 0.9128; the per-class F1-scores for RNC, RNSC, and RNGC were 0.921, 0.885, and 0.933, respectively. Grad-CAM visualization analysis demonstrated that the model’s attention was primarily focused on soil particle distribution regions rather than non-soil background areas, confirming the effective learning of mixed-grain features. The study shows that the combined application of transfer learning and 13-fold data augmentation can significantly improve classification performance under limited sample size conditions (an improvement of 15.33 percentage points compared to the baseline of 76.13%), demonstrating promising potential for engineering applications. Full article
Show Figures

Figure 1

27 pages, 44141 KB  
Article
A Multidimensional Framework for Diagnosing Streetscape Perception in Historic-District Renewal Using Street-View Imagery and Deep Learning
by Jilong Li, Hongyang Chen, Pan Liao, Liang Sun and Ruxin Liu
Land 2026, 15(9), 1640; https://doi.org/10.3390/land15091640 - 3 Sep 2026
Viewed by 211
Abstract
The renewal of historic districts needs to address human-centered issues, such as cultural expression, spatial experience, and visual comfort, at the street scale. However, existing assessment methods predominantly rely on field surveys, expert judgment, or individual visual indicators, making it difficult to produce [...] Read more.
The renewal of historic districts needs to address human-centered issues, such as cultural expression, spatial experience, and visual comfort, at the street scale. However, existing assessment methods predominantly rely on field surveys, expert judgment, or individual visual indicators, making it difficult to produce reproducible and spatially explicit diagnostic evidence across extensive street networks. This study proposes a multidimensional streetscape perception diagnostic framework that evaluates three dimensions: cultural character recognition (CCR), spatial order (SO), and visual comfort (VC). Taking the Pengcheng Qili historic district in Xuzhou, China, as a case study, the framework integrates street-view imagery, subjective pairwise comparisons, the Bradley–Terry model, and ResNet50-based deep learning prediction. Based on 2243 sampling points and 8693 street-view images, three perception–prediction models were developed and evaluated using five-fold stratified cross-validation, followed by independent external validation using additional historic-district images. The outputs were subsequently mapped onto the street network and examined using spatial statistical analysis and multidimensional profile classification. The results show that the three perception dimensions exhibit distinct spatial patterns and significant spatial clustering. CCR forms localized clusters around historical nodes and heritage-rich areas, whereas SO and VC show clearer corridor-like and network-like patterns. The proposed framework organizes streetscape perception predictions into interpretable multidimensional spatial profiles, thereby providing spatial evidence for conservation-oriented renewal and fine-scale governance of historic districts. Full article
(This article belongs to the Special Issue Big Data-Driven Urban Spatial Perception)
Show Figures

Figure 1

41 pages, 5727 KB  
Article
Synthetic-Data-Augmented Corrosion-Severity Grading of Grounding Connectors: A Colorimetric Benchmark and Kinetics-Aware Ranking
by Junjie Chen, Tao Liu, Zhigao Wang, Jigang Huang, Xinsheng Lan, Lin Zhang, Lutong Yang and Mei Wang
Processes 2026, 14(17), 2833; https://doi.org/10.3390/pr14172833 - 3 Sep 2026
Viewed by 234
Abstract
Corrosion-severity grading of grounding-grid connectors from optical images supports proactive power-infrastructure maintenance. Existing approaches rely on single-time-point, manually thresholded hue–saturation–value (HSV) metrics and static multi-criteria decision-making (MCDM) frameworks that cannot capture corrosion dynamics. In this paper we present a pipeline that (1) defines [...] Read more.
Corrosion-severity grading of grounding-grid connectors from optical images supports proactive power-infrastructure maintenance. Existing approaches rely on single-time-point, manually thresholded hue–saturation–value (HSV) metrics and static multi-criteria decision-making (MCDM) frameworks that cannot capture corrosion dynamics. In this paper we present a pipeline that (1) defines a four-class corrosion grade from an HSV area fraction (Scorr) measured on RGBA optical images, and validates those labels against a baseline-referenced CIEDE2000 metric zero-referenced to each connector’s as-received appearance; (2) generates 240 color-prior-constrained procedural synthetic images from 53 real images across six connector types; (3) fine-tunes a ResNet-18 to estimate corrosion coverage continuously, deriving the reported severity class from that estimate rather than predicting it directly; and (4) fits power-law kinetics C(t) = k·tn to the Scorr time series, propagates bootstrap uncertainty into a Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) framework, and reports kinetics-aware rankings as rank probabilities. The label validation quantifies two limitations of single-threshold HSV grading: a material-color offset that scores an unexposed copper connector at Scorr = 0.442, and insensitivity to achromatic corrosion products covering roughly 80% of the aluminum and galvanized-steel surface. Ablation experiments replicated over five random seeds show that neither contribution claimed from a single run survives replication: synthetic augmentation changes macro-F1 by +0.050 (p = 0.46) under the adopted checkpoint-selection rule and by −0.059 (p = 0.43) under the rule used in the original experiments, and the monotonicity-consistency loss by −0.011 (p = 0.87) and +0.001 (p = 0.99) respectively; the previously reported single-run values of 0.208 and 0.494 are draws from opposite tails of the same seed distributions (0.403 ± 0.140 and 0.344 ± 0.073). The one formulation that improves significantly is the continuous one adopted here, which raises Spearman agreement with the independent metric from 0.316 ± 0.150 to 0.698 ± 0.108 (p = 0.005). Measured against controls, a classifier that never sees the image reaches macro-F1 = 0.425 and, after Holm–Bonferroni correction, no deep configuration is distinguishable from it; none exceeds a one-dimensional linear rule on Scorr (0.664); and under leave-one-material-out cross-validation the network does not improve on Scorr used directly as a predictor (ρ = +0.627 against +0.744, paired p = 0.14). Time-resolved energy-dispersive X-ray spectroscopy (EDS) provides a partial chemical consistency check, with welding at ρ = 0.82 (raw p = 0.023), but no material survives Holm correction across the six tested. A U-Net segmentation head supervised only by synthesis-derived masks attains Dice = 0.85 in-domain and collapses to a 0.033 output range on real images, 5% of the HSV metric’s range; the photometric-stability advantage previously claimed for it is an artifact of that collapse and is withdrawn. Kinetics-aware MCDM with propagated uncertainty resolves 9 of 15 pairwise orderings, placing stainless steel above welding at 30 chamber days with probability 1.000 and reversing the static ranking. The pipeline, code and fixed data split are fully reproducible (random seed 42). Full article
(This article belongs to the Section AI-Enabled Process Engineering)
Show Figures

Figure 1

23 pages, 2525 KB  
Article
A Deep Mixed-Image Augmentation Strategy for Few-Shot Image Classification
by Rui Wang and Xiaomin Liu
Computers 2026, 15(9), 578; https://doi.org/10.3390/computers15090578 - 3 Sep 2026
Viewed by 116
Abstract
Few-shot image classification suffers from severe data scarcity and unstable generalization. Existing data augmentation strategies still have three major limitations: pixel-level fusion strategies are incompatible with the support–query structure of episodic learning, category selection for cropping-based augmentation is overly simplistic, and most approaches [...] Read more.
Few-shot image classification suffers from severe data scarcity and unstable generalization. Existing data augmentation strategies still have three major limitations: pixel-level fusion strategies are incompatible with the support–query structure of episodic learning, category selection for cropping-based augmentation is overly simplistic, and most approaches rely on a single augmentation method, limiting robustness. To address these issues, this study proposes a deep mixed data augmentation framework that jointly enhances both the support set and the query set. The method first performs global pixel-level fusion to construct fused support and query sets. A Hopfield network then turns fused-support similarities into a pairing matrix H, which assigns a different-class gallery partner for query-side cropping–mixing. Finally, cropping–mixing produces an enhanced query set for model training. The framework is validated using ResNet18+BDC as the backbone. Experimental results on MiniImageNet demonstrate that the proposed method is competitive in few-shot classification, attaining a five-seed test mean of 73.25%/81.88% under 5-way 1-shot and 5-shot. A single complementary run on FC100 attains 66.63%/77.80% and is not a same-backbone ranking against heterogeneous published protocols. Full article
Show Figures

Figure 1

21 pages, 15039 KB  
Article
Class Distribution-Aware Adversarial Training for Semantic Segmentation of Imbalanced Remote Sensing Imagery
by Ying Yu, Chunping Wang, Renke Kou and Qiang Fu
Electronics 2026, 15(17), 3965; https://doi.org/10.3390/electronics15173965 - 2 Sep 2026
Viewed by 126
Abstract
Class imbalance is a fundamental challenge in semantic segmentation of remote sensing imagery, causing deep convolutional networks to systematically neglect minority categories such as vehicles and small water bodies. To address this, we propose a class distribution-aware adversarial training framework that explicitly penalizes [...] Read more.
Class imbalance is a fundamental challenge in semantic segmentation of remote sensing imagery, causing deep convolutional networks to systematically neglect minority categories such as vehicles and small water bodies. To address this, we propose a class distribution-aware adversarial training framework that explicitly penalizes the model’s bias toward dominant classes at the representation level. The framework comprises two key components. First, a Hierarchical Attention U-Net (HAU-Net) integrates a Parallel Hybrid Attention Module (PHAM) at multiple scales, combining spatial and channel attention to amplify feature responses for small and sparsely distributed objects. Second, a distribution discriminator is trained to distinguish the predicted per-class probability distribution from a balanced uniform prior; through a minimax game, the segmentation network is forced to produce class-balanced predictions without manual loss re-weighting. Extensive experiments on the ISPRS Potsdam and Vaihingen datasets demonstrate that our method achieves an overall mIoU of 73.82% and delivers substantial improvements on severely under-represented minority classes, including a +12.11% IoU gain on vehicles and a +19.26% IoU gain on clutter/background, compared to our non-adversarial baseline. Ablation studies confirm that both the hierarchical attention design and the adversarial training contribute independently to these improvements. Our framework offers an effective solution to class imbalance in remote sensing segmentation, with the potential to be integrated into diverse segmentation architectures. Full article
(This article belongs to the Special Issue Applications of Image Processing and Sensor Systems)
Show Figures

Figure 1

15 pages, 20486 KB  
Article
A Multimodal Dual-Stream Framework for Sheep Behavior Recognition Using Skeletal and Local Visual Fusion
by Chuanzhong Xuan, Junze Jia, Suhui Liu and Zhaohui Tang
Animals 2026, 16(17), 2759; https://doi.org/10.3390/ani16172759 - 2 Sep 2026
Viewed by 202
Abstract
Intelligent sheep behavior monitoring is vital for modern husbandry, but faces severe challenges in natural pastures due to high-density flock occlusion. Traditional 2D skeleton-based networks often suffer from depth ambiguity and feature collapse, misclassifying static tremors as dynamic displacement. To overcome this, we [...] Read more.
Intelligent sheep behavior monitoring is vital for modern husbandry, but faces severe challenges in natural pastures due to high-density flock occlusion. Traditional 2D skeleton-based networks often suffer from depth ambiguity and feature collapse, misclassifying static tremors as dynamic displacement. To overcome this, we propose a robust multimodal dual-stream framework using skeletal and local visual fusion. The architecture features an upstream spatial perception stage utilizing YOLOv11m-Pose. To reduce annotation costs and improve robustness, we introduce an Active Hard-Example Mining mechanism, explicitly retaining difficult samples with severe overlapping or edge truncation. For downstream behavioral decisions, a multimodal dual-stream architecture processes the targets. The Spatio–Temporal Kinematic Stream employs a Kinematic Denoising Engine, incorporating a 1D Gaussian filter and displacement dead-zone gate to purify 2D coordinates before feeding them into a BiLSTM network. Concurrently, the Spatial Visual Stream uses a ResNet-50 backbone on cropped RGB patches to capture essential spatial context, addressing the limitations of pure coordinates. Finally, a weighted Softmax layer integrates both streams. Experiments on a complex real-world dataset validate this approach. A baseline kinematic-only model achieved just 69.05% overall accuracy and 68.18% walking precision. In contrast, our dual-stream fusion network achieved 93.26% overall accuracy, elevating walking precision to 97.14% and the eating F1-score to 94.29%. By effectively decoupling similar static and dynamic behaviors, this study demonstrates the indispensability of local visual features, establishing a high-precision baseline for smart livestock monitoring. Full article
(This article belongs to the Section Animal System and Management)
Show Figures

Figure 1

18 pages, 1388 KB  
Article
Enhancing Bone Marrow Lesion Segmentation Through Dual-Channel Deep Neural Networks and Test-Time Augmentation
by Shihua Qin, Hetali Tank, Qiong Wang, Kevin Wang, Jeffery Driban, Timothy McAlindon, Ming Zhang and Juan Shan
Electronics 2026, 15(17), 3950; https://doi.org/10.3390/electronics15173950 - 2 Sep 2026
Viewed by 179
Abstract
Bone marrow lesion (BML) volume is an essential biomarker for understanding knee osteoarthritis (KOA). However, automatic BML segmentation remains challenging due to the irregular shapes and indistinct boundaries of these lesions in knee magnetic resonance images (MRI). To improve BML segmentation, this study [...] Read more.
Bone marrow lesion (BML) volume is an essential biomarker for understanding knee osteoarthritis (KOA). However, automatic BML segmentation remains challenging due to the irregular shapes and indistinct boundaries of these lesions in knee magnetic resonance images (MRI). To improve BML segmentation, this study investigated two established strategies in the specific context of BML segmentation: (1) integrating bone segmentation as an additional output channel in deep neural networks to facilitate BML segmentation, and (2) incorporating test-time augmentation (TTA) to reduce uncertainty during testing. The added bone segmentation channel provides auxiliary anatomical information that may facilitate BML localization. TTA was used to improve boundary alignment and reduce false positives by generating more robust predictions. Multiple State-of-the-Art deep neural networks for segmentation were employed as the baseline models to compare performance before and after implementing the proposed strategies. A 10-fold cross-validation was conducted on a dataset of knee MR scans from 300 participants. Segmentation performance was evaluated using the Dice similarity coefficient (DSC) for overlap accuracy and the 95% Hausdorff Distance (HD95) for boundary alignment. Paired t-tests were used to assess the significance of improvements from the proposed strategies. Both strategies produced improvements in segmentation performance, although the magnitude and statistical significance of the improvements varied across architectures. The DSC improved from 63.1% to 64.8% for Residual U-Net, 64.2% to 65.8% for Swin UNETR, 61.5% to 66.5% for Attention U-Net, and 66.6% to 69.0% for UNet++. These gains were accompanied by improvements in boundary accuracy and reductions in false positives, reflected in lower HD95 values. Comparison with additional medical image segmentation models under the same evaluation framework showed that the dual-channel UNet++ with TTA achieved the highest BML DSC of 69.0%, followed by U-Mamba at 68.6% and nnU-Net at 65.2%, while U-Net + InceptionResNet-v2 achieved 57.7%. These findings support the potential value of dual-channel and TTA strategies for automated BML analysis, while further validation on independent datasets is needed to assess their broader generalizability and clinical utility. Full article
(This article belongs to the Special Issue Image Processing Based on Convolution Neural Network, 3rd Edition)
Show Figures

Figure 1

Back to TopTop