Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (247)

Search Parameters:
Keywords = weak annotations

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
33 pages, 3660 KB  
Article
Flow Matching for Generating Weakly Labeled Bags of Foundation-Model Mammography Representations
by Nikola Jovišić, Milica Škipina, Vanja Švenda, Dubravko Ćulibrk, Boris Antić and Branko Brkljač
AI 2026, 7(9), 375; https://doi.org/10.3390/ai7090375 (registering DOI) - 18 Sep 2026
Abstract
Annotated medical imaging data remain scarce, and labels are often weak and noisy: in mammography, an examination comprises several high-resolution views, yet the diagnostic outcome is recorded only at the breast level. Such problems are naturally cast as Multiple Instance Learning (MIL), where [...] Read more.
Annotated medical imaging data remain scarce, and labels are often weak and noisy: in mammography, an examination comprises several high-resolution views, yet the diagnostic outcome is recorded only at the breast level. Such problems are naturally cast as Multiple Instance Learning (MIL), where the model must infer instance-level structure from bag-level labels alone. Although contemporary foundation encoders supply strong general-purpose embeddings, augmenting MIL data in this representation space remains an open problem as established techniques act on one instance at a time and ignore the statistical dependencies binding a bag together. We address this with SetFlow, a generative model that learns the distribution of complete MIL bags directly in a frozen encoder’s embedding space. SetFlow couples flow-matching training with a Set Transformer-inspired backbone, making it invariant to instance ordering while modeling intra-bag relationships. Generation is conditioned jointly on class label and per-instance scale, yielding coherent, semantically faithful bags rather than isolated vectors. Evaluating on two large public mammography datasets and two encoders, we assess distributional fidelity, nearest-neighbor behavior, and downstream augmentation utility. We show that generated bags reproduce real-data statistics and improve classification in certain configuration, with performance gains varying on the amount of synthetic data. An architecture ablation confirms each design choice contributes to performance. Full article
Show Figures

Figure 1

23 pages, 1425 KB  
Article
MSRA-CCL: Multi-Source Reliability-Aware Aggregation and Cross-Domain Consistency Learning for Weakly Supervised Sentiment Classification
by Jiaxu Wang, Bowen Rong, Chengcheng Li, Huiying Xu and Xinzhong Zhu
Electronics 2026, 15(18), 4230; https://doi.org/10.3390/electronics15184230 - 17 Sep 2026
Viewed by 73
Abstract
Sentiment classification often relies on costly in-domain annotations, while weak-supervision sources provide low-cost but inconsistent labels. We propose Multi-Source Reliability-Aware Aggregation and Cross-domain Consistency Learning (MSRA-CCL) for weakly supervised sentiment classification. MSRA-CCL integrates four heterogeneous sources: VADER, TextBlob, a cross-domain TF–IDF classifier, and [...] Read more.
Sentiment classification often relies on costly in-domain annotations, while weak-supervision sources provide low-cost but inconsistent labels. We propose Multi-Source Reliability-Aware Aggregation and Cross-domain Consistency Learning (MSRA-CCL) for weakly supervised sentiment classification. MSRA-CCL integrates four heterogeneous sources: VADER, TextBlob, a cross-domain TF–IDF classifier, and a cross-domain RoBERTa model. Their predictions, confidence, entropy, probability margins, and agreement are encoded into a 27-dimensional evidence representation. A reliability-aware gate learns instance-dependent soft targets from source-domain validation data without using target-domain training labels; however, labeled target-domain validation data are used only for student-checkpoint selection. The target-domain RoBERTa student is trained with confidence filtering, class balancing, soft-label supervision, and R-Drop consistency regularization. On TweetEval and DynaSent, MSRA-CCL achieves Macro-F1 scores of 0.6909 and 0.6893. The results demonstrate improved use of heterogeneous weak supervision. Full article
Show Figures

Figure 1

24 pages, 8042 KB  
Article
SDE-Net: A Strip-Directional Dynamic-Scale and Edge-Aware Network for SAR Oil Spill Segmentation
by Yifei Shen, Yijing Liu, Guoru Li, Shuxi Chen, Yuanzhi Zhang and Shentao Wang
J. Mar. Sci. Eng. 2026, 14(18), 1715; https://doi.org/10.3390/jmse14181715 - 15 Sep 2026
Viewed by 111
Abstract
Marine oil spill monitoring plays a critical role in environmental protection and emergency response. Synthetic aperture radar (SAR) provides all-day, all-weather, and wide-area imaging capabilities and has become a primary sensor for operational oil spill surveillance. Nevertheless, SAR oil spill segmentation remains challenging. [...] Read more.
Marine oil spill monitoring plays a critical role in environmental protection and emergency response. Synthetic aperture radar (SAR) provides all-day, all-weather, and wide-area imaging capabilities and has become a primary sensor for operational oil spill surveillance. Nevertheless, SAR oil spill segmentation remains challenging. Real slicks are easily confused with look-alike dark formations; oil film boundaries are weak and fragmented under speckle noise; and small and sparse targets, such as ships, are easily overlooked amid the dominant sea surface background and complex coastal structures. To address these challenges, this paper proposes SDE-Net, an encoder–decoder network built on a ConvNeXtV2-Tiny backbone and a UPerNet-style decoder with three task-oriented modules. The Strip-Directional Local Enhancement (SDLE) module refines elongated low-contrast cues in the shallow lateral features of C2 and C3 while limiting the indiscriminate enhancement of speckle-contaminated responses. The Dynamic Scale Pyramid Context (DSPC) module redesigns pyramid-based context aggregation at the deepest stage by introducing an input-conditioned softmax scale gate that adaptively reweights the contributions of predefined contextual branches. The Edge-Aware Dynamic Gated FPN (EDG-FPN) decoder combines DySample-based feature alignment and dynamic gated fusion with a boundary-aware gate driven by an edge prediction branch, modulating high-level semantic propagation in boundary-sensitive regions without requiring additional edge annotations. On the Oil Spill Detection Dataset, SDE-Net achieves a 72.94% mIoU and 82.88% mDice, outperforming the UPerNet baseline equipped with the same ConvNeXtV2-Tiny backbone by 4.25 and 3.84 percentage points, respectively. The results demonstrate improved oil spill delineation and the preservation of sparse small-class structures under complex SAR sea surface conditions. Full article
(This article belongs to the Special Issue Oil Spills in the Marine Environment)
Show Figures

Figure 1

18 pages, 3432 KB  
Article
Metabolomics-Based Identification of α-Glucosidase Inhibitors from Pometia pinnata Stem Bark Using LC-HRMS and Molecular Docking
by Husniati Husniati, Berna Elya, Muhammad Hanafi, Puspa Dewi Narrij Lotulung, Faris Hermawan, Rifaldi Rifaldi, Dela Rosa and Alfi Khatib
Molecules 2026, 31(18), 3233; https://doi.org/10.3390/molecules31183233 - 13 Sep 2026
Viewed by 156
Abstract
Pometia pinnata J.R. Forst. & G. Forst. is traditionally used throughout tropical Asia and the Pacific to manage diabetes-associated hyperglycemia. However, the metabolites responsible for its α-glucosidase inhibitory activity (AGI) remain poorly characterized. This study aimed to identify putative AGI-associated metabolites from P. [...] Read more.
Pometia pinnata J.R. Forst. & G. Forst. is traditionally used throughout tropical Asia and the Pacific to manage diabetes-associated hyperglycemia. However, the metabolites responsible for its α-glucosidase inhibitory activity (AGI) remain poorly characterized. This study aimed to identify putative AGI-associated metabolites from P. pinnata stem bark through metabolomics-based prioritization and tentative annotation using untargeted LC–HRMS, followed by molecular docking to assess their interactions with α-glucosidase. Thirty ethyl acetate–methanol gradient fractions were analyzed by orthogonal partial least squares (OPLS) to prioritize LC–HRMS features associated with AGI activity, followed by molecular docking of the tentatively annotated metabolites against Saccharomyces cerevisiae α-glucosidase (3A4A) and human maltase-glucoamylase (3TOP). The 75% ethyl acetate in methanol fraction showed the strongest AGI, with an IC50 of 5.53 μg/mL. Five AGI-associated metabolites, namely scopoletin, 3,4-dihydroxybenzaldehyde, fisetin, 4-methoxycinnamic acid, and lindetannin, were tentatively annotated in P. pinnata stem bark based on LC-HRMS/MS data. To the best of our knowledge, these annotations have not previously been reported in P. pinnata stem bark. Among these candidates, fisetin showed the most favorable binding interactions with both target enzymes. Metabolomics-guided isolation, followed by NMR analysis, confirmed the structure of scopoletin, although the isolated compound showed weak AGI activity (IC50 > 200 μg/mL). The marked difference between the parent fraction and isolated scopoletin indicates that scopoletin alone is unlikely to account for the observed activity and that other constituents may contribute. Nevertheless, this study provides a promising metabolomics-guided framework for prioritizing and tentatively annotating candidate AGI-associated metabolites in the stem bark of P. pinnata. Full article
(This article belongs to the Section Natural Products Chemistry)
Show Figures

Figure 1

33 pages, 18538 KB  
Article
Boosting Multi-Class SAR Oriented Object Detection via Geo-Topology-Guided Diffusion Synthesis
by Zhen Wang, Gang Wan, Cheng Wang, Qinlong Lan and Yufei Guo
Remote Sens. 2026, 18(18), 3128; https://doi.org/10.3390/rs18183128 - 11 Sep 2026
Viewed by 170
Abstract
Oriented object detection in Synthetic Aperture Radar (SAR) imagery plays an important role in remote sensing, but its performance is usually limited by the shortage of high-quality annotated samples. This problem is particularly prominent in multi-class scenarios, where different targets exhibit significantly different [...] Read more.
Oriented object detection in Synthetic Aperture Radar (SAR) imagery plays an important role in remote sensing, but its performance is usually limited by the shortage of high-quality annotated samples. This problem is particularly prominent in multi-class scenarios, where different targets exhibit significantly different scattering characteristics, scale distributions, orientation variations, and background dependencies. Existing SAR sample synthesis methods are mostly designed for single-category targets or horizontal bounding box constraints, and suffer from insufficient category diversity, weak orientation controllability, and inadequate modeling of geo-topological relationships. To address these problems, this paper proposes the first diffusion-based sample generation method for multi-class SAR oriented object detection. A large-scale SAR image-geo-topological semantic text paired dataset is constructed, and a SAR text-to-image foundation model is pretrained based on Stable Diffusion, enabling the model to learn target categories, quantities, spatial distributions, and geo-topological relationships, thereby improving the geographic plausibility and scene consistency of generated results. Furthermore, a Direction Phase Shifting encoding strategy is proposed to alleviate the boundary discontinuity problem in rotation-angle representation and to achieve precise control of target location, scale, and orientation based on oriented bounding boxes. Meanwhile, an automated sample generation and label refinement pipeline is designed. More accurate oriented bounding boxes that better fit target contours are obtained through wavelet denoising and progressive SAM-2 segmentation, improving the annotation accuracy of synthetic samples. Experimental results show that our method can produce SAR images with diverse scattering characteristics, realistic background variations, and reasonable geo-topological relationships. The generated samples consistently improve detection performance on five baseline oriented object detectors, providing an effective data augmentation strategy for enhancing the robustness and generalization capability of SAR oriented object detection models. Full article
(This article belongs to the Special Issue Deep Learning for Target Detection in Radar Remote Sensing)
Show Figures

Figure 1

29 pages, 11223 KB  
Article
Confidence-Aware Semi-Supervised Vision–Language Contrastive Learning for Abnormal Behavior Recognition
by Haichuan Liu, Jianxin Sun and Xianmin Zhao
Information 2026, 17(9), 879; https://doi.org/10.3390/info17090879 - 10 Sep 2026
Viewed by 159
Abstract
Reliable abnormal behavior recognition from surveillance videos is hindered by the high cost of clip-level annotation, the scarcity of abnormal samples, and the context-dependent nature of behavioral semantics. Although vision–language models offer strong semantic transferability, their application under limited supervision remains susceptible to [...] Read more.
Reliable abnormal behavior recognition from surveillance videos is hindered by the high cost of clip-level annotation, the scarcity of abnormal samples, and the context-dependent nature of behavioral semantics. Although vision–language models offer strong semantic transferability, their application under limited supervision remains susceptible to noisy pseudo-labels and confirmation bias. We propose confidence-aware semi-supervised vision–language contrastive learning (CA-VLC), which jointly exploits limited labeled videos and abundant unlabeled videos. Building on an existing CLIP-initialized temporal backbone, CA-VLC combines behavior-only and context-enriched text prototypes through confidence- and agreement-guided semantic fusion. For unlabeled videos, the model generates predictions from weakly augmented views and selects reliable pseudo-labels using entropy-based confidence estimation and class-adaptive thresholds. Detached weak-view targets then supervise strongly augmented views through confidence-weighted self-training without requiring an additional teacher network. Furthermore, cross-view consistency regularization and confidence-aware contextual alignment suppress unreliable semantic cues and improve robustness to contextual noise. Experiments on CABR50 demonstrate consistent improvements across multiple labeled-data ratios, while evaluations on CABRZ6 and UCF-101 assess prompt-based transfer to predefined target label sets without target-domain fine-tuning. With 10% labeled videos, CA-VLC achieves 84.06% Top-1 accuracy and 83.51% Macro-F1, retaining 95.47% of its fully supervised Top-1 accuracy of 88.05%, thereby demonstrating its effectiveness for label-efficient abnormal behavior recognition. Full article
Show Figures

Graphical abstract

22 pages, 13745 KB  
Article
A Spatial Prior-Guided Feature Enhancement and Multi-Branch Complementary Learning Framework for Small Ship Detection in SAR Images
by Tao Liu, Yuanyuan Zhao, Zhenhua Li, Shuang Liu and Dong Li
Remote Sens. 2026, 18(18), 3106; https://doi.org/10.3390/rs18183106 - 10 Sep 2026
Viewed by 286
Abstract
Ship detection in Synthetic Aperture Radar (SAR) imagery is essential for maritime surveillance and situational awareness. Despite the advances of deep learning for SAR ship detection, small-target detection is still hindered by severe feature degradation from repeated down-sampling and insufficiently discriminative representations under [...] Read more.
Ship detection in Synthetic Aperture Radar (SAR) imagery is essential for maritime surveillance and situational awareness. Despite the advances of deep learning for SAR ship detection, small-target detection is still hindered by severe feature degradation from repeated down-sampling and insufficiently discriminative representations under weak scattering and complex background clutter. To alleviate this dilemma, a Spatial Prior-guided feature enhancement and Multi-branch Complementary learning framework is proposed, termed SPMC, for small ship detection in SAR images. Specifically, a Spatial Prior-Guided Feature Enhancement (SPFE) module is designed to derive spatial attention priors for multi-scale features from ground-truth annotations, thereby emphasizing target-related responses and strengthening small-ship representations. Second, a Multi-branch Complementary Classification (MCC) module is developed, which introduces multiple auxiliary classification heads to learn complementary discriminative information from different classification perspectives. Furthermore, a dual-weighted complementary regularization strategy is proposed to encourages different classifiers to focus on hard samples, thereby improving the discriminative capability for small ships. Extensive experiments on the HRSID and LS-SSDD benchmarks validate the effectiveness of the proposed framework. For extremely small ships with very limited image coverage, SPMC improves the baseline YOLOv11 detector by 3.25%/0.99% in AP50/AP0.5:0.95 on LS-SSDD and by 1.28%/1.21% on HRSID, demonstrating its effectiveness in challenging SAR small ship detection. Full article
Show Figures

Figure 1

27 pages, 8446 KB  
Article
Filtered Distillation from a Large Vision Teacher for Infrared Small Target Detection
by Zhanxu Jiang, Wenbin Chen, Zhi Li, Zhen Yuan, Min Wu, Gong Cheng and Ziwei Wang
Remote Sens. 2026, 18(18), 3079; https://doi.org/10.3390/rs18183079 - 8 Sep 2026
Viewed by 186
Abstract
Infrared target detection supports the continuous observation of traffic participants and low-altitude targets across aerial and fixed-view imaging settings, particularly under weak or changing illumination. However, infrared targets are often small, weakly textured, and easily confused with thermal noise and background clutter. Large [...] Read more.
Infrared target detection supports the continuous observation of traffic participants and low-altitude targets across aerial and fixed-view imaging settings, particularly under weak or changing illumination. However, infrared targets are often small, weakly textured, and easily confused with thermal noise and background clutter. Large pretrained vision models offer strong representation and generalization capabilities, but their parameter counts and computational costs make direct deployment on edge platforms with limited resources impractical. To transfer these capabilities to lightweight models, this paper proposes knowledge distillation at the label level based on filtered teacher detections. A large teacher adapted to the infrared domain first generates candidate boxes. Candidate boxes are selected using confidence thresholds, class reliability, and spatial relationships with ground truth annotations. The retained teacher boxes and the original annotations jointly form the student training targets, while the student retains its standard detection loss and original inference structure. On an independent sequence-level test set, the YOLOv5n baseline obtains an mAP@0.5 of 0.2277 and an mAP@0.5:0.95 of 0.1070, whereas filtered label distillation obtains 0.2547 and 0.1153, respectively, over three matched seeds. The filtering rules and thresholds are fixed before this evaluation, and the fixed Epoch-30 checkpoint is used for every run. Both student architectures are deployed on RK3588. Full article
Show Figures

Figure 1

35 pages, 3044 KB  
Article
Zero-Shot Annotation by Large Language Model with Serial Correction of Mixed Label Corruption for Weakly Supervised Financial News Classification
by Jianxin Sun and Haichuan Liu
Information 2026, 17(9), 864; https://doi.org/10.3390/info17090864 - 7 Sep 2026
Viewed by 161
Abstract
Multi-label classification of financial news is frequently affected by incomplete and noisy annotations, while obtaining expert-curated labels at scale is prohibitively expensive. This study proposes a weakly supervised classification framework that combines large language model (LLM) zero-shot annotation with a serial label-correction strategy. [...] Read more.
Multi-label classification of financial news is frequently affected by incomplete and noisy annotations, while obtaining expert-curated labels at scale is prohibitively expensive. This study proposes a weakly supervised classification framework that combines large language model (LLM) zero-shot annotation with a serial label-correction strategy. The framework first uses an LLM to generate initial weak labels and then refines them through a two-stage Correct→Clean procedure that recovers missing labels via centrality-weighted graph propagation before suppressing label noise. Systematic experiments on a financial subset of Reuters-21578 show that, under an extreme mixed-corruption setting with 80% missing labels and 15% noise labels, Correct→Clean increases the Micro-F1 from 0 to 0.6748. In an end-to-end evaluation, the proposed framework achieves a Micro-F1 of 0.8882 with reduced-dimensional features, recovering 88.69% of the performance gap to fully supervised learning. Additional experiments on the RCV1 Topics and AAPD datasets confirm that the advantage of Correct→Clean is consistently reproduced across domains and dataset sizes. These findings demonstrate that coupling LLM-generated annotations with ordered label correction offers an effective means of addressing the joint effects of missing and noisy labels, providing a promising approach to financial text classification when expert annotations are scarce. Full article
Show Figures

Figure 1

17 pages, 2700 KB  
Article
Leakage-Controlled Classification of Dataset-Derived Ordinal p53-Signature Size Categories in Fallopian-Tube Immunohistochemistry
by Ali Alhazmi
Diagnostics 2026, 16(17), 2869; https://doi.org/10.3390/diagnostics16172869 - 7 Sep 2026
Viewed by 241
Abstract
Background and Objectives: p53 signatures are segments of strongly p53-immunoreactive secretory epithelium in the fallopian tube. We tested whether fixed pathology-pretrained image representations could classify three ordered categories reconstructed from the public dataset’s reported cell-count intervals. Methods: The primary cohort comprised 113 canonical [...] Read more.
Background and Objectives: p53 signatures are segments of strongly p53-immunoreactive secretory epithelium in the fallopian tube. We tested whether fixed pathology-pretrained image representations could classify three ordered categories reconstructed from the public dataset’s reported cell-count intervals. Methods: The primary cohort comprised 113 canonical p53 20× fields from 67 patients: 59 Small (12–20 cells), 24 Medium (21–80), and 30 Large (>80). A 141-field partial-label cohort retained boundary-spanning intervals for sensitivity analysis. Twenty-six eligible feature/head pipelines were evaluated by three-repeat five-fold nested patient-grouped cross-validation. Within every outer fold, the pipeline was selected using inner validation only; primary performance metrics were computed within each repeat and averaged across the three repeats, while the resulting 339 held-out predictions were pooled only for the confusion-matrix visualization. Uncertainty was estimated using 2000 patient-cluster bootstrap replicates conditional on these recorded predictions. Results: The fold-wise nested selector achieved quadratic-weighted kappa (QWK) 0.687 (95% CI 0.547–0.785), macro-F1 0.628 (0.568–0.680), balanced accuracy 0.637 (0.581–0.693), macro-AUROC 0.805 (0.750–0.855), and macro-AUPRC 0.681 (0.633–0.754). Small, Medium, and Large sensitivities were 0.870 (0.801–0.928), 0.264 (0.159–0.386), and 0.778 (0.617–0.906); severe Small–Large errors occurred in 6.5% (2.9–10.6%). Equal patient weighting gave QWK 0.700. In a secondary locked UNI2-h/CORAL analysis, ordinal temperature scaling reduced ECE from 0.195 to 0.157 and negative log-likelihood from 1.291 to 0.778. Conclusions: Fixed representations provided useful internal discrimination of dataset-derived ordinal categories, with substantial fold variability and weak Medium-category sensitivity. The results do not validate lesion measurement, detection, diagnosis, or clinical use; independent pathologist-annotated evaluation is required. Full article
(This article belongs to the Special Issue Artificial Intelligence in Pathological Image Analysis, 3rd Edition)
Show Figures

Figure 1

19 pages, 4600 KB  
Review
Evaluating Sequence Foundation Models for Insect Olfactory Gene Discovery: Failure Modes, Evidence Standards, and a Benchmark Agenda
by Huiqin Li, Dian Zhou, Wei Pu, Guoxing Wu, Chun Xiao, Xi Gao and Junfu Yu
Insects 2026, 17(9), 928; https://doi.org/10.3390/insects17090928 - 4 Sep 2026
Viewed by 313
Abstract
Genome-scale searches for insect olfactory genes are vulnerable to rapid receptor divergence, tandem duplication, fragmented gene models, circular labels, and evolutionary leakage between training and test data. Sequence foundation models trained on DNA or proteins may extend homology-based annotation, but available evidence does [...] Read more.
Genome-scale searches for insect olfactory genes are vulnerable to rapid receptor divergence, tandem duplication, fragmented gene models, circular labels, and evolutionary leakage between training and test data. Sequence foundation models trained on DNA or proteins may extend homology-based annotation, but available evidence does not support their use as autonomous annotators or direct predictors of sensory function. This critical review organizes the problem around annotation failures rather than model catalogues. We distinguish six inference levels, from candidate locus detection to organism-level function, and define task-specific high-confidence reference criteria. We also propose a benchmark using biological hard negatives, sequence-cluster and taxonomic holdouts, conventional baselines, calibration, abstention, and workload-aware retrieval metrics. Evidence integration is defined operationally as a pipeline that retains model scores alongside similarity, domains, topology, phylogeny, expression, and assay evidence, without promoting a weak signal to a stronger biological claim. A published ESM-2-supported analysis of remote insect chemoreceptor homology illustrates both the potential and the current evidence boundary. A convincing benefit will require an incremental gain over transparent baselines under frozen biologically realistic splits, with reliable uncertainty and a clear route from ambiguous predictions to review or experiment. Full article
Show Figures

Figure 1

24 pages, 2624 KB  
Article
Detection-Guided ROI-Constrained Diffusion for Weakly Supervised White Blood Cell Segmentation: A Retrospective Internal and External Dataset Evaluation
by Julius Bamwenda, Mehmet Siraç Özerdem, Orhan Ayyildiz, Veysi Akpolat and İrem Akpolat
J. Clin. Med. 2026, 15(17), 6846; https://doi.org/10.3390/jcm15176846 - 3 Sep 2026
Viewed by 302
Abstract
Background: Accurate white blood cell (WBC) segmentation is important for quantitative microscopic image analysis, but conventional supervised approaches depend on labor-intensive pixel-level annotations and may exhibit reduced robustness across datasets. This study proposes a detector-guided, region of interest (ROI)-constrained framework that combines [...] Read more.
Background: Accurate white blood cell (WBC) segmentation is important for quantitative microscopic image analysis, but conventional supervised approaches depend on labor-intensive pixel-level annotations and may exhibit reduced robustness across datasets. This study proposes a detector-guided, region of interest (ROI)-constrained framework that combines automatic localization, pseudo-mask-based weak supervision, and diffusion-assisted segmentation. Methods: You Only Look Once version 13 Nano (YOLOv13-N) was used to localize WBCs and define ROIs, within which segmentation was performed using the proposed diffusion-assisted model. Automatically generated pseudo-masks served as segmentation-training targets, while expert masks were retained for reference evaluation. Experiments used a verified Dicle cohort of 14,721 records, partitioned into 10,305 training, 2208 validation, and 2208 held-out test records. Performance was evaluated separately at ROI-conditional and end-to-end levels. Generalization was assessed on 11,200 Raabin-WBC records using the frozen framework without external tuning. Results: The automatically generated pseudo-masks achieved a mean Dice score of approximately 0.702, demonstrating usable but imperfect weak supervision. On the held-out Dicle test set, the proposed framework achieved a Dice score of 0.7188, intersection over union (IoU) of 0.5796, precision of 0.8966, and recall of 0.6422 under ROI-conditional evaluation. End-to-end performance was 0.6705 Dice, 0.5408 IoU, 0.8378 precision, and 0.5986 recall, demonstrating the influence of localization on overall performance. Comparison with reference segmentation architectures showed that the proposed framework did not maximize internal Dice, while achieving the highest reported ROI-conditional precision. Frozen external evaluation on Raabin-WBC revealed further degradation under dataset shift, with failure analysis identifying detector/ROI transfer as an important end-to-end bottleneck. Conclusions: The proposed framework demonstrates the feasibility of WBC segmentation using detector-guided ROI processing and pseudo-mask-based weak supervision, reducing reliance on expert pixel-level segmentation targets. The findings further show that robust cross-dataset localization is critical to end-to-end generalization and provide a clear direction for improving weakly supervised WBC segmentation across heterogeneous microscopy datasets. Full article
Show Figures

Figure 1

25 pages, 10320 KB  
Article
UAV-Based Transmission Tower Inspection Using Hierarchical Multitask Learning Under Heterogeneous Supervision
by Hongzhu Song, Kaiyue Liu, Keqin Jia, Liyan Liu, Xiaomeng Wu, Dezhi Meng and Ruisheng Ma
Appl. Sci. 2026, 16(17), 8772; https://doi.org/10.3390/app16178772 - 3 Sep 2026
Viewed by 198
Abstract
Unmanned aerial vehicle (UAV) inspection of transmission towers requires joint analysis of corridor geometry, tower components, and localized defects, whereas available datasets provide incompatible annotations at different scales. Tower-HMT is a hierarchical multitask network that shares a ConvNeXt-Tiny encoder and feature pyramid across [...] Read more.
Unmanned aerial vehicle (UAV) inspection of transmission towers requires joint analysis of corridor geometry, tower components, and localized defects, whereas available datasets provide incompatible annotations at different scales. Tower-HMT is a hierarchical multitask network that shares a ConvNeXt-Tiny encoder and feature pyramid across corridor parsing, component parsing, missing-bolt localization, and component-condition recognition. Four public real-image datasets are organized by native supervision rather than merged into a flat label space. Task-conditioned feature modulation separates dataset statistics; topology and boundary losses preserve thin conductors and lattice edges; and a detached tower-probability gate supplies structural context to a high-resolution bolt head. Source-image groups define non-overlapping training, validation, and test partitions. On held-out data, the network achieved foreground mIoU values of 0.597 for tower-conductor parsing and 0.561 for box-conditioned component parsing, an AP50 of 0.123 for missing-bolt localization, and a macro-F1 of 0.925 for component-condition recognition. Boundary, class-wise, robustness, threshold-sensitivity, and latency results quantify the effects and limitations of hierarchical learning. The framework provides a reproducible real-image baseline for review-oriented inspection, while low missing-bolt localization accuracy and weak component masks remain the principal constraints. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

19 pages, 20007 KB  
Article
Lightweight Underwater Marine-Debris Detection for Sustainable Ocean Monitoring Using Receptive-Field Aggregation and Residual Channel-Spatial Recalibration
by Yuhua He, Zhiqiang Huang and Yun Guo
Sustainability 2026, 18(17), 8986; https://doi.org/10.3390/su18178986 - 2 Sep 2026
Viewed by 229
Abstract
Marine debris threatens aquatic habitats and complicates inspections in ports, seabed environments, and offshore infrastructure. Previous lightweight detectors remain vulnerable to weak texture, blurred boundaries, and cluttered multi-scale features, while direct network expansion conflicts with restricted onboard resources. This study adapts YOLO11n by [...] Read more.
Marine debris threatens aquatic habitats and complicates inspections in ports, seabed environments, and offshore infrastructure. Previous lightweight detectors remain vulnerable to weak texture, blurred boundaries, and cluttered multi-scale features, while direct network expansion conflicts with restricted onboard resources. This study adapts YOLO11n by integrating Receptive-Field Aggregation (RFA) with a Residual Channel-Spatial Recalibration (RCSA) implementation based on dynamic residual groups. Experiments used a public 15-class dataset with 10,884 training images and 1001 model-selection validation images containing 1892 annotated objects. All principal checkpoints were trained for 100 epochs. Across seeds 42, 2026, and 3407, RFA + RCSA achieved validation precision 0.861 ± 0.016, recall 0.801 ± 0.012, mean average precision at IoU 0.5 (mAP@0.5) 0.848 ± 0.001, and mAP@0.5:0.95 0.511 ± 0.002. On an audited group-disjoint holdout (498 images; 963 instances), the corresponding means were 0.820 ± 0.033, 0.748 ± 0.014, 0.779 ± 0.018, and 0.467 ± 0.007. The detector contains 4.19 M parameters and requires 8.91 giga floating-point operations (GFLOPs). These results position it as a lightweight candidate for resource-constrained remotely operated vehicle (ROV) perception; they do not establish real-time embedded deployment. Full article
(This article belongs to the Section Sustainable Oceans)
Show Figures

Figure 1

49 pages, 799 KB  
Article
HIEF: An Interpretable Evidence-Fusion Framework for Phishing Email Detection with Decomposable Decision Uncertainty and a Preliminary English–Spanish Evaluation
by Carolina Del-Valle-Soto, Carlos-Santiago Cruz-Diaz, Manuel Cardona, Hiram Ponce, Leonardo J. Valdivia and Paolo Visconti
Algorithms 2026, 19(9), 741; https://doi.org/10.3390/a19090741 - 1 Sep 2026
Viewed by 270
Abstract
(1) Background: Phishing remains a pervasive and economically damaging cyberthreat. The dominant detection paradigm has moved toward deep neural and transformer-based classifiers, a literature that reports high accuracy and that does not, in general, expose a per-decision justification, whereas interpretability and auditability are [...] Read more.
(1) Background: Phishing remains a pervasive and economically damaging cyberthreat. The dominant detection paradigm has moved toward deep neural and transformer-based classifiers, a literature that reports high accuracy and that does not, in general, expose a per-decision justification, whereas interpretability and auditability are increasingly required in regulated environments; no comparison against transformer-scale detectors is made in this paper. This work asks how far a fully interpretable detector can close the accuracy gap to an opaque text classifier while preserving per-decision explanations, and what such a detector returns that accuracy alone does not measure. (2) Methods: HIEF, an interpretable evidence-fusion framework, is presented. Each email is represented by eighteen human-readable signals: fourteen structural and linguistic cues and four lexical aggregates derived from a published sparse log-odds lexicon. The signals are fused by three transparent layers, namely an L1-regularized logistic model, a shallow interaction-rule tree, and a calibrated Dempster–Shafer stage that reports belief, disbelief and ignorance masses together with an order-invariant global conflict coefficient derived in closed form. A logistic meta-learner fitted on out-of-fold component scores integrates the three layers. The evidential layer uses a type-aware calibration in which discrete signals are calibrated on their attainable values and continuous signals by isotonic regression. Evaluation uses 38,908 public emails, 38,512 of them after exact-duplicate removal, with near-duplicate control, group-aware partitioning, ten repeated splits, a source-held-out protocol, a two-class cross-source test set, a component ablation and a human audit of 100 messages annotated independently by two evaluators. (3) Results: Under group-aware partitioning, HIEF attains an F1 of 0.855 and the strongest term frequency–inverse document frequency (TF–IDF) baseline 0.954; a compact character n-gram neural reference model, evaluated over the same ten partitions, attains 0.973. The linear layer alone attains 0.872, so the two fusion layers do not improve accuracy over it, and the paired difference of 0.017 excludes zero. Type-aware calibration raises the evidential layer from 0.771 to 0.780 and more than halves its partition-to-partition standard deviation, but does not make it competitive; the weakness, therefore, lies in the fusion formulation rather than in the binning. What the evidential layer does supply is a decomposable account of decision uncertainty: the ignorance mass separates errors from correct decisions, 0.265 against 0.175. The human audit reaches an inter-annotator Cohen’s kappa of 0.950 over the five categories before adjudication, and shows that the permissive corpus label agrees with human phishing judgment at a Cohen’s kappa between 0.18 and 0.21, against 0.70 to 0.77 for the automatic strict rule; the audited block is annotated by two of the authors and its human positives are confined to the advance-fee family, so the audit is a bounded comparison of label assignments and not an independent annotation study. (4) Conclusions: HIEF is positioned as an uncertainty and explanation framework rather than as an accuracy-improving fusion method, since the measured accuracy cost of the fusion layers is not compensated by an accuracy gain. Quantifying how much of the performance reported on these widely used corpora is attributable to template leakage and to label permissiveness is a contribution independent of the detector itself. Cross-source operation has not been demonstrated: specificity falls to 0.041 on an unseen collection, so all evaluation reported here is proof-of-concept and no operational deployment claim is made. The Spanish-language evaluation rests on a small and entirely positive subset and is reported as preliminary. Full article
Show Figures

Figure 1

Back to TopTop