Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (187)

Search Parameters:
Keywords = semantic distortion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
29 pages, 6482 KB  
Article
A Synergistic Knowledge Graph and LLM-Driven Framework for Intelligent Process Decision-Making Systems
by Deguo Yao, Zhaoze Sun, Jie Gao, Haoyu Cao and Xiaoyue Li
Appl. Syst. Innov. 2026, 9(8), 171; https://doi.org/10.3390/asi9080171 - 13 Aug 2026
Viewed by 89
Abstract
To address the problems of complex process knowledge sources, heterogeneous representations, dispersed semantic associations, and limited reusability in the domain of machining distortion of thin-walled parts, this study proposes a knowledge graph construction method for the workpiece machining distortion domain, together with an [...] Read more.
To address the problems of complex process knowledge sources, heterogeneous representations, dispersed semantic associations, and limited reusability in the domain of machining distortion of thin-walled parts, this study proposes a knowledge graph construction method for the workpiece machining distortion domain, together with an intelligent decision-making framework driven by the collaboration of knowledge graphs and large language models. First, a domain ontology model is established around core concepts, including workpiece objects, deformation-driving factors, analytical resources, analytical methods, and optimization knowledge, thereby providing a unified semantic foundation for domain knowledge organization. Second, considering the characteristics of domain texts, such as dense technical terminology, ambiguous entity boundaries, and complex relation expressions, a dual-channel knowledge extraction method integrating BERT-BiLSTM-CRF and Universal Information Extraction (UIE) is developed to achieve high-precision extraction of entities and relations from unstructured texts. Knowledge fusion is further carried out through cross-validation, entity disambiguation, coreference resolution, and semantic alignment, and the extracted knowledge is ultimately stored and organized in Neo4j. Furthermore, an intelligent decision-making framework based on the collaboration of knowledge graphs and large language models is constructed. In this framework, a LoRA-tuned Qwen model is employed for user intent recognition and key information extraction, RapidFuzz WRatio is adopted for similar-node retrieval, and local subgraph construction, Label Propagation-based community detection, Betweenness Centrality-based key-node analysis, and evidence fusion are integrated to support process recommendation and intelligent question answering. Based on the proposed framework, an intelligent decision-making system is further developed for process recommendation and intelligent question answering in machining distortion scenarios. Experimental results show that the proposed dual-channel knowledge extraction model achieves an F1-score of 0.88, demonstrating its effectiveness in knowledge acquisition for the machining distortion domain. The constructed knowledge graph contains 4639 entities and 5822 relations, enabling a systematic representation of machining distortion knowledge. Case studies further demonstrate that the proposed method can generate interpretable recommendation results under complex process constraints in real industrial query scenarios. Overall, the proposed approach provides a feasible pathway for the structured organization, intelligent retrieval, and decision support of workpiece machining distortion knowledge. Full article
Show Figures

Figure 1

22 pages, 3131 KB  
Article
Reliability-Aware Gaussian Residual Counterpart Generation for Robust Multi-View Clustering with Noisy Correspondence
by Xin Liu, Lican Dai and Boyuan Zheng
Computation 2026, 14(8), 185; https://doi.org/10.3390/computation14080185 - 12 Aug 2026
Viewed by 129
Abstract
Multi-view clustering (MvC) aims to discover cluster structures by exploiting complementary information across views. Most existing MvC methods assume that same-index observations across views describe the same semantic instance. In practice, however, index-aligned observations can be semantically unrelated. This inconsistency between observed index-level [...] Read more.
Multi-view clustering (MvC) aims to discover cluster structures by exploiting complementary information across views. Most existing MvC methods assume that same-index observations across views describe the same semantic instance. In practice, however, index-aligned observations can be semantically unrelated. This inconsistency between observed index-level correspondence and underlying semantic correspondence is known as noisy correspondence (NC). Learning from such mismatched pairs imposes erroneous cross-view constraints and distorts clustering. Many existing methods only suppress unreliable pairs. This discriminative strategy avoids incorrect alignment but also excludes suspicious pairs from cross-view learning. To reuse these pairs without enforcing incorrect correspondence, we propose Reliability-Aware Gaussian Residual Counterpart Generation. Using reliability estimates derived from cross-view losses, the framework retains observed counterparts for reliable pairs and routes unreliable pairs to counterpart generation. For each unreliable pair, prototype-level semantic transport locates a matched target-view prototype. A Gaussian residual model estimated from reliable target-view samples captures variations around this prototype. The framework samples a residual from this model and adds it to the prototype center, yielding a semantically matched yet diverse counterpart. Random walk-based intra-view contrastive learning further preserves neighborhood structures. Experiments on Scene15, LandUse21, Reuters, and CCV20 achieve the best average ACC, NMI, and ARI across the evaluated NC ratios. Ablation and transfer studies further support the effectiveness of the proposed design. Full article
(This article belongs to the Special Issue Computational Methods for Multi-View Representation Learning)
Show Figures

Figure 1

15 pages, 421 KB  
Article
Representation and Geometric Collapse in Spatiotemporal EEG Classifiers: A Mathematical Diagnostic Framework
by Ahmed El Badaoui, Hicham Ben Alla, Manal Hilali, Said Ben Alla and Abdellah Ezzati
Signals 2026, 7(4), 76; https://doi.org/10.3390/signals7040076 - 4 Aug 2026
Viewed by 196
Abstract
Spatiotemporal deep learning models, like Graph Neural Networks (GNNs), Transformers, and selective State Space Models (Mamba), have achieved impressive performance in electroencephalogram (EEG) decoding and affective computing. However, their generalization performance often degrades severely under subject-independent Leave-One-Subject-Out (LOSO) cross-validation protocols. This generalization drop [...] Read more.
Spatiotemporal deep learning models, like Graph Neural Networks (GNNs), Transformers, and selective State Space Models (Mamba), have achieved impressive performance in electroencephalogram (EEG) decoding and affective computing. However, their generalization performance often degrades severely under subject-independent Leave-One-Subject-Out (LOSO) cross-validation protocols. This generalization drop is often attributed to generic domain shifts in standard Brain–Computer Interface (BCI) literature, and is tackled by parameter-heavy adaptations. In contrast, this paper proposes a unified mathematical diagnostic framework to audit and measure the underlying representation and geometric collapse in spatiotemporal brain–computer interfaces. More concretely, we formalize: (1) Topological over-smoothing under volume conduction through Graph Dirichlet Energy bounds indicating GCNs as low-pass filters that smooth localized electrode variations; (2) representation collapse through the Normalized Rank Uniformity Index (NRUI) based on the Shannon Entropy of latent covariance eigenvalues, that distinguishes between dimensional and semantic collapse; and (3) geometric manifold distortions under subject domain shifts on the Symmetric Positive Definite (SPD) Riemannian manifold under the Affine-Invariant Riemannian Metric (AIRM) projection. Auditing these diagnostic metrics on canonical models across DEAP, DREAMER, and SEED, we demonstrate why standard spatiotemporal architectures suffer from performance collapse in cross-subject configurations. We provide BCI engineers with a tangible mathematical blueprint to design robust, collapse-resistant decoders. Full article
Show Figures

Figure 1

23 pages, 10620 KB  
Article
A Novel Coverless Image Steganography Scheme Based on LLM-Guided Image Generation
by Yung-Chen Chou, Jun-Yi Liu, Yuan-Yu Tsai and Chun-Hsiu Yeh
Electronics 2026, 15(15), 3415; https://doi.org/10.3390/electronics15153415 - 2 Aug 2026
Viewed by 206
Abstract
With the advancement of deep learning-based steganalysis and prevalence of lossy compression mechanisms in social network transmission, traditional steganography based on cover modification (such as LSB substitution) faces dual challenges of security and robustness. This study proposes a novel steganographic framework based on [...] Read more.
With the advancement of deep learning-based steganalysis and prevalence of lossy compression mechanisms in social network transmission, traditional steganography based on cover modification (such as LSB substitution) faces dual challenges of security and robustness. This study proposes a novel steganographic framework based on generative artificial intelligence and predefined semantic mapping. Unlike embedding ciphertext in pixel noise, this method utilizes a shared mapping protocol (Codebook) to transform abstract information into concrete visual elements (such as characters, actions, scenes, and styles), and constructs stego-images through generative models. Experimental results show that when both parties share the same key, the system achieves full semantic recovery. Using an explicitly specified pipeline (Gemini 2.5 Flash Image for synthesis and Gemini 2.5 Flash for parsing), we evaluate the scheme on an enlarged, randomly sampled scenario set spanning three to six active semantic dimensions. Across these scenarios, we report per-dimension accuracy, the end-to-end full-recovery rate with 95% confidence intervals, and the partial-recovery rate under controlled JPEG compression, re-scaling, Gaussian noise, and cropping, rather than a single aggregate figure. The results indicate that carrying information at the semantic level yields graceful degradation under common channel distortions together with high visual camouflage, while also revealing that recovery reliability decreases as more semantic dimensions are activated simultaneously. We therefore present these findings as a proof of concept and explicitly separate demonstrated results from hypotheses left to future work. We therefore present these findings as a proof of concept and explicitly separate demonstrated results from hypotheses left to future work. Full article
(This article belongs to the Special Issue New Trends in Cybersecurity and Privacy Protection)
Show Figures

Figure 1

32 pages, 11357 KB  
Article
Transforming Misleading Multimodal Reviews into Decision-Ready Evidence for E-Commerce: A Sarcasm Detection, LLM-Driven Semantic Correction, and Non-Compensatory Decision-Support Framework
by Hongze Hao and Chien-Chih Wang
J. Theor. Appl. Electron. Commer. Res. 2026, 21(8), 245; https://doi.org/10.3390/jtaer21080245 - 1 Aug 2026
Viewed by 319
Abstract
Multimodal online reviews, which integrate textual content with user-uploaded images, have emerged as a significant information source for consumer decision-making and enterprise service enhancement in electronic commerce. Sarcastic reviews, characterized by positive textual statements that are contradicted by negative visual evidence, systematically distort [...] Read more.
Multimodal online reviews, which integrate textual content with user-uploaded images, have emerged as a significant information source for consumer decision-making and enterprise service enhancement in electronic commerce. Sarcastic reviews, characterized by positive textual statements that are contradicted by negative visual evidence, systematically distort sentiment analysis and subsequent evaluations. Most existing research on multimodal sarcasm detection is limited to binary classification. In contrast, many conventional review-based decision models permit cross-criterion trade-offs, allowing favorable performance on secondary attributes to obscure deficiencies in tolerance-sensitive attributes. To bridge the gap between misleading review content and reliable decision support, this study introduces an integrated three-module framework, the RDSF-MSD-SC. The first module, the Deep Feature Fusion Sarcasm Detection Model (DFF-SDM), combines RoBERTa-based contextual text analysis with CLIP-based cross-modal alignment to detect reviews in which linguistic and visual cues diverge. The second module, the Large Language Model-Driven Structured Semantic Correction Mechanism (LLM-SSCM), reinterprets identified sarcastic reviews and translates their implicit negative sentiment into structured evaluation units (problem dimension, problem severity, and correction confidence) that can be directly utilized by decision models. The third module, Non-Compensatory Decision-Making based on Contextual Tolerance and Mismatch Rate (NCDM-CTMR), aggregates the resulting severity matrix using context-specific tolerance vectors, penalizing breaches of tolerance rather than averaging them out. An empirical analysis of 8623 multimodal reviews from eight major U.S. passenger airlines demonstrates that DFF-SDM achieves an F1-score of 0.86, LLM-SSCM attains an expert-assessed correction quality score of 0.92, and NCDM-CTMR identifies Delta as the most context-appropriate airline with zero mismatch in a time-sensitive business travel scenario. Delta maintained the top ranking in 99.82% of the 20,000 Monte Carlo simulations conducted under a simplified mismatch formulation. Ablation analysis showed that the full pipeline recovered negative evidence from approximately 87% of the 2847 expert-labeled sarcastic reviews, a rate bounded by DFF-SDM’s recall of 0.87; the no-correction and no-detection variants recovered approximately 0%. Compared to the WSM, TOPSIS, and VIKOR baselines, the proposed framework yields a more threshold-sensitive differentiation among mid-tier alternatives. Full article
(This article belongs to the Special Issue Human–Technology Synergies in AI-Driven E-Commerce Environments)
Show Figures

Figure 1

18 pages, 63175 KB  
Article
OrbitGS: High-Fidelity 3D Reconstruction of On-Orbit Non-Cooperative Targets via Physically Decoupled Gaussian Splatting
by Ligang Li, Ziyan Qin, Fan Zhang, Wenbo Zhou, Yi Li and Yang Li
Remote Sens. 2026, 18(15), 2489; https://doi.org/10.3390/rs18152489 - 31 Jul 2026
Viewed by 320
Abstract
High-precision 3D reconstruction of on-orbit non-cooperative targets is essential for space situational awareness. However, extreme space environments induce severe imaging degradations, including high-dynamic-range (HDR) illumination, rapid motion blur, and platform jitter. Traditional 3D Gaussian Splatting (3DGS) conflates these optical distortions with geometric optimization, [...] Read more.
High-precision 3D reconstruction of on-orbit non-cooperative targets is essential for space situational awareness. However, extreme space environments induce severe imaging degradations, including high-dynamic-range (HDR) illumination, rapid motion blur, and platform jitter. Traditional 3D Gaussian Splatting (3DGS) conflates these optical distortions with geometric optimization, leading to pathological structural inflation and the loss of thin appendages like solar panels. To overcome this, we propose OrbitGS, a physically decoupled 3DGS framework. OrbitGS integrates physical imaging priors via a Kinematics-Driven Degradation Synthesizer (KDDS) to deterministically extract view-specific degradation kernels. Furthermore, a blur-decoupled rendering strategy with intensity-aware weighting mitigates HDR variations, while a semantic-aware densification scheme mathematically penalizes abnormal primitive expansion. Evaluations on the SPE3R dataset demonstrate that OrbitGS effectively disentangles optical degradations from the geometric representation. Quantitatively, evaluated across seven space targets under moderate (200-view) and extreme sparse (50-view) settings, our framework achieves state-of-the-art robustness against extreme degradations. Notably, it avoids the catastrophic structural blow-ups observed in baseline methods, yielding an average geometric F1-score of 0.80 and a Chamfer Distance of 0.818, alongside a rendering Structural Similarity Index (SSIM) of 0.82 and a Learned Perceptual Image Patch Similarity (LPIPS) of 0.16. By preserving delicate structures under severe degradation, OrbitGS provides a robust, high-fidelity 3D reconstruction solution for complex orbital environments. Full article
Show Figures

Figure 1

40 pages, 5562 KB  
Article
Scene-Prompt-Driven Dynamic Routing Expert Network for Open-World Person Re-Identification
by Peng Dong, Hongbin Liu, Xiuyi Guo, Jilong Li, Yitong Zhou and Baoxu Wang
Mathematics 2026, 14(15), 2716; https://doi.org/10.3390/math14152716 - 31 Jul 2026
Viewed by 322
Abstract
Open-world person re-identification (ReID) faces severe spatially asymmetric interferences, such as partial occlusion and illumination distortion. Existing models adopting static weight fusion irreversibly corrupt identity representations when processing localized noise. To address this, we propose a Scene-Prompt-Driven Dynamic Routing Expert Network (SPDR-Net). Guided [...] Read more.
Open-world person re-identification (ReID) faces severe spatially asymmetric interferences, such as partial occlusion and illumination distortion. Existing models adopting static weight fusion irreversibly corrupt identity representations when processing localized noise. To address this, we propose a Scene-Prompt-Driven Dynamic Routing Expert Network (SPDR-Net). Guided by spatial topological prior constraints, SPDR-Net treats person semantic parts as independent local experts. First, a semantic expert module decouples and refines local features using human spatial topology priors. Second, a Prompt-Guided Dynamic Routing (PGDR) network extracts high-level scene contexts to dynamically evaluate each expert’s reliability, assigning routing weights to attenuate noise propagation. Finally, a global–local fusion module superimposes high-purity local features onto a global identity anchor, maintaining topological integrity to generate a unified descriptor for robust person re-identification. Extensive experiments on seven public benchmarks and a newly constructed real-world street dataset (SD-ReID) demonstrate that SPDR-Net achieves highly competitive performance against state-of-the-art methods and exhibits superior robustness under severe spatial-asymmetric interferences. Furthermore, end-to-end multi-camera closed-loop tests verify its robust decision-making capability and high engineering application value in real-world security systems. Full article
(This article belongs to the Special Issue Mathematical Computation for Pattern Recognition and Computer Vision)
Show Figures

Figure 1

30 pages, 23735 KB  
Article
SGDC-UIE: A Semantic Guidance Network with Degradation Consistency for Underwater Image Enhancement
by Rui Ming, Jianshan Zhang, Taotao Lai, Haibo Luo and Jiancheng Yang
J. Mar. Sci. Eng. 2026, 14(15), 1366; https://doi.org/10.3390/jmse14151366 - 25 Jul 2026
Viewed by 338
Abstract
Underwater images often suffer from color distortion, low contrast, and structural blurring caused by wavelength-dependent absorption and scattering, which degrade both visual observation and downstream perception. Existing underwater image enhancement methods usually learn image-level restoration mappings, while the relationships among semantic regions, degradation [...] Read more.
Underwater images often suffer from color distortion, low contrast, and structural blurring caused by wavelength-dependent absorption and scattering, which degrade both visual observation and downstream perception. Existing underwater image enhancement methods usually learn image-level restoration mappings, while the relationships among semantic regions, degradation patterns, and restoration responses are not fully exploited. In this paper, we propose a Semantic Guidance Network with Degradation Consistency for Underwater Image Enhancement (SGDC-UIE). Specifically, SGDC-UIE first extracts dense semantic responses from a frozen DINOv3 prior and converts them into foreground, boundary, and background region gates. These gates are then used to guide pseudo-physical degradation estimation, producing attenuation-like, transmission-like, illumination, structure, and background-light priors for region-aware restoration. These pseudo-physical priors are learned, bounded conditioning variables rather than calibrated estimates of underwater optical parameters. Based on these degradation conditions, a dual-branch restoration network corrects low-frequency color and illumination degradation while recovering high-frequency structural details through semantic-aware wavelet restoration. The color-restored and structure-restored outputs are further integrated by a degradation-consistent fusion gate, which adaptively balances visual fidelity and task-relevant structure preservation. In addition, grouped supervision with quality-anchor replay stabilizes task-aware fine-tuning and reduces visual-quality drift. Extensive experiments on paired and no-reference underwater enhancement benchmarks, semantic segmentation, and underwater object detection show that SGDC-UIE achieves competitive restoration quality and improves the usability of enhanced images for downstream perception. Full article
Show Figures

Figure 1

17 pages, 6118 KB  
Article
Relation-Aware Dual-View Graph Contrastive Learning with Huber Covariance Whitening
by Ahmed El Badaoui, Abdellah Ezzati, Said Ben Alla, Manal Hilali and Hicham Ben Alla
AI 2026, 7(8), 276; https://doi.org/10.3390/ai7080276 - 23 Jul 2026
Viewed by 344
Abstract
Self-supervised graph collaborative filtering suffers from two geometric pathologies. In Dimensional Collapse, embeddings quietly collapse into a low-rank subspace, a well-documented but poorly solved problem. Semantic Collapse is more complex: stiff L2-squared orthogonalization penalties that are supposed to push the embeddings [...] Read more.
Self-supervised graph collaborative filtering suffers from two geometric pathologies. In Dimensional Collapse, embeddings quietly collapse into a low-rank subspace, a well-documented but poorly solved problem. Semantic Collapse is more complex: stiff L2-squared orthogonalization penalties that are supposed to push the embeddings apart end up ripping through the heavy-tailed community overlaps that contain the collaborative signal. Earlier studies attempted to mitigate sparsity by injecting static noise or structural perturbations, but such interventions did not pinpoint the root cause, i.e., the distortions in the global covariance geometry itself. In this paper, we propose a Huber-Contrastive Graph Convolutional Network (HCGCN) that combines a spatial message-passing backbone with an O(1) contrastive augmentation overhead and a Relation-Aware Dual-View Gated Contrastive Network. The main novelty is a Huber Covariance Whitening module that imposes a geometry-aware threshold on the cross-correlation matrix of augmented views—below the threshold, the gradients follow an L2 penalty (enforcing uniformity); above it, the penalty flattens to L1 (protecting genuine semantic clusters from gradient explosion). This theoretically motivated dual-regime penalty actively preserves the macro-semantic topology of the graph while aggressively stamping out spurious noise correlations. The HCGCN is evaluated on the Yelp2018, Amazon-Book, and MovieLens datasets and performs significantly better than state-of-the-art baselines like LightGCN, SGL, SimGCL, and NESCL, especially under severe cold-start settings where covariance regulation proves most critical. Full article
(This article belongs to the Section AI Systems: Theory and Applications)
Show Figures

Graphical abstract

27 pages, 10988 KB  
Article
Text Image Super-Resolution via Fusion of OCR Priors and Cross-Scale Attention
by Xinyu Qiu, Jingchao Liu and Chen Fang
Symmetry 2026, 18(7), 1221; https://doi.org/10.3390/sym18071221 - 20 Jul 2026
Viewed by 433
Abstract
Text image super-resolution aims to improve the readability of low-quality text images while preserving character structures, stroke details, and semantic consistency. Compared with natural image super-resolution, this task is more sensitive to structural distortion because small changes in stroke topology may lead to [...] Read more.
Text image super-resolution aims to improve the readability of low-quality text images while preserving character structures, stroke details, and semantic consistency. Compared with natural image super-resolution, this task is more sensitive to structural distortion because small changes in stroke topology may lead to incorrect text recognition. To address this problem, this paper proposes an OCR prior-guided cross-scale framework for text image super-resolution. Specifically, character-level semantic priors extracted from a pretrained OCR model are introduced to provide structural guidance for degraded text reconstruction. A gated feature modulation mechanism is designed to adaptively regulate the contribution of OCR priors, reducing the influence of unreliable semantic predictions. A cross-scale dynamic attention module is also developed to aggregate multi-granularity visual features, enabling the model to jointly recover fine stroke boundaries and global character structures. In addition, a sequence-aware calibration module is introduced to improve structural consistency along the logical reading order of text. Experiments on mixed text image benchmarks and the TextZoom dataset show that the proposed method achieves competitive or better performance among the compared methods in terms of PSNR, SSIM, and recognition-oriented metrics. Additional ablation, OCR prior robustness, and computational complexity analyses further indicate that the proposed framework improves text readability while maintaining a reasonable accuracy–complexity trade-off. The results also suggest that OCR priors are useful for text image reconstruction, but should be used as soft constraints when external recognition predictions are uncertain. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

12 pages, 13095 KB  
Proceeding Paper
A Hybrid Synthetic Dataset Generation for Robust Document Recognition Using Image Rendering and Domain Randomization
by Plamen Nakov, Petar Petrov, Georgi Kotov, Milena Lazarova and Ognyan Nakov
Eng. Proc. 2026, 150(1), 10; https://doi.org/10.3390/engproc2026150010 - 16 Jul 2026
Viewed by 333
Abstract
Automated recognition of identity documents is a critical component in digital identity verification systems. The development of robust recognition models is often constrained by the limited availability of large, diverse, high-quality, and publicly accessible ID card datasets. Collecting and annotating real-world ID card [...] Read more.
Automated recognition of identity documents is a critical component in digital identity verification systems. The development of robust recognition models is often constrained by the limited availability of large, diverse, high-quality, and publicly accessible ID card datasets. Collecting and annotating real-world ID card images is time-consuming, and often restricted due to privacy, legal, and security concerns. The paper proposes a novel approach for generating a large-scale synthetic dataset for ID card recognition by merging real ID card images with a texture dataset through a structured data fusion pipeline that introduces realistic visual variations as illumination effects, geometric distortions, and noise patterns while preserving the semantic integrity of the original ID card content. Full article
Show Figures

Figure 1

24 pages, 10388 KB  
Article
Adaptive Content and Style Fusion for Text-to-Image Generations
by Yi-Fang Lee, Chun-Chieh Lee, Chi-Hung Chuang, Chih-Lung Lin and Kuo-Chin Fan
Electronics 2026, 15(13), 2800; https://doi.org/10.3390/electronics15132800 - 25 Jun 2026
Viewed by 456
Abstract
Text-to-image generation aims to produce images that match the semantic content of a text prompt. In style transfer tasks, the model must further integrate reference styles while preserving prompt semantics. However, balancing semantic consistency and style fidelity remains challenging. Existing methods commonly rely [...] Read more.
Text-to-image generation aims to produce images that match the semantic content of a text prompt. In style transfer tasks, the model must further integrate reference styles while preserving prompt semantics. However, balancing semantic consistency and style fidelity remains challenging. Existing methods commonly rely on fixed feature weights and lack adaptive control, which often leads to style over-injection and content distortion. To address these issues, we propose a novel framework that performs dynamic regulation at both the feature and temporal levels. At the feature level, we propose an Entropy-Aware Adaptive Fusion (EAAF) module. It incorporates a bidirectional distribution transformation mechanism to enhance the statistical correlation between content and style features. The module further uses information entropy as a dynamic control signal to adaptively adjust the strength of style injection, thereby achieving a balance between semantic consistency and style fidelity. At the temporal level, we design a Progressive Feature Reweighting (PFR) strategy. By applying stage-wise weighting to content and style features at different diffusion steps, this strategy effectively improves structural stability and color consistency. In addition, our framework is modular and can be integrated into existing diffusion-based style transfer models without additional fine-tuning or retraining. Experimental results demonstrate that applying our approach to current state-of-the-art models, such as StyleStudio and CSGO, significantly enhances their performance, particularly in maintaining strong prompt alignment while achieving high-fidelity style transfer. Full article
(This article belongs to the Special Issue Recent Advances in Object Detection and Computer Vision)
Show Figures

Figure 1

18 pages, 8476 KB  
Article
Dual-Pathway Wavelet-Attention Framework for Image-Only AI-Generated Image Quality Assessment
by Yang Li, Yu Zheng and Dong Sui
Mathematics 2026, 14(13), 2249; https://doi.org/10.3390/math14132249 - 23 Jun 2026
Viewed by 315
Abstract
AI-generated images (AIGIs) often contain perceptual defects that differ from the distortions commonly studied in conventional no-reference image quality assessment (NR-IQA). This work investigates image-only AIGC image quality assessment, where no prompt text is used and the quality score must be inferred from [...] Read more.
AI-generated images (AIGIs) often contain perceptual defects that differ from the distortions commonly studied in conventional no-reference image quality assessment (NR-IQA). This work investigates image-only AIGC image quality assessment, where no prompt text is used and the quality score must be inferred from visual evidence such as artifacts, structure, and semantic plausibility. We propose a dual-pathway wavelet-attention framework built on a Swin Transformer V2-Base backbone. The artifact pathway employs a Noise Perceptive Attention Module (NPAM) with fixed Haar wavelet decomposition to describe generation-related sub-band degradation cues, whereas the image-perception pathway models semantic, structural, and contextual quality evidence using multi-scale attention, global–local spatial-channel attention, and pyramid pooling. The two pathways are integrated through adaptive fusion and a spatially weighted regression head with an auxiliary global prediction. Experiments on AGIQA-1K, AGIQA-3K, and AIGCIQA2023 demonstrate competitive in-domain performance, including SRCC values of 0.8418 on AGIQA-3K and 0.8445 on the quality dimension of AIGCIQA2023. The evaluation further covers individual module ablations, score-fusion variants, seed stability, qualitative error analysis, and cross-database transfer, revealing both the contribution of the proposed components and the remaining difficulty of source-disjoint generalization. Full article
(This article belongs to the Section E1: Mathematics and Computer Science)
Show Figures

Figure 1

26 pages, 4854 KB  
Article
Class-Aware Semantic Calibration for Cross-Scene Hyperspectral Image Classification
by Boshan Shi, Yanbo Liu, Youqiang Zhang and Guo Cao
Remote Sens. 2026, 18(12), 1976; https://doi.org/10.3390/rs18121976 - 14 Jun 2026
Viewed by 300
Abstract
Cross-scene Hyperspectral Image (HSI) classification faces substantial domain shifts caused by sensor heterogeneity, acquisition variation, and scene diversity. While benchmark annotations are assigned to individual center pixels, local patches often contain implicit multi-label semantics due to spectral mixing and spatial overlap. This mismatch [...] Read more.
Cross-scene Hyperspectral Image (HSI) classification faces substantial domain shifts caused by sensor heterogeneity, acquisition variation, and scene diversity. While benchmark annotations are assigned to individual center pixels, local patches often contain implicit multi-label semantics due to spectral mixing and spatial overlap. This mismatch distorts prediction structure, exacerbates generalization errors, and limits the effectiveness of standard domain generalization (DG) techniques focused solely on feature or prediction invariance. We propose Class-Aware Semantic Calibration (CASC), a systematic semantic structure calibration framework that addresses three complementary distortions induced by mismatched patch supervision: (i) Balance corrects class frequency bias via reweighted supervision; (ii) Separability enhances boundary decision stability through margin-based logit calibration; and (iii) Independence reduces domain-specific spurious co-occurrence via prediction covariance decorrelation. To preserve calibrated semantics under pseudo-source shift, we further introduce a complementary DualAlign (DA) module, which jointly aligns feature statistics and prediction distributions, enforcing consistency at both representation and semantic levels. Extensive experiments on three cross-scene benchmarks (Houston, Pavia, and WHU-Hi) demonstrate that CASC-DA consistently improves performance over strong baselines, achieving an average gain of 3.0% in overall accuracy and 4.9% in Kappa coefficient compared with the best-performing baseline on each dataset. These results underscore the importance of semantic structure calibration for domain-generalized HSI classification. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

23 pages, 19029 KB  
Article
CETransUNet: An Intelligent Landslide Identification Method Based on Collaborative Optimization of Global Context and Dual Attention Mechanisms
by Tianli Sun, Chengsheng Yang, Jifeng Wu, Zewei Liu, Ziqian Wang and Xiaoqiang Cheng
Remote Sens. 2026, 18(12), 1974; https://doi.org/10.3390/rs18121974 - 13 Jun 2026
Cited by 1 | Viewed by 368
Abstract
Accurate landslide identification is crucial for enhancing emergency response capabilities during destructive geological hazards. Although deep-learning-based semantic segmentation has demonstrated effectiveness, substantial variations in landslide scales and environmental similarities continue to challenge existing methods. This paper systematically constructs a new co-seismic landslide dataset [...] Read more.
Accurate landslide identification is crucial for enhancing emergency response capabilities during destructive geological hazards. Although deep-learning-based semantic segmentation has demonstrated effectiveness, substantial variations in landslide scales and environmental similarities continue to challenge existing methods. This paper systematically constructs a new co-seismic landslide dataset for the Yarlung Zangbo River basin based on the 2017 Nyingchi earthquake, effectively filling a critical regional data gap. This paper proposes CETransUNet (coordinate attention and edge-guided attention transformer UNet), a novel landslide detection model that integrates ResNet and Transformer architectures. Specifically, a coordinate attention (CA) module is introduced within the skip connections between the encoder and decoder. This module encodes positional information along both horizontal and vertical spatial directions and dynamically re-weights the feature maps, thereby effectively suppressing background noise caused by semantic gaps and enhancing the model’s ability to localize landslide regions. Additionally, an edge-guided attention (EGA) module is incorporated into the decoder. This module extracts explicit edge priors from the input image using a Laplacian operator and imposes geometric constraints on the predictions via a boundary reverse attention mechanism, thereby significantly alleviating boundary ambiguity and morphological distortion of landslides. Evaluations across datasets from the Yarlung Zangbo River, Iburi-Tobu, and Bijie regions demonstrate that CETransUNet significantly outperforms state-of-the-art models—including TransUNet, SegFormer, and SwinUNet—in terms of IoU, MIoU, and F1-score. Overall, through the synergistic optimization of the coordinate attention and edge-guided attention modules, the CETransUNet model achieves synchronous enhancement of boundary integrity and geometric precision in complex scenarios, providing a reliable technical solution for large-scale intelligent landslide identification. Full article
Show Figures

Figure 1

Back to TopTop