Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (877)

Search Parameters:
Keywords = vision–language modeling

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
19 pages, 6457 KB  
Article
Real-Time Anomaly Detection on Edge Devices via VLM Prompt Optimization
by Sungmin Yu, Jongwon Moon and Hosub Yoon
Electronics 2026, 15(15), 3305; https://doi.org/10.3390/electronics15153305 (registering DOI) - 27 Jul 2026
Abstract
Real-time video anomaly detection (VAD) under realistic edge constraints—sub-second latency, ≤25 W power, no cloud dependency, and human-interpretable output—remains an open problem. Existing lightweight video convolutional neural networks (X3D, MoViNets) are bound to closed-set training distributions, while recent vision–language-model-based VAD methods (LAVAD, VERA, [...] Read more.
Real-time video anomaly detection (VAD) under realistic edge constraints—sub-second latency, ≤25 W power, no cloud dependency, and human-interpretable output—remains an open problem. Existing lightweight video convolutional neural networks (X3D, MoViNets) are bound to closed-set training distributions, while recent vision–language-model-based VAD methods (LAVAD, VERA, Holmes-VAD) achieve 80–89% area under the curve (AUC) but rely on datacenter-grade GPUs and Chain-of-Thought (CoT) reasoning that pushes per-segment latency well above one second. This paper reframes the design target from peak accuracy to practical edge deployability and contributes two tightly coupled designs: (i) an edge-optimized inference stack that compresses Qwen3-VL-2B with 4-bit Activation-aware Weight Quantization (INT4 AWQ) and serves it through a TensorRT-LLM C++ runtime on NVIDIA Jetson Orin NX (16 GB, 25 W); and (ii) a fully automatic, CoT-free verbalized prompt optimization in which an 8B optimizer iteratively refines a natural-language definition block Dt using class-balanced (stratified) development batches on a disjoint development subset, with no human editing and no runtime cost on the edge device. Three findings support this framing: (a) the inference stack reduces per-segment latency to 0.25 s, a 7.4× speed-up and 55% memory reduction over a Python/PyTorch baseline; (b) verbalized prompt optimization improves zero-shot AUC from 71.82% (manual prompt) to 76.39%, outperforming GPT-4- and Gemini-Pro-generated prompts (74.12% and 74.35%) under the same edge backbone; and (c) single-frame input attains the highest mean AUC among one-, five-, and eight-frame windows—statistically comparable to the five-frame setting—while offering the lowest latency, making it the preferred operating point under the edge budget. While the absolute AUC (76.39%) is below recent server-side methods (CLIP-TSA 87.58%, VadCLIP 88.02%, Holmes-VAD 89.51%), our framework is the only one in this comparison that operates entirely on a ≤25 W edge device, providing a deployment-oriented operating point on the accuracy–feasibility frontier of VLM-based VAD. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

42 pages, 3402 KB  
Review
Computational Recipe Intelligence: A Survey of Recipe Design, Generation, Recommendation, and Evaluation Protocols
by Zihan Song and Bai Li
Electronics 2026, 15(15), 3281; https://doi.org/10.3390/electronics15153281 - 25 Jul 2026
Viewed by 39
Abstract
Recent advances in artificial intelligence have transformed recipes from static textual artifacts into computational objects that support design, generation, and personalized recommendation. However, existing studies remain fragmented, typically addressing these tasks in isolation and lacking a unified analytical framework. This paper presents a [...] Read more.
Recent advances in artificial intelligence have transformed recipes from static textual artifacts into computational objects that support design, generation, and personalized recommendation. However, existing studies remain fragmented, typically addressing these tasks in isolation and lacking a unified analytical framework. This paper presents a comprehensive survey of recipe intelligence by systematically integrating recipe design, recipe generation, and recipe recommendation within a coherent computational framework. We review representative methods across these three tasks, including knowledge-based design systems, statistical and neural sequence generation models, cross-modal vision-language generation methods, pretraining- and retrieval-augmented generation approaches, and behavior-, context-, semantic-, and goal-driven recommendation methods. Particular attention is given to how these methods represent culinary knowledge, model user preferences, handle constraints, and support personalization. We further summarize commonly used datasets, data acquisition and processing strategies, and evaluation protocols, thereby clarifying the empirical foundations of current recipe intelligence research. Finally, we discuss key challenges and outline future directions toward unified, controllable, and interactive recipe intelligence systems. Full article
Show Figures

Figure 1

24 pages, 16370 KB  
Article
Unifying Inconsistent Emotion Labeling Criteria Across Datasets via Prototype-Guided Multimodal Alignment for Facial Expression Recognition
by Junjie Liu, Yufei Xie, Dong Zhang and Dah-Jye Lee
Electronics 2026, 15(15), 3279; https://doi.org/10.3390/electronics15153279 - 25 Jul 2026
Viewed by 59
Abstract
Facial expression recognition (FER) plays an important role in human–computer interaction and affective computing. Although combining multiple FER datasets in joint training can potentially improve model generalization, it is hindered by inconsistent emotion labeling criteria across datasets. To address this issue, we propose [...] Read more.
Facial expression recognition (FER) plays an important role in human–computer interaction and affective computing. Although combining multiple FER datasets in joint training can potentially improve model generalization, it is hindered by inconsistent emotion labeling criteria across datasets. To address this issue, we propose a Prototype-Guided Multimodal Alignment Joint Training framework for multi-dataset FER. The core idea is to leverage the image–text alignment knowledge learned by Vision–Language Models from the target dataset as a unified emotion labeling criterion across datasets, while deriving prototypes from visual features to serve as emotion anchors for reliable feature alignment in the latent space. Based on the consistency among semantic predictions, prototype-distance predictions, and auxiliary labels, semantically consistent auxiliary samples are selected for joint training under a unified labeling criterion. Extensive experiments show that the proposed framework, with the Amending Representation Module as the backbone network, achieves 93.74% accuracy on the RAF-DB dataset and 98.31% on the CAER-S dataset, attaining state-of-the-art performance. The experimental results demonstrate that establishing semantically consistent labeling criteria across datasets is an effective strategy for multi-dataset FER learning. Full article
Show Figures

Figure 1

52 pages, 38034 KB  
Article
An Augmented Reality and AI-Based System for Contextual Appliance Guidance: Implications for Cognitive Accessibility and Assistive Interaction
by Kimia Hafezi, Atra Hossein Tafreshi, Christian Napoli, Cristian Randieri and Samuele Russo
Brain Sci. 2026, 16(8), 783; https://doi.org/10.3390/brainsci16080783 (registering DOI) - 24 Jul 2026
Viewed by 67
Abstract
Background: Modern household appliances often present complex interfaces that can be difficult to use, especially for older adults, people with visual impairments, and users with mild cognitive difficulties. In such cases, interacting with appliances may require sustained attention, visuospatial search, working memory, and [...] Read more.
Background: Modern household appliances often present complex interfaces that can be difficult to use, especially for older adults, people with visual impairments, and users with mild cognitive difficulties. In such cases, interacting with appliances may require sustained attention, visuospatial search, working memory, and sequential action planning, while traditional user manuals often provide limited contextual support. Methods: To address this issue, this study presents a proof-of-concept augmented reality (AR) and artificial intelligence (AI)-based system for contextual appliance guidance. The proposed architecture integrates visual sensing, deep learning, and large language models to detect appliance controls, interpret user queries, retrieve relevant information from user manuals, and provide step-by-step guidance directly on the real interface. A YOLOv8 model trained on a custom dataset was used for button detection, YOLO-Seg was employed to enhance visual highlighting through segmentation, and BoT-SORT was used to maintain detection consistency across frames. A Unity-based mobile application displayed real-time AR overlays with customizable visual settings for accessibility needs, such as low vision and color blindness that may be relevant for future accessibility-oriented applications. In addition to its technical pipeline, the system is conceptually relevant as a potential form of external cognitive support because it transforms static manual instructions into situated, sequential, and visually grounded guidance.Results: Experimental results showed promising technical performance for button detection and segmentation, while a preliminary user evaluation in a non-clinical sample suggested good usability, clarity, and acceptability of the interface. Conclusions: These findings support the technical feasibility and preliminary usability of the approach, while cognitive workload, confidence, functional autonomy, and clinical benefit were not directly measured. Targeted validation in older adults, people with visual impairments, and clinical populations is therefore still needed. Future developments will include multimodal feedback, read-aloud guidance, and more specific evaluation of workload, confidence, and functional autonomy. Full article
(This article belongs to the Section Neural Engineering, Neuroergonomics and Neurorobotics)
24 pages, 2860 KB  
Article
CGF-Net: A Multi-View Contrastive Learning Model for Encrypted Traffic Classification
by Yanlin He and Ning Hu
Electronics 2026, 15(15), 3249; https://doi.org/10.3390/electronics15153249 - 23 Jul 2026
Viewed by 118
Abstract
For the task of encrypted traffic classification, existing approaches commonly rely on a single feature view, such as side-channel characteristics or raw packet bytes, often incorporating techniques inspired by natural language processing and computer vision for modeling and classification. In recent years, multi-view [...] Read more.
For the task of encrypted traffic classification, existing approaches commonly rely on a single feature view, such as side-channel characteristics or raw packet bytes, often incorporating techniques inspired by natural language processing and computer vision for modeling and classification. In recent years, multi-view learning has gained increasing attention due to its ability to enhance discriminative power and generalization performance by capturing complementary information from different perspectives. However, heterogeneous feature views often exhibit distributional discrepancies, which makes direct multi-view integration difficult. To address this issue, we propose CGF-Net, a multi-view contrastive learning framework for encrypted traffic classification. The proposed model is inspired by cross-modal contrastive learning and employs lightweight adaptation to learn representations from both behavioral and content views. During pre-training, CGF-Net performs instance-level cross-view contrastive learning by treating the behavioral and content views of the same network flow as a positive pair, thereby aligning heterogeneous representations at the flow-instance level. In addition, a lightweight fine-tuning module together with a gating-based fusion mechanism is introduced to improve the collaborative modeling capability of multi-view representations. Extensive experiments on four public datasets show that CGF-Net achieves ACC scores of 95.81%, 93.38%, 96.38%, and 95.72% on CSTNET-TLS1.3, CipherSpectrum, ISCXVPN2016, and ISCXTor2016, respectively. Compared with the best-performing baseline on each dataset, CGF-Net improves the average ACC and F1-score by 0.78 and 0.69 percentage points, respectively, demonstrating the effectiveness of the proposed model. Full article
(This article belongs to the Section Networks)
Show Figures

Figure 1

37 pages, 12135 KB  
Article
A Hierarchical VLM-to-TD3 Framework with Novel Object Coordinate Estimation and Persistent Spatial Memory for Semantically Guided Indoor Navigation
by Yernar Akhmetbek, Ayaulym Parmash, Temirlan Meiramkhanov, Azamat Yesmukhametov, Aigul Meirmanova and Darkhan Zholtayev
Sensors 2026, 26(15), 4661; https://doi.org/10.3390/s26154661 - 23 Jul 2026
Viewed by 211
Abstract
Autonomous semantic indoor navigation requires robust low-level control and high-level understanding of objects and spatial context in cluttered and partially occluded environments. While deep reinforcement learning (DRL) methods such as twin delayed deep deterministic policy gradient (TD3) enable reactive obstacle avoidance, they typically [...] Read more.
Autonomous semantic indoor navigation requires robust low-level control and high-level understanding of objects and spatial context in cluttered and partially occluded environments. While deep reinforcement learning (DRL) methods such as twin delayed deep deterministic policy gradient (TD3) enable reactive obstacle avoidance, they typically struggle with long-horizon semantic navigation, where object-location memory and language-level reasoning are required. We present a lightweight hierarchical two-stage framework that, to the best of our knowledge, is introduced for the first time to integrate a locally deployed vision–language model (VLM), semantic object coordinate memory, and a TD3-based DRL controller for language-conditioned indoor navigation. In Stage 1, the robot performs semantic exploration using odometry, 2D LiDAR, and VLM-based object recognition to build a geometric map and store detected object categories with their estimated world coordinates in a structured javaScript object notation (JSON) semantic memory. In Stage 2, a natural language query is used to retrieve the target object coordinates from memory and pass them to a TD3 target point navigation policy, which performs mapless navigation using odometry and RealSense RGB-D perception. The proposed framework combines open-vocabulary VLM-based object coordinate estimation, LiDAR mapping, RGB-D perception, and language grounding within a unified semantic memory representation. Experiments in a ROS-integrated realistic simulation demonstrate consistent goal-reaching performance and improved navigation efficiency compared with an Artificial Potential Field baseline using the same VLM and a DRL + GPT-4o mini configuration. We also compare the proposed VLM-based recognition module with YOLO-World v2.6 and Grounding DINO, showing that the VLM-based approach provides more reliable semantic grounding and target-coordinate estimation in the tested indoor navigation scenarios, particularly for flexible natural language object queries. Full article
(This article belongs to the Special Issue AI-Powered Vision Sensing for Autonomous Driving)
Show Figures

Figure 1

31 pages, 15220 KB  
Article
SRGFormer: Semantic Role-Guided Graph Reasoning for Referring Remote Sensing Image Segmentation
by Libang Liu, Jianxiang Li, Yaqin Li, Cao Yuan, Lili Fan, Xinyu Xiong and Wei Huang
Sensors 2026, 26(14), 4657; https://doi.org/10.3390/s26144657 - 22 Jul 2026
Viewed by 157
Abstract
Referring remote sensing image segmentation (RRSIS) aims to segment a target instance from remote sensing imagery according to a natural-language expression. It provides a flexible way to retrieve and localize specific objects in remote sensing scenes, benefiting intelligent Earth observation applications. Although existing [...] Read more.
Referring remote sensing image segmentation (RRSIS) aims to segment a target instance from remote sensing imagery according to a natural-language expression. It provides a flexible way to retrieve and localize specific objects in remote sensing scenes, benefiting intelligent Earth observation applications. Although existing methods have achieved promising progress by strengthening vision–language alignment, most of them still represent the expression as a holistic language feature and rely on convolution-dominated decoding for mask prediction. Such a paradigm tends to entangle target category, inter-object relation, and spatial position cues, making it difficult to distinguish the intended instance from multiple same-class distractors in complex remote sensing scenes. To address this limitation, we propose SRGFormer, a graph reasoning framework for RRSIS. Specifically, a semantic role decomposition (SRD) module decomposes the referring expression into target, relation, and position semantics, providing explicit linguistic priors for instance-level localization. Guided by the decomposed relation semantics, a semantic-relational graph transformer (SRGT) performs relation-aware graph reasoning over fused multi-scale visual features, enabling long-range dependency modeling among spatially distributed candidate instances. Furthermore, a progressive mask refinement (PMR) module continuously injects the decomposed semantic priors into semantic modulation, query initialization, and iterative mask decoding, thereby alleviating semantic fading during mask generation. Extensive experiments demonstrate that SRGFormer achieves substantial improvements on RefSegRS, attaining 66.08% mIoU and 76.93% oIoU (surpassing the prior state of the art by 3.96% and 2.83%, respectively) along with a notable 15.95% gain in Pr@0.7. Experiments on the additional RRSIS-D benchmark further demonstrate the general applicability of our approach, where SRGFormer maintains competitive performance (65.87% mIoU and 24.61% Pr@0.9) against existing methods. These results demonstrate that the proposed framework improves target localization and fine-grained mask prediction in complex remote sensing scenes. Full article
Show Figures

Figure 1

31 pages, 6714 KB  
Article
A Lightweight Vision-Language-Action Policy with Progress-Aware Hybrid Execution for UAV Waypoint Navigation in AirSim
by Yiqing Xu, Haifeng Lin, Yujin Yang, Ji’An Xia and Zidong Han
Sensors 2026, 26(14), 4655; https://doi.org/10.3390/s26144655 - 22 Jul 2026
Viewed by 179
Abstract
Offline action prediction does not by itself guarantee reliable closed-loop flight for unmanned aerial vehicle (UAV) vision-language-action (VLA) models. We study a controlled AirSim Blocks waypoint task using 100 expert episodes and 3385 RGB-D, instruction, state, and action records. Checkpoints are selected only [...] Read more.
Offline action prediction does not by itself guarantee reliable closed-loop flight for unmanned aerial vehicle (UAV) vision-language-action (VLA) models. We study a controlled AirSim Blocks waypoint task using 100 expert episodes and 3385 RGB-D, instruction, state, and action records. Checkpoints are selected only on val-seen data, after which val-unseen is evaluated once. A 132,840-parameter policy reaches 0.9278±0.0019 final-test action accuracy across three training seeds, yet raw VLA control fails in closed loop. We therefore embed its action proposals in progress-aware hybrid execution with explicit recovery and near-goal precision. The strongest checkpoint reaches 59/60 goals, but crossing three independently trained checkpoints with three target seeds yields a more conservative 147/180 successes (81.7%) with zero recorded collisions and marked checkpoint sensitivity. Substantial overrides and fallback-tagged steps further show that the reported closed-loop outcomes are properties of the hybrid system, not of the learned policy alone. These findings are restricted to the controlled AirSim Blocks benchmark and do not demonstrate real-UAV deployment, sim-to-real transfer, or field robustness. Full article
(This article belongs to the Section Sensors and Robotics)
Show Figures

Figure 1

23 pages, 19255 KB  
Article
CLIFF: A Multi-Modal Remote Sensing Model for Geological Hazard Monitoring Based on Bitemporal UAV Images
by Quanxi Zhou, Qianxiao Su, Xinran Wei, Wencan Mao, Yili Ren, Yunfei Chen, Jianzhong Bi, Mingjun Zhao and Manabu Tsukada
Remote Sens. 2026, 18(14), 2432; https://doi.org/10.3390/rs18142432 - 22 Jul 2026
Viewed by 241
Abstract
UAV-based remote sensing excels in rapid response, high timeliness, simple operation, and high degrees of automation, and has been widely applied for geological hazard monitoring. Deep learning methods based on unitemporal UAV images can only analyze the static appearance of a scene, while [...] Read more.
UAV-based remote sensing excels in rapid response, high timeliness, simple operation, and high degrees of automation, and has been widely applied for geological hazard monitoring. Deep learning methods based on unitemporal UAV images can only analyze the static appearance of a scene, while bitemporal change detection can capture the dynamic evolution of hazards; however, due to diverse geological landforms and topography, environmental noises such as vegetation cover, and dynamic weather conditions, change detection of geological hazards from UAV images based on traditional deep learning technology is not always effective. Therefore, there is an urgent need to utilize large vision-language models (LVLMs) to further improve the accuracy and robustness of the change detection model. Motivated by this, this paper proposes a novel remote sensing model for geological hazard monitoring, referred to as CLIFF (CLIP-BIT-EfficientNet), based on the multi-modal LVLM Contrastive Language–Image Pre-training (CLIP), the change detection network Bitemporal Image Transformer (BIT), and the classification network EfficientNet, along with corresponding datasets and model fine-tuning strategies. The proposed transfer fusion module bridges the CLIFF and BIT networks by aligning their feature distributions and dimensions, allowing the general knowledge of the LVLM and the task-specific knowledge of the learnable branch to reinforce each other. Furthermore, this integrated pipeline addresses the scarcity of labeled hazard data by allowing the BIT to train on larger public datasets, while fine-tuning EfficientNet on smaller hazard-classification datasets within the change area, making the approach more efficient and reliable than direct classification methods. Experimental results show that the proposed CLIFF algorithm outperforms state-of-the-art deep learning algorithms such as LightCDNet and ChangeFormer, with an IoU of 75.74% and an F1 score of 0.8689 for change detection. Meanwhile, CLIFF has an overall accuracy rate of 86.89% in identifying geological hazards along gas pipelines, such as crude oil spills, collapses, landslides, and floods, with per-class accuracies of 87.32% and 86.17% for crude oil spills and landslides, respectively. Full article
Show Figures

Figure 1

19 pages, 813 KB  
Article
Cross-Modal Variance-Aware KV Cache Optimization for Efficient Multimodal Long-Context Inference
by Shenglong Liu, Yanli Lv, Siyao An, Chenghuan Yu and Yiwei Ru
Electronics 2026, 15(14), 3206; https://doi.org/10.3390/electronics15143206 - 21 Jul 2026
Viewed by 194
Abstract
Multimodal large language models (MLLMs) face substantial memory bottlenecks when processing long visual contexts, such as videos and high-resolution images. Existing methods that allocate visual KV cache budgets using cross-modal attention entropy mainly estimate the distributional breadth of text–vision interaction and may overlook [...] Read more.
Multimodal large language models (MLLMs) face substantial memory bottlenecks when processing long visual contexts, such as videos and high-resolution images. Existing methods that allocate visual KV cache budgets using cross-modal attention entropy mainly estimate the distributional breadth of text–vision interaction and may overlook how visual relevance varies across query positions in the encoded multimodal context. We propose a cross-modal query-position variance-aware KV cache optimization method for efficient multimodal long-context inference. The proposed method combines cross-modal attention entropy with prefill-stage query-position variance computed from cross-modal attention to estimate layer-wise visual KV cache preferences. Based on this preference score, visual KV cache budgets are allocated across layers, and variance-aware token pruning is applied to retain high-importance KV states while directly evicting redundant visual tokens without feature merging. Experiments on the MileBench benchmark using LLaVA-v1.5-7B show that, under deterministic single-run evaluation and while retaining only 20% of the visual KV cache, the proposed method produces point-estimate performance close to the full-cache reference and higher point estimates on several fine-grained reasoning and retrieval subtasks. Additional representative-subtask evaluations under different visual cache budgets and on InternVL2.5-8B further provide preliminary point-estimate evidence that the proposed allocation signal is not restricted to a single cache ratio or backbone. System profiling further shows that the 20% cache setting reduces measured KV cache GPU memory from 1.28 GiB to 0.26 GiB and decoding latency from 100.28 ms/token to 92.85 ms/token. These results suggest that cross-modal query-position variance may help preserve sparse, query-dependent visual cues under low-cache-budget multimodal inference. Full article
Show Figures

Figure 1

31 pages, 9859 KB  
Article
Unified UAV Open-Vocabulary Semantic Segmentation: Benchmark Construction and LLM-Guided Text–Visual Enhancement
by Kun Wang, Wei Li, Xiaopeng Liu, Duping Huang and Jiale Yang
Remote Sens. 2026, 18(14), 2408; https://doi.org/10.3390/rs18142408 - 20 Jul 2026
Viewed by 249
Abstract
Open-vocabulary semantic segmentation (OVSS) of unmanned aerial vehicle (UAV) imagery aims to recognize arbitrary text-specified categories in aerial scenes, but existing OVSS models often suffer from UAV-domain shifts. To provide a reproducible testbed, this paper constructs a unified UAV OVSS benchmark by reorganizing [...] Read more.
Open-vocabulary semantic segmentation (OVSS) of unmanned aerial vehicle (UAV) imagery aims to recognize arbitrary text-specified categories in aerial scenes, but existing OVSS models often suffer from UAV-domain shifts. To provide a reproducible testbed, this paper constructs a unified UAV OVSS benchmark by reorganizing multiple UAV segmentation datasets into cross-dataset transfer settings with explicit category harmonization and seen/unseen vocabulary analysis. Based on this benchmark, we propose UAV-OVSeg, a Cost Aggregation-style dense matching framework enhanced in two complementary directions: an LLM-guided Category Expansion Module that converts raw category names into structured UAV-aware descriptions, and a DINO-enhanced Geometric Feature Fusion Module that injects local structure into dense visual–text matching. Under SynDrone training, UAV-OVSeg achieves 65.5% mean intersection over union (mIoU) and 78.0% mean accuracy (mACC), improving upon the CAT-Seg baseline by 2.6 mIoU and 3.2 mACC. Under Aeroscapes training, it achieves 64.5% mIoU and 76.8% mACC, improving upon CAT-Seg by 2.6 mIoU and 3.1 mACC. Additional analyses of LLM variants, object scales, prompt sensitivity, boundary quality, and computational cost further verify the effectiveness and reproducibility of the proposed UAV-oriented text–visual enhancement strategy. Full article
Show Figures

Figure 1

24 pages, 1891 KB  
Article
Cross-Attention-Driven Propose-and-Select Generative Data Augmentation for Few-Shot Image Classification
by Ying Liu, Liaomo Zheng and Shiyu Wang
Sensors 2026, 26(14), 4590; https://doi.org/10.3390/s26144590 - 20 Jul 2026
Viewed by 240
Abstract
Generative data augmentation based on diffusion models has emerged as a promising approach for few-shot image classification. Existing methods, such as DA-Fusion, typically follow a “generate-once, use-directly” paradigm, which often suffers from uncontrollable generation quality, unstable semantic consistency, insufficient global diversity, and high [...] Read more.
Generative data augmentation based on diffusion models has emerged as a promising approach for few-shot image classification. Existing methods, such as DA-Fusion, typically follow a “generate-once, use-directly” paradigm, which often suffers from uncontrollable generation quality, unstable semantic consistency, insufficient global diversity, and high sample redundancy. To address these limitations, we propose a two-stage Propose-and-Select framework for controllable data augmentation. This framework curates high-quality synthetic data offline, ensuring that no additional training overhead is introduced to downstream models. For selector optimization, our method eliminates the need for additional human annotations by leveraging the zero-shot prior knowledge of a vision–language model (CLIP) to construct relative-quality pseudo-labels. Furthermore, we develop an adaptive-temperature listwise ranking distillation objective to transfer quality-aware supervision effectively. We also introduce a multi-objective consistency regularization strategy to stabilize training and improve convergence. Under a strictly controlled augmentation budget, where all methods are provided with the same number of synthetic samples, the proposed approach consistently outperforms existing diffusion-based augmentation baselines across both few-shot classification benchmarks, achieving an accuracy of 79.58% on PASCAL VOC and 80.74% on the fine-grained Oxford 102 Flowers dataset. These results demonstrate the effectiveness of the proposed generation-selection paradigm in improving the quality, diversity, and semantic relevance of synthetic samples, thereby enhancing downstream few-shot classification performance. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

21 pages, 10432 KB  
Article
Projection-Free CLIP-Scale EEG Latents via a U-Net-Style Autoencoder
by Jeyoung Lee, Jaekwan Ahn, Jaeseung Sim and Hochul Kang
Sensors 2026, 26(14), 4583; https://doi.org/10.3390/s26144583 - 20 Jul 2026
Viewed by 234
Abstract
Electroencephalography is emerging as a promising conditioning modality for generative visual models. However, existing representation learning approaches often rely on high-capacity masked autoencoders and complex projection networks. When constrained to compact embedding dimensions to match vision-language models, these heavy transformer-based bottlenecks frequently suffer [...] Read more.
Electroencephalography is emerging as a promising conditioning modality for generative visual models. However, existing representation learning approaches often rely on high-capacity masked autoencoders and complex projection networks. When constrained to compact embedding dimensions to match vision-language models, these heavy transformer-based bottlenecks frequently suffer from representation collapse and lose critical signal dynamics. To address this, we propose a lightweight and projection-free autoencoder that directly outputs compact, Contrastive Language–Image Pre-training (CLIP)-scale latent vectors trained toward the CLIP embedding space. Our model adopts a U-Net-style architecture combining one-dimensional convolutional residual blocks for temporal dynamics and inter-channel attention modules for spatial dependencies, alongside skip connections to ensure stable reconstruction. Extensive experiments on visual perception datasets demonstrate that our approach successfully tracks complex signal amplitudes without collapsing. Under strict dimensional constraints, the proposed model achieves superior signal reconstruction fidelity across time and frequency domains using significantly fewer parameters than traditional masked autoencoder baselines. Furthermore, latent space visualizations and zero-shot retrieval tasks reveal that while the baseline collapses toward unstructured, near-chance representations, our architecture preserves emerging, partial semantic organization and retrieves several times above chance. This indicates that the proposed design preserves signal structure while exhibiting preliminary, above-chance semantic alignment, enabling integration into brain-driven generative pipelines. Full article
(This article belongs to the Special Issue Biosignal Sensing Analysis (EEG, EMG, ECG, PPG) (3rd Edition))
Show Figures

Figure 1

38 pages, 3059 KB  
Review
Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions
by Duc Tri Phan and Hong Duc Nguyen
J. Imaging 2026, 12(7), 324; https://doi.org/10.3390/jimaging12070324 - 17 Jul 2026
Viewed by 167
Abstract
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion [...] Read more.
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion blur, and limited field-of-view or traditional fixed-camera multi-view setups (assuming static geometry and controlled environments), egocentric multi-view systems leverage body-worn rigs to enable a more robust and flexible 3D understanding in open-world, mobile scenarios. In this work, we present a systematic survey of advancements in cross-view feature fusion, geometric consistency enforcement, open-world detection, human–object interaction (HOI) modeling, action segmentation, 3D reconstruction, and novel-view synthesis specifically tailored to wearable multi-camera platforms. Key datasets released between 2024 and 2026—including HOT3D (833 min of synchronized multi-view hand/object interactions from Project Aria and Quest 3), MultiEgo (first multi-egocentric dataset for 4D social scene reconstruction), and Ego-1K (large-scale 12-camera rig for dynamic 3D video synthesis) are thoroughly examined alongside an analysis of integrations with large language models (LLMs) and vision–language models that drive performance gains, typically in the 15–30% range over single-view baselines in hand tracking, HOI recognition, and reconstruction fidelity, although we show through a consolidated meta-analysis that this gain is task-dependent: larger for geometry-bottlenecked tasks such as in-hand object lifting, and smaller, method-dependent, or occasionally negative for semantic-recognition tasks such as keystep recognition under naive view fusion. These methods cover work in multi-view stereo, cross-view learning, and novel-view synthesis while addressing several real-time wearable constraints. Practical applications such as immersive Augmented Reality/Virtual Reality (AR/VR), assistive robotics, and healthcare monitoring are also discussed together with the challenges in motion calibration, benchmark diversity, and edge deployment ability. Thus, in this review, we attempt to fill a critical gap by focusing exclusively on wearable multi-view systems in an open-world setting, synthesizing the latest literature to chart future directions toward more embodied and continual learning agents. Full article
(This article belongs to the Special Issue Techniques in Multi-View Image Analysis)
Show Figures

Figure 1

19 pages, 1153 KB  
Article
Hijacking the “Safety Bifurcation Point”: Coupled Path Hijacking Attacks on Multimodal Vision–Language Models
by Ranyi Peng and Jingfu Bao
Electronics 2026, 15(14), 3143; https://doi.org/10.3390/electronics15143143 - 16 Jul 2026
Viewed by 271
Abstract
The generalization of safety mechanisms from text-only LLMs to Large Vision–Language Models (LVLMs) remains poorly understood, especially concerning their vulnerability to typographic adversarial prompts. In this paper we ask where, inside a model, the safe-versus-unsafe decision is causally made. Focusing on decoder-only LVLMs, [...] Read more.
The generalization of safety mechanisms from text-only LLMs to Large Vision–Language Models (LVLMs) remains poorly understood, especially concerning their vulnerability to typographic adversarial prompts. In this paper we ask where, inside a model, the safe-versus-unsafe decision is causally made. Focusing on decoder-only LVLMs, we show that safety is encoded as a non-linear, distributed representation that bifurcates early at a distinct Safety Bifurcation Point (SBP)—e.g., Layer 9 in LLaVA-1.5 and Layer 5 in Qwen-VL—and we expose a Correlation–Causality Gap: late layers are highly linearly separable yet causally inert. To leverage this insight we introduce a multistage, causal (interpretability-guided) framework to locate and verify safety representations, and we propose CPH (Coupled Path Hijacking), a white-box attack that jointly patches MLP and self-attention activations at the SBP to deterministically hijack the safe computational path. On the two models for which we report full tabulated results, CPH attains high oracle-scored attack success rates of 99.43% (LLaVA-1.5) and 98.57% (Qwen-VL); we additionally observe 99.71% on InstructBLIP-Vicuna, whose Vicuna backend is likewise decoder-only. We contrast this decoder-only vulnerability with encoder–decoder designs (e.g., a T5 backend or a Q-Former front-end), in which cross-modal grounding precedes decoding; we advance the greater robustness of such architectures as an architectural hypothesis for future validation, not as a tabulated result of this study. These findings indicate that VLM safety is architecture-dependent and motivate architecture-aware defenses for trustworthy multimodal AI. Full article
Show Figures

Figure 1

Back to TopTop