Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (302)

Search Parameters:
Keywords = audio sensors

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
35 pages, 1900 KB  
Article
EELLM: An Emotion-Enhanced Large Language Model for Multimodal Emotion Perception in IoT-Enabled Smart Sensor Networks
by Lijiao Yang, Ming Cao, Ting Yang, Jinting Liu and Bachong Ma
Electronics 2026, 15(15), 3382; https://doi.org/10.3390/electronics15153382 - 1 Aug 2026
Viewed by 286
Abstract
AI-empowered smart sensor networks are driving Internet of Things (IoT) systems from passive data acquisition toward human-centric dynamic perception. In such scenarios, multimodal emotion recognition is expected to infer subtle affective states from heterogeneous audio, visual, and textual sensing streams. However, existing multimodal [...] Read more.
AI-empowered smart sensor networks are driving Internet of Things (IoT) systems from passive data acquisition toward human-centric dynamic perception. In such scenarios, multimodal emotion recognition is expected to infer subtle affective states from heterogeneous audio, visual, and textual sensing streams. However, existing multimodal emotion recognition and emotion-oriented multimodal large-language-model (MLLM) methods still face several limitations for fine-grained emotion perception. Temporal asynchrony weakens cross-modal correspondence, redundant or heterogeneous features blur emotion-discriminative cues, unreliable sensing streams reduce robustness, and directly injecting all multimodal tokens into a large language model increases decoding redundancy. To address these issues, this paper proposes an Emotion-Enhanced Large Language Model (EELLM), a unified multimodal collaborative interaction framework for emotion perception in IoT-enabled smart sensor networks. Specifically, EELLM employs a dynamic time warping (DTW)-based Cross-Modal Alignment Module (DCAM) to mitigate temporal inconsistency, a Gaussian maximum mean discrepancy (MMD)-based Multimodal Feature Interaction Module (GMFIM) to disentangle shared and private representations and suppress fusion redundancy, a Modality Reliability-Aware Gating module (MRG) to adaptively weight heterogeneous modalities, and an Emotion-Salient Token Compression strategy (ESTC) to retain emotion-discriminative prefix tokens before instruction-guided LLaMA decoding. Extensive experiments on the trimodal MER2023 and MER2024 benchmarks and the visual-only DFEW benchmark demonstrate the effectiveness of EELLM. Under the adopted evaluation settings, EELLM achieves an F1 score of 0.9068 on MER2023, an average score of 67.10 on MER2024, and a UAR of 68.41 on DFEW, presenting competitive emotion perception performance across trimodal and visual-only benchmark settings. In addition, EELLM improves the recognition accuracy of the underrepresented disgust category on DFEW to 24.03%, showing better class-balanced emotion perception. These results indicate that EELLM provides an effective and efficient solution for fine-grained multimodal emotion perception in resource-sensitive intelligent sensing scenarios. Full article
Show Figures

Figure 1

52 pages, 6054 KB  
Article
Intelligent Inclusive Navigation System for a University Digital Ecosystem
by Aibol Tileukhan, Gulmira Bekmanova, Valentina Franzoni, Alibek Barlybayev, Lena Zhetkenbay, Altynbek Sharipbay, Zhanar Lamasheva, Assel Omarbekova and Aizhan Nazyrova
Computers 2026, 15(8), 480; https://doi.org/10.3390/computers15080480 - 28 Jul 2026
Viewed by 283
Abstract
Indoor navigation remains challenging for students with visual impairments because GPS is unavailable indoors and building layouts are often complex. This paper presents a wearable marker-assisted navigation system integrating QR code localization, SSD MobileNet V3 obstacle detection, TFmini-S LiDAR ranging, A*-based dynamic route [...] Read more.
Indoor navigation remains challenging for students with visual impairments because GPS is unavailable indoors and building layouts are often complex. This paper presents a wearable marker-assisted navigation system integrating QR code localization, SSD MobileNet V3 obstacle detection, TFmini-S LiDAR ranging, A*-based dynamic route planning, and audio feedback on a Raspberry Pi 5. The main contribution is an analytical framework relating marker spacing to predicted localization uncertainty and defining a latency budget for obstacle warnings. A confidence-weighted sensor-fusion method is developed analytically but was not implemented in the evaluated prototype, in which the QR code, camera, and LiDAR channels operated independently. The proposed fusion method and the simulated multi-floor planning extension require further experimental validation. Controlled tests produced a mean positioning error below 1.2 m, a LiDAR ranging MAE of 8.3 cm, and an object-detection throughput of 6–9 FPS. A pilot field evaluation covered nine routes totalling 901 m across two buildings and included one participant with self-reported vision loss of approximately 95%. All route trials were completed, although some required researcher assistance. The system remains a proof of concept and has not yet been evaluated against a baseline or with a sufficiently large target-user sample. Full article
Show Figures

Figure 1

8 pages, 5029 KB  
Article
Single Applications of Commercial Mammal Deterrents Fail to Prevent Chewing Damage to Passive Acoustic Sensors
by Brooke D. Goodman, Lauren M. Chronister, Tessa A. Rhinehart, R. Patrick Lyon and Justin Kitzes
Sensors 2026, 26(15), 4704; https://doi.org/10.3390/s26154704 - 24 Jul 2026
Viewed by 308
Abstract
Large sensor arrays are an increasingly popular sampling method among ecologists. To last in the field, sensor housing needs to be resistant to damage from both weather and animals. The popular AudioMoth acoustic recorder does not have integral weather-resistant housing and is deployed [...] Read more.
Large sensor arrays are an increasingly popular sampling method among ecologists. To last in the field, sensor housing needs to be resistant to damage from both weather and animals. The popular AudioMoth acoustic recorder does not have integral weather-resistant housing and is deployed by users in a wide variety of protective cases. One inexpensive way to protect AudioMoths is to deploy them in plastic bags, which offer moderate weather resistance but are susceptible to chewing damage from small mammals. In this study, we test the effectiveness of commercially available mammal deterrents in preventing such chewing damage. We deployed 115 treatment-control pairs across two grids in temperate forests in Pennsylvania. Bag treatments consisted of Liquid Fence, Bonide, and a cayenne and Vaseline mixture. For all deterrents, there was no statistically significant difference in the proportion or severity of mammal chewing damage between treatments and controls. Counter to expectations, for all three treatments, more of the bags treated with a deterrent were damaged by mammal chewing than the paired control bags. Our results strongly suggest that single applications of these three deterrents have no useful effect on preventing mammal chewing damage to sensor housing in the field. Full article
(This article belongs to the Section Remote Sensors)
Show Figures

Figure 1

20 pages, 3202 KB  
Article
M2WPR-Net: Robust Multimodal Weld Quality Assessment via Cross-Modal Attention
by Ao Han, Tongyu Zhao, Yanjun Pei, Haining Chen, Jun Zhou, Hailei Yuan and Pan Hu
Information 2026, 17(7), 687; https://doi.org/10.3390/info17070687 - 15 Jul 2026
Viewed by 348
Abstract
Robust monitoring of weld pool dynamics is critical for automated arc welding; however, single-modality sensors are frequently constrained by severe optical interference and high-frequency environmental noise. To address these limitations, we propose M2WPR-Net, a novel multimodal framework that synergizes visual and acoustic signals [...] Read more.
Robust monitoring of weld pool dynamics is critical for automated arc welding; however, single-modality sensors are frequently constrained by severe optical interference and high-frequency environmental noise. To address these limitations, we propose M2WPR-Net, a novel multimodal framework that synergizes visual and acoustic signals for simultaneous weld width regression and physical quality classification. The architecture employs a dual-stream ResNet50 backbone to process heterogeneous sensory data. Specifically, the visual stream utilizes a Convolutional Block Attention Module (CBAM) to suppress intense arc glare and localize the weld pool. Concurrently, the acoustic stream transforms 1D audio sequences into 2D Gramian Angular Summation Field (GASF) textures, which are subsequently refined by Squeeze-and-Excitation (SE) networks to isolate target frequency channels. A central contribution of this study is a bidirectional cross-modal attention mechanism based on Query–Key–Value (Q-K-V) matrix operations. Overcoming the shortcomings of static feature concatenation, this module dynamically aligns the modalities, enabling acoustic cues to guide visual feature extraction and vice versa, thereby mitigating information bottlenecks. Optimized via a joint multi-task loss function, the proposed M2WPR-Net significantly outperforms existing single-modal and conventional fusion baselines. Experimental results demonstrate that the network achieves a Mean Absolute Error (MAE) of 0.18 mm for width prediction and a 93.5% accuracy in penetration state classification, confirming its resilience and practical applicability in complex industrial welding environments. Full article
(This article belongs to the Special Issue Advances in Computer Graphics and Visual Computing)
Show Figures

Figure 1

12 pages, 1594 KB  
Study Protocol
Detecting Distress in Cognitively Impaired People to Prevent Suffering: Protocol for an Observational Feasibility Study of a Radar-Based Technology Augmented with Photoplethysmographic Sensors and Audio Signals (SURREAL)
by Christopher Boehlke, Fabian Buergi, Jens Eckstein, Marc Stawiski, Simone Hemm, Wolfgang Hasemann and Jan Gaertner
Sensors 2026, 26(14), 4484; https://doi.org/10.3390/s26144484 - 15 Jul 2026
Viewed by 436
Abstract
Background: With the increasing prevalence of multimorbidity, the demand for palliative and end-of-life care is expected to rise substantially in the coming decades. Digital health technologies may enable automated detection of clinically relevant distress, including symptoms such as pain, breathlessness (dyspnea), anxiety/panic, [...] Read more.
Background: With the increasing prevalence of multimorbidity, the demand for palliative and end-of-life care is expected to rise substantially in the coming decades. Digital health technologies may enable automated detection of clinically relevant distress, including symptoms such as pain, breathlessness (dyspnea), anxiety/panic, nausea, and agitation. Remote detection of such distress in cognitively impaired patients who are unable to reliably call for help could enable timely intervention when patients are unattended. Methods: This observational feasibility study will collect multimodal data from a non-invasive sensor system consisting of 3D radar, a photoplethysmographic sensor (wearable), and a microphone. Sensor data will be linked to distress events identified by nurses or physicians during routine clinical care using structured proxy assessments. Adults (≥18 years) admitted to the Palliative Care Center Basel who are unable to reliably call for help due to cognitive impairment will be included based on written informed consent provided by a legal proxy. Aim: The aim of this study is to evaluate the feasibility of multimodal sensor-based monitoring for detecting clinician-identified distress events and to explore associations between sensor-derived variables and distress, informing future validation studies and the development of automated detection approaches in palliative care. Full article
(This article belongs to the Collection Medical Applications of Sensor Systems and Devices)
Show Figures

Figure 1

10 pages, 319 KB  
Article
Effect of Continuous Audio Biofeedback During Postoperative Partial Weight Bearing in Older Patients—An Exploratory Study
by Léa Staub, Arlene Vivienne von Aesch, Johannes Dominik Bastian and Heiner Baur
J. Clin. Med. 2026, 15(14), 5498; https://doi.org/10.3390/jcm15145498 - 14 Jul 2026
Viewed by 252
Abstract
Background/Objectives: Adherence to partial weight-bearing prescriptions (PWBP) is challenging. The use of audio biofeedback (AB) can potentially help to better implement PWBP. This study aimed to investigate the effect of continuous AB provided by sensor insoles on partial weight-bearing load during functional [...] Read more.
Background/Objectives: Adherence to partial weight-bearing prescriptions (PWBP) is challenging. The use of audio biofeedback (AB) can potentially help to better implement PWBP. This study aimed to investigate the effect of continuous AB provided by sensor insoles on partial weight-bearing load during functional tasks. Methods: Twenty older patients received a single AB training for PWBP management postoperatively. The prescribed limb load was measured (ground reaction force) during four activities (walking, walking with a 5 kg backpack, sitting–standing–sitting, and standing) with force–sensor insoles with continuous AB (n = 10) or without AB (n = 10). Individual deviation from the prescribed load and the influence of age and cognitive function were analyzed. Results: The intervention group (continuous AB) managed PWBP better for three out of four activities: The relative deviation was 104.4% ± 144.2 vs. 164.5% ± 164.5 for the 5 kg backpack walk, 57.6% ± 107.2 vs. 86.8% ± 122.4 for the sit–stand–sit task, and 8.0% ± 102.3 vs. 38.8% ± 131.7 for standing. For the 3 min walking activity, the relative deviation was 127.7% ± 121.9 vs. 116.0% ± 133.0 in favor of the control group. Mean differences were not statistically significant for any of the activities. Conclusions: After a single training session with continuous AB, PWBP can only be insufficiently met. Studies with multiple training sessions seem necessary to further test the potential of continuous AB in the management of PWBP. Full article
(This article belongs to the Special Issue The “Orthogeriatric Fracture Syndrome”—Issues and Perspectives)
Show Figures

Figure 1

22 pages, 12731 KB  
Article
MxArray: A Modular, Multiplexed, and Massive MEMS-Based Acoustic Array
by Ricardo Moreno, Jorge Ortigoso-Narro, Daniel de la Prida, Luis A. Azpicueta-Ruiz, Borja Genovés Guzmán and Marco Raiola
Sensors 2026, 26(12), 3899; https://doi.org/10.3390/s26123899 - 19 Jun 2026
Viewed by 1631
Abstract
While state-of-the-art massive acoustic arrays typically rely on costly, specialized FPGA architectures or rigid proprietary hardware, there is a growing need for modular, high-density sensing in complex aeroacoustics environments. This paper presents the electronic and acoustic design of a multiplexed, modular, scalable, and [...] Read more.
While state-of-the-art massive acoustic arrays typically rely on costly, specialized FPGA architectures or rigid proprietary hardware, there is a growing need for modular, high-density sensing in complex aeroacoustics environments. This paper presents the electronic and acoustic design of a multiplexed, modular, scalable, and low-cost massive acoustic array (MxArray) founded on an embedded Linux system. The AM3358 SoC microprocessor collects audio data through its multichannel audio peripheral, where it simultaneously receives four Time-Division Multiplexing streams of 16 microphones each. This multiplexed scheme enables the handling of 64 microphones per module, whose acquisition synchronization is set with the Precision Time Protocol and a pulse injection hardware. The combination of both BeagleBone Black and microphones based on Micro-Electro-Mechanical Systems yields a cost-effective solution with built-in Ethernet connectivity and accessible software development through an embedded Linux environment with audio libraries for hardware control. Sensors are arranged in an Underbrink Spiral pattern on a four-layer printed-circuit board. The perforated thin layout minimizes any airborne disturbance, exploiting a distribution that simultaneously achieves a low sidelobe level and a narrow main lobe when used with a beamforming algorithm. Measurement results for the developed module are presented, as well as an evaluation of a full-scale system comprising 16 modules (1024 microphones) arranged in a honeycomb pattern. The resulting instrument offers a practical and scalable solution for applications that require a large number of simultaneous microphone measurements, such as beamforming technology for aeroacoustics applications. Full article
(This article belongs to the Special Issue Acoustic Sensors and Their Applications—2nd Edition)
Show Figures

Figure 1

22 pages, 2231 KB  
Article
Simulation and Analysis of a Silicon Membrane-Supported Beam–Island Diaphragm for Graphene Piezoresistive MEMS Microphones in High-SPL Acoustic Sensing
by Shengsheng Wei, Chunyuan Li, Yipeng Wang, Junqiang Wang and Mengwei Li
Micromachines 2026, 17(6), 719; https://doi.org/10.3390/mi17060719 - 13 Jun 2026
Viewed by 445
Abstract
High sound pressure level (SPL) acoustic sensing requires miniaturized microphones that can operate under large acoustic loading while maintaining mechanical linearity, sufficient sensing response, and broadband audio frequency behavior. This work targets high-SPL operation and numerically investigates a graphene piezoresistive MEMS microphone based [...] Read more.
High sound pressure level (SPL) acoustic sensing requires miniaturized microphones that can operate under large acoustic loading while maintaining mechanical linearity, sufficient sensing response, and broadband audio frequency behavior. This work targets high-SPL operation and numerically investigates a graphene piezoresistive MEMS microphone based on a membrane-supported beam–island diaphragm. The proposed structure retains a continuous membrane for acoustic load bearing, while the upper beam–island topology redirects deformation-induced strain toward beam root regions where graphene piezoresistors are placed. This design is intended to increase the local strain available for piezoresistive readout without simply relying on larger global diaphragm deflection. Finite-element analysis was used to optimize the diaphragm geometry and evaluate strain enhancement, pressure response linearity, modal behavior, and harmonic response. Under the 170 dB SPL reference condition, the optimized structure increases the peak structural strain from 47.83 με in a thickness-equivalent solid diaphragm to 562.53 με, achieving an approximately 11.8-fold enhancement in local sensing strain while maintaining a highly linear pressure response (R2 > 0.9999). Additionally, the results also show that the sensor exhibits a high first natural frequency of 64.07 kHz and a small response variation of approximately 0.94 dB within the 0–20 kHz target frequency range, indicating excellent dynamic stability and high-fidelity signal transduction characteristics. To connect the structural response with piezoresistive readout, first-order electromechanical output estimation was further performed using representative graphene gauge factors, quarter-bridge readout assumptions, contact resistance correction, and Johnson-noise-limited signal-to-noise ratio estimation. A ±5% geometric tolerance check further indicates that the membrane side length is the most fabrication-sensitive parameter, while the selected design remains generally robust except for reduced linearity margin under positive membrane side-length deviation. These results demonstrate the potential of the proposed graphene-based MEMS microphone for high-SPL broadband acoustic sensing applications in harsh and high-intensity acoustic environments. Full article
Show Figures

Figure 1

31 pages, 30018 KB  
Article
Sensors-Driven Multimodal Deepfake Detection: A Cross-Attention Fusion Approach with Adaptive Modality Gating
by Syeda Sitara Waseem, Noman Shabbir, Syed Rizwan Hassan and KangYoon Lee
Sensors 2026, 26(12), 3695; https://doi.org/10.3390/s26123695 - 10 Jun 2026
Cited by 1 | Viewed by 626
Abstract
Deepfakes threaten sensor-based authentication systems, including biometric sensors, surveillance cameras, and IoT edge devices. Unimodal detectors remain vulnerable to modality-specific attacks. We propose a multimodal deepfake detection framework optimized for resource-constrained edge devices, featuring a novel cross-modal attention fusion mechanism with adaptive gating. [...] Read more.
Deepfakes threaten sensor-based authentication systems, including biometric sensors, surveillance cameras, and IoT edge devices. Unimodal detectors remain vulnerable to modality-specific attacks. We propose a multimodal deepfake detection framework optimized for resource-constrained edge devices, featuring a novel cross-modal attention fusion mechanism with adaptive gating. The architecture combines enhanced Res2Net for audio, temporal 3D CNN with SE attention for video, and bidirectional cross-modal attention with quality-based gates. On our benchmark (5472 audio + 1842 video samples), the fusion model achieves 96.7% accuracy, 96.6% F1-score, 0.988 AUC-ROC, and 3.3% EER. Adversarial testing shows 92.3% accuracy under the Fast Gradient Sign Method (FGSM) attack. The model has a 30.3 MB footprint and runs at 20 FPS on edge hardware. Modality contribution analysis reveals adaptive weighting (72% audio for TTS forgery, 78% video for lip-synced attacks). Cross-dataset evaluation on FakeAVCeleb achieves 92.3% overall accuracy, confirming generalization. Full article
Show Figures

Figure 1

23 pages, 3023 KB  
Article
Design of an Adaptive Augmented Reality Guidance System for Mechanical Assembly
by Aleeha Zafar and Magesh Chandramouli
Electronics 2026, 15(11), 2478; https://doi.org/10.3390/electronics15112478 - 4 Jun 2026
Viewed by 493
Abstract
This paper presents the design and development of an adaptive augmented reality (AR) assistance system for complex mechanical assembly tasks. Integrating a wrist-worn optical heart rate sensor to evaluate the user’s cognitive state, the system is intended to run as a standalone application [...] Read more.
This paper presents the design and development of an adaptive augmented reality (AR) assistance system for complex mechanical assembly tasks. Integrating a wrist-worn optical heart rate sensor to evaluate the user’s cognitive state, the system is intended to run as a standalone application on the Meta Quest 3 headset. The system displays instructions and visual cues directly overlaid on the user’s physical workspace and constantly monitors their heart rate variability through the sensor as an estimate of their cognitive load. When the system detects an overload, it dynamically adjusts the presentation of information—for example, it slows down pacing, simplifies instructions, or switches to a different interaction modality (audio)—as an attempt to reduce the overload. The paper makes three contributions: first, it provides a documented standalone integration of physiological sensing with adaptive interface logic on a mixed reality headset without external compute infrastructure; second, it provides a systematic characterization of platform-specific tracking incompatibilities on the Meta Quest 3, documenting the progression through four spatial registration strategies and the specific failure condition that triggered each transition; third, it reports spatial interface design observations from iterative developer testing in the current prototype configuration, including panel height ranges not previously reported in the AR interface literature at this level of specificity. The paper also discusses the within-subjects evaluation protocol that is planned for final system testing with actual users. The work is intended as an engineering and design contribution that establishes the foundation for subsequent empirical evaluation of adaptive AR guidance in industrial assembly contexts. Full article
Show Figures

Figure 1

13 pages, 15333 KB  
Communication
Noise Optimization of VCO-ADCs Based on Ring Oscillators with Cascoded Inverter Delay Cells
by Javier Granizo, Ruben Garvi, Javier Fernandez, Jorge de la Torre and Luis Hernandez
Electronics 2026, 15(11), 2299; https://doi.org/10.3390/electronics15112299 - 26 May 2026
Viewed by 1009
Abstract
A key component of VCO-ADCs is the ring oscillator, which determines the circuit and quantization noise of the converter. The input-referred thermal and flicker noise of a VCO-ADC stems from the VCO driver source and the VCO phase noise. On the other hand, [...] Read more.
A key component of VCO-ADCs is the ring oscillator, which determines the circuit and quantization noise of the converter. The input-referred thermal and flicker noise of a VCO-ADC stems from the VCO driver source and the VCO phase noise. On the other hand, quantization noise depends on the oscillation frequency of the VCO with respect to the sampling frequency. An optimal VCO-ADC design should balance oscillation frequency with flicker and thermal contributions of the VCO. In this paper, we show a simple modification of the conventional stages used in VCO-ADC ring oscillators. The modification consists of including two extra transistors in series, isolating the inverter from the power rails when switching. This modification allows one to significantly increase the oscillation frequency while having similar phase noise contributions compared to other ring oscillator architectures with the same area and power. Full article
(This article belongs to the Section Microelectronics)
Show Figures

Figure 1

56 pages, 31726 KB  
Review
Theoretical Framework, Technical Evolution, and Future Prospects of Cross-Modal Mapping and Controllable Image Generation Under Multi-Source Heterogeneous Collaboration
by Mingju Chen, Zhihao Lin, Xiaofei Song, Yangming Luo, Xueyang Duan, Senyuan Li and Chen Xie
Sensors 2026, 26(10), 2972; https://doi.org/10.3390/s26102972 - 8 May 2026
Viewed by 1045
Abstract
The rapid evolution of diffusion models has shifted visual synthesis from text-only inputs to precisely controlled generation driven by multi-source heterogeneous sensor signals (e.g., audio, 3D, and physiological data). This paper presents a systematic review of cross-modal mapping and controllable generation under multi-source [...] Read more.
The rapid evolution of diffusion models has shifted visual synthesis from text-only inputs to precisely controlled generation driven by multi-source heterogeneous sensor signals (e.g., audio, 3D, and physiological data). This paper presents a systematic review of cross-modal mapping and controllable generation under multi-source collaboration. More precisely, we propose a unified “cross-modal mapping and injection” taxonomy by abstracting the intervention logic of heterogeneous signals. Fundamentally, we analyze these mechanisms in a backbone-agnostic manner, delineating the architectural transition from legacy U-Net dependencies to scalable architectures like Diffusion Transformers (DiTs) and tracing the technical evolution from single-source atomic driving to complex multi-source collaborative paradigms. Our mechanistic analysis reveals that seamless feature fusion heavily relies on gradient conflict resolution, rigorous arbitration, and dynamic disentanglement under multi-constraint scenarios. Furthermore, by systematizing current evaluation metrics, we identify intrinsic quality-controllability trade-offs through performance game analysis (e.g., Pareto optimization), yielding a scientifically grounded technical selection guide. The study concludes that overcoming current generation limitations necessitates integrating Hardware-in-the-Loop (HIL) deployment, PDE-driven physical constraints, and causal inference, laying the foundation for next-generation robust and real-time generative models. Full article
(This article belongs to the Section State-of-the-Art Sensors Technologies)
Show Figures

Figure 1

29 pages, 4742 KB  
Article
DistSense: A Distributed P2P System for Privacy-Preserving and Robust Audiovisual Activity Recognition in Smart Homes
by José Manuel Torres, Luis P. Mota, Rui S. Moreira, Christophe Soares and Pedro Sobral
Appl. Sci. 2026, 16(9), 4407; https://doi.org/10.3390/app16094407 - 30 Apr 2026
Cited by 1 | Viewed by 769
Abstract
Ambient Assisted Living (AAL) systems have become increasingly relevant as aging populations intensify the demand for technologies that promote autonomy, safety, and quality of life. However, the widespread adoption of audiovisual sensing in smart homes raises critical concerns regarding data protection, privacy, and [...] Read more.
Ambient Assisted Living (AAL) systems have become increasingly relevant as aging populations intensify the demand for technologies that promote autonomy, safety, and quality of life. However, the widespread adoption of audiovisual sensing in smart homes raises critical concerns regarding data protection, privacy, and user trust. Ensuring secure processing while maintaining accurate activity recognition remains a key challenge. This work introduces DistSense, a distributed Peer-to-Peer (P2P) system designed to enhance activity detection in domestic environments through collaborative inference among intelligent audiovisual sensors. DistSense prioritizes privacy by performing local processing, sharing only high-level events, and leveraging distributed ledger mechanisms to ensure data integrity and auditability and support cross-device validation. This collaborative strategy reduces false positives caused by occlusions, illumination variability, and acoustic noise. To assess the system, functional tests were conducted for each module, followed by two use cases evaluated in both simulated and real edge hardware environments. The trained models achieved 88% accuracy for audio and 80% for video, and the system demonstrated effective performance in detecting daily activities and domestic hazards under varying noise conditions. Results indicate that DistSense successfully balances security, user acceptance, and inference robustness, positioning it as a viable solution for privacy-preserving activity monitoring in smart home contexts. Full article
Show Figures

Figure 1

544 KB  
Proceeding Paper
Design and Implementation of Facial Recognition Smart Glasses for Visually Challenged Persons
by Alfonzo Janrick Eneria and Ramon Garcia
Eng. Proc. 2026, 134(1), 99; https://doi.org/10.3390/engproc2026134099 - 21 Apr 2026
Viewed by 598
Abstract
We designed and implemented facial recognition smart glasses to assist visually impaired individuals in recognizing people and navigating their environment safely and independently. The smart glasses utilize a Raspberry Pi 4 as the processing unit, integrating a Pi Camera for facial recognition and [...] Read more.
We designed and implemented facial recognition smart glasses to assist visually impaired individuals in recognizing people and navigating their environment safely and independently. The smart glasses utilize a Raspberry Pi 4 as the processing unit, integrating a Pi Camera for facial recognition and a USB camera for object detection. A face recognition library is employed to extract 128-dimensional facial embeddings using a convolutional neural network, enabling real-time face identification at close range (100–500 cm) under proper lighting conditions. Object detection is performed using a YOLOv5-based model, while ultrasonic sensors provide proximity alerts through audio feedback. Real-time processing is optimized to minimize latency and protect user privacy. The smart glasses were tested on participants with varying levels of visual impairment, including low vision, legal blindness, and total blindness. The system achieved an overall facial recognition accuracy of 88.89% and an object detection accuracy of 61.11%. The results demonstrate the viability of edge-AI wearable devices in assistive technology, with user feedback highlighting strengths in audio feedback and recognition accuracy, as well as areas for improvement, such as device comfort and low-light performance. Full article
Show Figures

Figure 1

29 pages, 3416 KB  
Article
Enhancing Collaborative AI Learning: A Blockchain-Secured, Edge-Enabled Platform for Multimodal Education in IIoT Environments
by Ahsan Rafiq, Eduard Melnik, Alexey Samoylov, Alexander Kozlovskiy and Irina Safronenkova
Big Data Cogn. Comput. 2026, 10(4), 123; https://doi.org/10.3390/bdcc10040123 - 17 Apr 2026
Viewed by 1538
Abstract
As industries deploy more connected devices in factories, warehouses, and smart facilities, the need for artificial intelligence (AI) systems that can operate securely in distributed, data-intensive environments is growing. Traditional centralized learning and online education platforms struggle when students and systems have to [...] Read more.
As industries deploy more connected devices in factories, warehouses, and smart facilities, the need for artificial intelligence (AI) systems that can operate securely in distributed, data-intensive environments is growing. Traditional centralized learning and online education platforms struggle when students and systems have to process real-time streams (sensors, video, text) with strict latency and privacy requirements. To address this challenge, a blockchain-secured, edge-enabled multimodal federated learning framework tailored for Industrial IoT (IIoT) environments is proposed. The model integrates four key layers: (i) a blockchain layer that provides credentialing, transparency, and token-based incentives; (ii) a multimodal community layer that supports group formation, peer consensus, and cross-modal learning across text, images, audio, and sensor data; (iii) an edge computing layer that enables low-latency task offloading and secure training within Intel SGX enclaves; and (iv) a data layer that applies pre-processing, differential privacy, and synthetic augmentation to safeguard sensitive information. Experiments on industrial multimodal datasets demonstrate 42% faster model aggregation, 78.9% multimodal accuracy, and 1.9% accuracy loss under ε = 1.0 differential privacy. This shows a scalable and practical path for decentralized AI training in next-generation IIoT systems, confirming the possibility of technical support for educational processes. However, the conducted research requires a validation of pedagogical effectiveness. Full article
Show Figures

Figure 1

Back to TopTop