Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (135)

Search Parameters:
Keywords = network audio technology

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
21 pages, 14302 KB  
Article
Audio-Based Device for Automated Surgical Counting, ToolSafe
by Michael R. Gardner, Latifa A. Aladdal, Lama Alshammari, Fatima Aldalgan, Maram A. Alomair, Shahad Alomair and Amani Alrashed
Appl. Sci. 2026, 16(11), 5181; https://doi.org/10.3390/app16115181 - 22 May 2026
Viewed by 483
Abstract
Manual counting of surgical tools, known as surgical counting, is a time-consuming and error-prone task that increases the risk of retained surgical instruments and extends operating room (OR) time. Presently, in hospitals around the world, surgical counting is often performed manually with paper [...] Read more.
Manual counting of surgical tools, known as surgical counting, is a time-consuming and error-prone task that increases the risk of retained surgical instruments and extends operating room (OR) time. Presently, in hospitals around the world, surgical counting is often performed manually with paper or tablet checklists, often leading to delays, increased infection risk, and financial cost. RFID, barcode-based, and computer vision solutions exist but are expensive and have challenges with sterilization and signal interference. This paper presents ToolSafe, a low-cost, portable system that classifies surgical tools by their acoustic signatures when dropped into a detection box. A pilot dataset of 4004 audio samples from four tool types (n = 996, tissue forceps; n = 1005, iris scissors; n = 1006, scalpel handle; n = 997, testing needle) was collected using ToolSafe. A convolutional neural network (CNN) was evaluated using stratified five-fold cross-validation on the laboratory dataset, with a k-nearest neighbors (KNN) classifier implemented as a control model. In each fold, both models were trained on 80% of the data and tested on the remaining 20%, ensuring that all samples were used for both training and evaluation. The CNN achieved a mean (±standard deviation) classification accuracy of 99.55% (±0.19%) across the validation folds, outperforming the KNN model, which achieved a mean accuracy of 97.28% (±0.50%). The difference was statistically significant according to a paired t-test across folds (p = 0.0003), indicating CNN’s superior performance on the dataset. For a run of 100 additional samples using the Raspberry Pi-based system, spectrogram generation averaged 0.121 s (±0.025 s), CNN inference averaged 0.180 s (±0.033 s), and total end-to-end latency averaged 1.851 s (±0.253 s) per tool. This pilot study proposes a possible technological solution for surgical counting that reduces human error and enhances patient safety. ToolSafe may be subsequently improved by increasing the number of surgical tools used in the training dataset, testing under more robust OR-like environments, and comparing to other classification algorithms. Further refinement and incorporation of ToolSafe in operating room workflows have the potential to reduce patient risks from extended surgical times and retained surgical instruments. Full article
Show Figures

Figure 1

12 pages, 1542 KB  
Article
A Pilot Study of Telerobotic Radical Thyroidectomy for Thyroid Cancer Using a 5G Network
by Bing Wang, Chen Li, Zheng Wan, Jian Zhu, Meng Wang, Yanbing Jian, Zelong Yang, Xin Miao, Linlin Zhang, Fei Kuang, Lin Liu, Guolou Li, Qingqing He, Jing Yao and Wen Tian
J. Clin. Med. 2026, 15(10), 3591; https://doi.org/10.3390/jcm15103591 - 8 May 2026
Viewed by 654
Abstract
Background: The incidence of thyroid cancer has increased globally. In recent years, robotic surgical systems have been applied in thyroid surgery, and the rapid development of fifth-generation (5G) communication technology has laid a solid foundation for the smooth implementation of remote surgery. [...] Read more.
Background: The incidence of thyroid cancer has increased globally. In recent years, robotic surgical systems have been applied in thyroid surgery, and the rapid development of fifth-generation (5G) communication technology has laid a solid foundation for the smooth implementation of remote surgery. Objective: The aim was to explore the feasibility and safety of telerobotic radical thyroidectomy using 5G communication technology to treat thyroid cancer. Methods: From August 2024 to October 2024, telerobotic radical thyroidectomy was performed on seven female patients using a 5G wireless network and a dedicated line network (or ordinary wired broadband) spanning 22–2200 km. The patients’ clinical and information transmission data were analyzed. Results: All patients (papillary thyroid carcinoma, female, with an average age of 44.0 ± 4.6 years) underwent uneventful surgical procedures without any transfer to open surgery or complications. The average surgical duration was 91.3 ± 11.8 min, the average blood loss was 11.4 ± 4.8 mL, and the average postoperative hospital stay was 3.6 ± 0.8 days. All subjects were successfully discharged within 5 days after surgery. The average total latency time of the intraoperative network was 137.5 (range, 121–159) ms, and there were no adverse events, such as network disconnection, frame loss, or network attacks. The operator worked smoothly without any obvious delay or lag, and the recorded audio and video are clear. Conclusions: Telerobotic radical thyroidectomy for thyroid cancer over a 5G network demonstrates promising feasibility and safety. With stable network transmission and a clear surgical field, the precise operations required in thyroid surgery can be performed reliably. These findings suggest that this technology can facilitate high-quality surgical care in remote areas, contributing to a more balanced distribution of medical resources. Full article
Show Figures

Figure 1

26 pages, 1976 KB  
Article
Assisted Navigation for Visually Impaired People Using 3D Audio and Stereoscopic Cameras
by José Francisco Lucio-Naranjo, Daniel Sanaguano Moreno, Roberto A. Tenenbaum, Erick P. Herrera-Granda, Luis Bravo-Moncayo and Henry Paz-Arias
Appl. Sci. 2026, 16(9), 4405; https://doi.org/10.3390/app16094405 - 30 Apr 2026
Viewed by 615
Abstract
This paper presents a prototype for an assistive navigation system that integrates three-dimensional audio spatialization with computer vision to improve the mobility of visually impaired individuals. The system uses stereoscopic depth perception and real-time point cloud reconstruction alongside a modified YOLO convolutional neural [...] Read more.
This paper presents a prototype for an assistive navigation system that integrates three-dimensional audio spatialization with computer vision to improve the mobility of visually impaired individuals. The system uses stereoscopic depth perception and real-time point cloud reconstruction alongside a modified YOLO convolutional neural network for object detection and auralization techniques with head-related impulse response functions. Twenty participants (ten who were visually impaired and ten who were blindfolded) navigated controlled obstacle scenarios while wearing a chest-mounted camera and specialized headphones. The prototype achieved 95.00% precision in object classification across eleven obstacle categories and a 33.19% recall, indicating conservative detection behavior. The processing efficiency was 0.042489 s per image, which exceeds real-time requirements. User evaluation revealed an average collision rate of 0.5 per scenario and a mean completion time of 48 s. Statistical analysis showed no significant difference in collision rates between participant groups (p=0.172), though visually impaired participants demonstrated faster completion times (p=0.003). Integrating segmented, convolution-based audio processing with stereoscopic depth estimation enabled users to perceive obstacle locations through spatial sound cues, establishing a foundation for advancing assistive navigation technologies without extensive training. Full article
(This article belongs to the Section Acoustics and Vibrations)
Show Figures

Figure 1

544 KB  
Proceeding Paper
Design and Implementation of Facial Recognition Smart Glasses for Visually Challenged Persons
by Alfonzo Janrick Eneria and Ramon Garcia
Eng. Proc. 2026, 134(1), 99; https://doi.org/10.3390/engproc2026134099 - 21 Apr 2026
Viewed by 627
Abstract
We designed and implemented facial recognition smart glasses to assist visually impaired individuals in recognizing people and navigating their environment safely and independently. The smart glasses utilize a Raspberry Pi 4 as the processing unit, integrating a Pi Camera for facial recognition and [...] Read more.
We designed and implemented facial recognition smart glasses to assist visually impaired individuals in recognizing people and navigating their environment safely and independently. The smart glasses utilize a Raspberry Pi 4 as the processing unit, integrating a Pi Camera for facial recognition and a USB camera for object detection. A face recognition library is employed to extract 128-dimensional facial embeddings using a convolutional neural network, enabling real-time face identification at close range (100–500 cm) under proper lighting conditions. Object detection is performed using a YOLOv5-based model, while ultrasonic sensors provide proximity alerts through audio feedback. Real-time processing is optimized to minimize latency and protect user privacy. The smart glasses were tested on participants with varying levels of visual impairment, including low vision, legal blindness, and total blindness. The system achieved an overall facial recognition accuracy of 88.89% and an object detection accuracy of 61.11%. The results demonstrate the viability of edge-AI wearable devices in assistive technology, with user feedback highlighting strengths in audio feedback and recognition accuracy, as well as areas for improvement, such as device comfort and low-light performance. Full article
Show Figures

Figure 1

22 pages, 1747 KB  
Review
Talking Head Generation Through Generative Models and Cross-Modal Synthesis Techniques
by Hira Nisar, Salman Masood, Zaki Malik and Adnan Abid
J. Imaging 2026, 12(3), 119; https://doi.org/10.3390/jimaging12030119 - 10 Mar 2026
Viewed by 1755
Abstract
Talking Head Generation (THG) is a rapidly advancing field at the intersection of computer vision, deep learning, and speech synthesis, enabling the creation of animated human-like heads that can produce speech and express emotions with high visual realism. The core objective of THG [...] Read more.
Talking Head Generation (THG) is a rapidly advancing field at the intersection of computer vision, deep learning, and speech synthesis, enabling the creation of animated human-like heads that can produce speech and express emotions with high visual realism. The core objective of THG systems is to synthesize coherent and natural audio–visual outputs by modeling the intricate relationship between speech signals, facial dynamics, and emotional cues. These systems find widespread applications in virtual assistants, interactive avatars, video dubbing for multilingual content, educational technologies, and immersive virtual and augmented reality environments. Moreover, the development of THG has significant implications for accessibility technologies, cultural preservation, and remote healthcare interfaces. This survey paper presents a comprehensive and systematic overview of the technological landscape of Talking Head Generation. We begin by outlining the foundational methodologies that underpin the synthesis process, including generative adversarial networks (GANs), motion-aware recurrent architectures, and attention-based models. A taxonomy is introduced to organize the diverse approaches based on the nature of input modalities and generation goals. We further examine the contributions of various domains such as computer vision, speech processing, and human–robot interaction, each of which plays a critical role in advancing the capabilities of THG systems. The paper also provides a detailed review of datasets used for training and evaluating THG models, highlighting their coverage, structure, and relevance. In parallel, we analyze widely adopted evaluation metrics, categorized by their focus on image quality, motion accuracy, synchronization, and semantic fidelity. Operating parameters such as latency, frame rate, resolution, and real-time capability are also discussed to assess deployment feasibility. Special emphasis is placed on the integration of generative artificial intelligence (GenAI), which has significantly enhanced the adaptability and realism of talking head systems through more powerful and generalizable learning frameworks. Full article
Show Figures

Figure 1

26 pages, 530 KB  
Review
Generative AI as a General-Purpose Technology: Foundations, Applications, and Labor Market Implications Through 2030
by Maikel Leon
Big Data Cogn. Comput. 2026, 10(3), 69; https://doi.org/10.3390/bdcc10030069 - 27 Feb 2026
Cited by 3 | Viewed by 5088
Abstract
Generative Artificial Intelligence (AI) has transitioned from a research milestone to a general-purpose technology with wide-ranging implications for organizations, labor markets, and information systems. Thanks to improvements in deep learning, generative adversarial networks (GANs), variational autoencoders (VAEs), diffusion models, transformer-based language models, and [...] Read more.
Generative Artificial Intelligence (AI) has transitioned from a research milestone to a general-purpose technology with wide-ranging implications for organizations, labor markets, and information systems. Thanks to improvements in deep learning, generative adversarial networks (GANs), variational autoencoders (VAEs), diffusion models, transformer-based language models, and reinforcement learning from human feedback (RLHF), generative AI can now create high-quality text, images, audio, code, and other types of content. This review synthesizes the core technical foundations and best practices for training, evaluation, and governance, with an emphasis on scalability and human oversight. The paper examines applications across customer service, marketing, software development, healthcare, finance, law, logistics, and the creative industries, and assesses the labor implications of generative AI using a sociotechnical lens. This study also develops a disruption index that integrates task exposure, adoption rates, time savings, and skill complementarity. The paper concludes with actionable recommendations for policymakers, organizations, and workers, emphasizing the importance of reskilling, algorithmic transparency, and inclusive innovation. Taken together, these contributions situate generative AI within broader debates about automation, augmentation, and the future of work. Full article
(This article belongs to the Section Large Language Models and Embodied Intelligence)
21 pages, 920 KB  
Article
Audio Deepfake Detection via a Fuzzy Dual-Path Time-Frequency Attention Network
by Jinzi Li, Hexu Wang, Fei Xie, Xiaozhou Feng, Jiayao Chen, Jindong Liu and Juan Wang
Sensors 2025, 25(24), 7608; https://doi.org/10.3390/s25247608 - 15 Dec 2025
Cited by 3 | Viewed by 1627
Abstract
With the rapid advancement of speech synthesis and voice conversion technologies, audio deepfake techniques have posed serious threats to information security. Existing detection methods often lack robustness when confronted with environmental noise, signal compression, and ambiguous fake features, making it difficult to effectively [...] Read more.
With the rapid advancement of speech synthesis and voice conversion technologies, audio deepfake techniques have posed serious threats to information security. Existing detection methods often lack robustness when confronted with environmental noise, signal compression, and ambiguous fake features, making it difficult to effectively identify highly concealed fake audio. To address this issue, this paper proposes a Dual-Path Time-Frequency Attention Network (DPTFAN) based on Pythagorean Hesitant Fuzzy Sets (PHFS), which dynamically characterizes the reliability and ambiguity of fake features through uncertainty modeling. It introduces a dual-path attention mechanism in both time and frequency domains to enhance feature representation and discriminative capability. Additionally, a Lightweight Fuzzy Branch Network (LFBN) is designed to achieve explicit enhancement of ambiguous features, improving performance while maintaining computational efficiency. On the ASVspoof 2019 LA dataset, the proposed method achieves an accuracy of 98.94%, and on the FoR (Fake or Real) dataset, it reaches an accuracy of 99.40%, significantly outperforming existing mainstream methods and demonstrating excellent detection performance and robustness. Full article
(This article belongs to the Section Sensor Networks)
Show Figures

Figure 1

34 pages, 583 KB  
Review
Artificial Intelligence Applications in Chronic Obstructive Pulmonary Disease: A Global Scoping Review of Diagnostic, Symptom-Based, and Outcome Prediction Approaches
by Alberto Pinheira, Manuel Casal-Guisande, Cristina Represas-Represas, María Torres-Durán, Alberto Comesaña-Campos and Alberto Fernández-Villar
Biomedicines 2025, 13(12), 3053; https://doi.org/10.3390/biomedicines13123053 - 11 Dec 2025
Cited by 8 | Viewed by 3109
Abstract
Background: Chronic Obstructive Pulmonary Disease (COPD) represents a significant global health burden, characterized by complex diagnostic and management challenges. Artificial Intelligence (AI) presents a powerful opportunity to enhance clinical decision-making and improve patient outcomes by leveraging complex health data. Objectives: This [...] Read more.
Background: Chronic Obstructive Pulmonary Disease (COPD) represents a significant global health burden, characterized by complex diagnostic and management challenges. Artificial Intelligence (AI) presents a powerful opportunity to enhance clinical decision-making and improve patient outcomes by leveraging complex health data. Objectives: This scoping review aims to systematically map the existing literature on AI applications in COPD. The primary objective is to identify, categorize, and summarize research into three key domains: (1) Diagnosis, (2) Clinical Symptoms, and (3) Clinical Outcomes. Methods: A scoping review was conducted following the Arksey and O’Malley framework. A comprehensive search of major scientific databases, including PubMed, Scopus, IEEE Xplore, and Google Scholar, was performed. The Population–Concept–Context (PCC) criteria included patients with COPD (Population), the use of AI (Concept), and applications in healthcare settings (Context). A global search strategy was employed with no geographic restrictions. Studies were included if they were original research articles published in English. The extracted data were charted and classified into the three predefined categories. Results: A total of 120 studies representing global distribution were included. Most datasets originated from Asia (predominantly China and India) and Europe (notably Spain and the UK), followed by North America (USA and Canada). There was a notable scarcity of data from South America and Africa. The findings indicate a strong trend towards the use of deep learning (DL), particularly Convolutional Neural Networks (CNNs) for medical imaging, and tree-based machine learning (ML) models like CatBoost for clinical data. The most common data types were electronic health records, chest CT scans, and audio recordings. While diagnostic applications are well-established and report high accuracy, research into symptom analysis and phenotype identification is an emerging area. Key gaps were identified in the lack of prospective validation and clinical implementation studies. Conclusions: Current evidence shows that AI offers promising applications for COPD diagnosis, outcome prediction, and symptom analysis, but most reported models remain at an early stage of maturity due to methodological limitations and limited external validation. Future research should prioritize rigorous clinical evaluation, the development of explainable and trustworthy AI systems, and the creation of standardized, multi-modal datasets to support reliable and safe translation of these technologies into routine practice. Full article
(This article belongs to the Section Molecular and Translational Medicine)
Show Figures

Graphical abstract

20 pages, 2845 KB  
Article
From Gaze to Music: AI-Powered Personalized Audiovisual Experiences for Children’s Aesthetic Education
by Jiahui Liu, Jing Liu and Hong Yan
Behav. Sci. 2025, 15(12), 1684; https://doi.org/10.3390/bs15121684 - 4 Dec 2025
Cited by 3 | Viewed by 1088
Abstract
The cultivation of aesthetic appreciation through engagement with exemplary artworks constitutes a fundamental pillar in fostering children’s cognitive and emotional development, while simultaneously facilitating multidimensional learning experiences across diverse perceptual domains. However, children in early stages of cognitive development frequently encounter substantial challenges [...] Read more.
The cultivation of aesthetic appreciation through engagement with exemplary artworks constitutes a fundamental pillar in fostering children’s cognitive and emotional development, while simultaneously facilitating multidimensional learning experiences across diverse perceptual domains. However, children in early stages of cognitive development frequently encounter substantial challenges when attempting to comprehend and internalize complex visual narratives and abstract artistic concepts inherent in sophisticated artworks. This study presents an innovative methodological framework designed to enhance children’s artwork comprehension capabilities by systematically leveraging the theoretical foundations of audio-visual cross-modal integration. Through investigation of cross-modal correspondences between visual and auditory perceptual systems, we developed a sophisticated methodology that extracts and interprets musical elements based on gaze behavior patterns derived from prior pilot studies when observing artworks. Utilizing state-of-the-art deep learning techniques, specifically Recurrent Neural Networks (RNNs), these extracted visual–musical correspondences are subsequently transformed into cohesive, aesthetically pleasing musical compositions that maintain semantic and emotional congruence with the observed visual content. The efficacy and practical applicability of our proposed method were validated through empirical evaluation involving 96 children (analyzed through objective behavioral assessments using eye-tracking technology), complemented by qualitative evaluations from 16 parents and 5 experienced preschool educators. Our findings show statistically significant improvements in children’s sustained engagement and attentional focus under AI-generated, artwork-matched audiovisual support, potentially scaffolding deeper processing and informing future developments in aesthetic education. The results demonstrate statistically significant improvements in children’s sustained engagement (fixation duration: 58.82 ± 7.38 s vs. 41.29 ± 6.92 s, p < 0.001, Cohen’s d ≈ 1.29), attentional focus (AOI gaze frequency increased 73%, p < 0.001), and subjective evaluations from parents (mean ratings 4.56–4.81/5) when visual experiences are augmented by AI-generated, personalized audio-visual experiences. Full article
(This article belongs to the Section Cognition)
Show Figures

Figure 1

23 pages, 2079 KB  
Article
Enhanced Image Security via Dynamic Chaotic Fuzzy Cellular Neural Networks and Voice Authentication
by Maha Ayad Alenizi, Kalpana Muthusamy and Seng Huat Ong
Symmetry 2025, 17(12), 2056; https://doi.org/10.3390/sym17122056 - 2 Dec 2025
Cited by 1 | Viewed by 639
Abstract
The vulnerability of transmitted digital images has become a pressing concern due to recent advancements in multimedia technology. Conventional encryption methods often fail to meet the requirements of large-scale real-time multimedia security. In order to strengthen color image encryption, in this paper, we [...] Read more.
The vulnerability of transmitted digital images has become a pressing concern due to recent advancements in multimedia technology. Conventional encryption methods often fail to meet the requirements of large-scale real-time multimedia security. In order to strengthen color image encryption, in this paper, we propose a novel encryption method that combines fuzzy cellular neural networks with dynamic audio-based biometric data, which aligns with the principle of symmetric encryption. To make the encryption process specific to each user and hard to replicate, the method uses speech characteristics—the peak frequency and zero-crossing rate—extracted from the user’s voice. By integrating these voice features into the fuzzy cellular neural network structure, the method expands the set of potential keys and enhances protection against brute-force, statistical, and chosen-plaintext attacks. Compared to conventional methods that rely solely on chaotic maps or neural networks, this approach provides a larger key space, higher entropy, and better disruption of pixel correlation. The encryption quality is validated through experimental results using the NPCR, UACI, PSNR, and SSIM metrics. Securing multimedia transmissions contributes to the broader vision of a protected society, where technological progress promotes safety, trust, and fair access to knowledge. Full article
Show Figures

Figure 1

19 pages, 4574 KB  
Article
Multi-Service Multiplexing System Based on Visible Light Communication
by Yangyu Zhang
Sensors 2025, 25(23), 7207; https://doi.org/10.3390/s25237207 - 26 Nov 2025
Cited by 1 | Viewed by 928
Abstract
As the Internet of Things (IoT) and communication technologies continue to evolve, the value of multi-service multiplexing in visible light communication (VLC) systems has been increasingly recognized, particularly in addressing the scarcity of wireless spectrum resources. This study reconstructed the stereo transmission protocol [...] Read more.
As the Internet of Things (IoT) and communication technologies continue to evolve, the value of multi-service multiplexing in visible light communication (VLC) systems has been increasingly recognized, particularly in addressing the scarcity of wireless spectrum resources. This study reconstructed the stereo transmission protocol through methods such as dynamic level control, designed a timer interrupt service routine with a double buffer, and reassigned channel status bits in the frame processing function. Consequently, a multi-service multiplexing system based on VLC was designed and implemented. The system enables hybrid transmission of audio signals (1–21.6 kHz) and character data (300–1200 bps) via a single channel, accurately reproducing both voice and text input over a 3.2 m communication range. The system, benefiting from the directional nature of visible light communication, exhibits inherent robustness to multipath-induced interference in dominant line-of-sight (LoS) scenarios and can be easily integrated into existing lighting networks. Featuring a simple architecture and cost-effective design, this solution shows promise for deployment in RF-sensitive areas requiring multi-service communication. Full article
(This article belongs to the Collection Visible Light Communication (VLC))
Show Figures

Figure 1

37 pages, 16007 KB  
Review
Speech Separation Using Advanced Deep Neural Network Methods: A Recent Survey
by Zeng Wang and Zhongqiang Luo
Big Data Cogn. Comput. 2025, 9(11), 289; https://doi.org/10.3390/bdcc9110289 - 14 Nov 2025
Cited by 3 | Viewed by 6460
Abstract
Speech separation, as an important research direction in audio signal processing, has been widely studied by the academic community since its emergence in the mid-1990s. In recent years, with the rapid development of deep neural network technology, speech processing based on deep neural [...] Read more.
Speech separation, as an important research direction in audio signal processing, has been widely studied by the academic community since its emergence in the mid-1990s. In recent years, with the rapid development of deep neural network technology, speech processing based on deep neural networks has shown outstanding performance in speech separation. While existing studies have surveyed the application of deep neural networks in speech separation from multiple dimensions including learning paradigms, model architectures, loss functions, and training strategies, current achievements still lack systematic comprehension of the field’s developmental trajectory. To address this, this paper focuses on single-channel supervised speech separation tasks, proposing a technological evolution path “U-Net–TasNet–Transformer–Mamba” as the main thread to systematically analyze the impact mechanisms of core architectural designs on separation performance across different stages. By reviewing the transition process from traditional methods to deep learning paradigms and delving into the improvements and integration of deep learning architectures at various stages, this paper summarizes milestone achievements, mainstream evaluation frameworks, and typical datasets in the field, while also providing prospects for future research directions. Through this detailed-focused review perspective, we aim to provide researchers in the speech separation field with a clearly articulated technical evolution map and practical reference. Full article
Show Figures

Figure 1

23 pages, 2166 KB  
Article
Performance Analysis of Switch Buffer Management Policy for Mixed-Critical Traffic in Time-Sensitive Networks
by Ling Zheng, Yingge Feng, Weiqiang Wang and Qianxi Men
Mathematics 2025, 13(21), 3443; https://doi.org/10.3390/math13213443 - 29 Oct 2025
Cited by 1 | Viewed by 1625
Abstract
Time-sensitive networking (TSN), a cutting-edge technology enabling efficient real-time communication and control, provides strong support for traditional Ethernet in terms of real-time performance, reliability, and deterministic transmission. In TSN systems, although time-triggered (TT) flows enjoy deterministic delay guarantees, audio video bridging (AVB) and [...] Read more.
Time-sensitive networking (TSN), a cutting-edge technology enabling efficient real-time communication and control, provides strong support for traditional Ethernet in terms of real-time performance, reliability, and deterministic transmission. In TSN systems, although time-triggered (TT) flows enjoy deterministic delay guarantees, audio video bridging (AVB) and best effort (BE) traffic still share link bandwidth through statistical multiplexing, a process that remains nondeterministic. This competition in shared memory switches adversely affects data transmission performance. In this paper, a priority queue threshold control policy is proposed and analyzed for mixed-critical traffic in time-sensitive networks. The core of this policy is to set independent queues for different types of traffic in the shared memory queuing system. To prevent low-priority traffic from monopolizing the shared buffer, its entry into the queue is blocked when buffer usage exceeds a preset threshold. A two-dimensional Markov chain is introduced to accurately construct the system’s queuing model. Through detailed analysis of the queuing model, the truncated chain method is used to decompose the two-dimensional state space into solvable one-dimensional sub-problems, and the approximate solution of the system’s steady-state distribution is derived. Based on this, the blocking probability, average queue length, and average queuing delay of different priority queues are accurately calculated. Finally, according to the optimization goal of the overall blocking probability of the system, the optimal threshold value is determined to achieve better system performance. Numerical results show that this strategy can effectively allocate the shared buffer space in multi-priority traffic scenarios. Compared with the conventional schemes, the queue blocking probability is reduced by approximately 40% to 60%. Full article
Show Figures

Figure 1

37 pages, 10732 KB  
Review
Advances on Multimodal Remote Sensing Foundation Models for Earth Observation Downstream Tasks: A Survey
by Guoqing Zhou, Lihuang Qian and Paolo Gamba
Remote Sens. 2025, 17(21), 3532; https://doi.org/10.3390/rs17213532 - 24 Oct 2025
Cited by 16 | Viewed by 9122
Abstract
Remote sensing foundation models (RSFMs) have demonstrated excellent feature extraction and reasoning capabilities under the self-supervised learning paradigm of “unlabeled datasets—model pre-training—downstream tasks”. These models achieve superior accuracy and performance compared to existing models across numerous open benchmark datasets. However, when confronted with [...] Read more.
Remote sensing foundation models (RSFMs) have demonstrated excellent feature extraction and reasoning capabilities under the self-supervised learning paradigm of “unlabeled datasets—model pre-training—downstream tasks”. These models achieve superior accuracy and performance compared to existing models across numerous open benchmark datasets. However, when confronted with multimodal data, such as optical, LiDAR, SAR, text, video, and audio, the RSFMs exhibit limitations in cross-modal generalization and multi-task learning. Although several reviews have addressed the RSFMs, there is currently no comprehensive survey dedicated to vision–X (vision, language, audio, position) multimodal RSFMs (MM-RSFMs). To tackle this gap, this article provides a systematic review of MM-RSFMs from a novel perspective. Firstly, the key technologies underlying MM-RSFMs are reviewed and analyzed, and the available multimodal RS pre-training datasets are summarized. Then, recent advances in MM-RSFMs are classified according to the development of backbone networks and cross-modal interaction methods of vision–X, such as vision–vision, vision–language, vision–audio, vision–position, and vision–language–audio. Finally, potential challenges are analyzed, and perspectives for MM-RSFMs are outlined. This survey from this paper reveals that current MM-RSFMs face the following key challenges: (1) a scarcity of high-quality multimodal datasets, (2) limited capability for multimodal feature extraction, (3) weak cross-task generalization, (4) absence of unified evaluation criteria, and (5) insufficient security measures. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

21 pages, 3700 KB  
Article
Lung Sound Classification Model for On-Device AI
by Jinho Park, Chanhee Jeong, Yeonshik Choi, Hyuck-ki Hong and Youngchang Jo
Appl. Sci. 2025, 15(17), 9361; https://doi.org/10.3390/app15179361 - 26 Aug 2025
Cited by 3 | Viewed by 3342
Abstract
Following the COVID-19 pandemic, public interest in healthcare has significantly in-creased, emphasizing the importance of early disease detection through lung sound analysis. Lung sounds serve as a critical biomarker in the diagnosis of pulmonary diseases, and numerous deep learning-based approaches have been actively [...] Read more.
Following the COVID-19 pandemic, public interest in healthcare has significantly in-creased, emphasizing the importance of early disease detection through lung sound analysis. Lung sounds serve as a critical biomarker in the diagnosis of pulmonary diseases, and numerous deep learning-based approaches have been actively explored for this purpose. Existing lung sound classification models have demonstrated high accuracy, benefiting from recent advances in artificial intelligence (AI) technologies. However, these models often rely on transmitting data to computationally intensive servers for processing, introducing potential security risks due to the transfer of sensitive medical information over networks. To mitigate these concerns, on-device AI has garnered growing attention as a promising solution for protecting healthcare data. On-device AI enables local data processing and inference directly on the device, thereby enhancing data security compared to server-based schemes. Despite these advantages, on-device AI is inherently limited by computational constraints, while conventional models typically require substantial processing power to maintain high performance. In this study, we propose a lightweight lung sound classification model designed specifically for on-device environments. The proposed scheme extracts audio features using Mel spectrograms, chromagrams, and Mel-Frequency Cepstral Coefficients (MFCC), which are converted into image representations and stacked to form the model input. The lightweight model performs convolution operations tailored to both temporal and frequency–domain characteristics of lung sounds. Comparative experimental results demonstrate that the proposed model achieves superior inference performance while maintaining a significantly smaller model size than conventional classification schemes, making it well-suited for deployment on resource-constrained devices. Full article
Show Figures

Figure 1

Back to TopTop