Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (447)

Search Parameters:
Keywords = quantized neural networks

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
32 pages, 5759 KB  
Article
A Multimodal TinyML-Based Predictive Maintenance Architecture for Industrial IoT in the 6G Era
by Carlos Exequiel Garay, Fernando Alberto Miranda Bonomi, Gonzalo Nicolás Mansilla, Mariano Fagre, Sergio Gustavo Guzmán, Pablo Alberto Ritorto, Franco Ismael Perez and Marcos Katz
Sensors 2026, 26(14), 4536; https://doi.org/10.3390/s26144536 - 17 Jul 2026
Viewed by 264
Abstract
Predictive maintenance (PdM) is central to Industry 5.0 strategies for reducing unplanned downtime in rotating machinery. This work proposes and evaluates, as a proof of concept on a controlled single-machine testbed, a multimodal TinyML edge architecture for PdM designed to remain compatible across [...] Read more.
Predictive maintenance (PdM) is central to Industry 5.0 strategies for reducing unplanned downtime in rotating machinery. This work proposes and evaluates, as a proof of concept on a controlled single-machine testbed, a multimodal TinyML edge architecture for PdM designed to remain compatible across the application plane’s evolution toward sixth-generation (6G) networks. Three complementary modalities run local inference on commercial off-the-shelf smart sensor nodes—vibration, acoustic, and thermography—with an embedded gateway bridging per-modality decisions to a serverless cloud back-end. Using real vibration data from a controlled static-unbalance protocol, five anomaly-detection model variants, operating on ten frequency-independent time-domain features extracted from 6 s windows, are benchmarked on the actual Cortex-M4F target; the INT8-quantized fully connected autoencoder, scored by per-window reconstruction error, reaches F1 = 0.9807 with 254 µs inference latency and a 6056 B Flash footprint, well within the microcontroller budget. In a second acquisition session with the remounted sensor, the frozen model retains perfect fault recall, and a short per-installation healthy-baseline recalibration restores F1 = 0.975 without any weight retraining. The acoustic modality is classified in-sensor on log-Mel filterbank energies by the Syntiant NDP120 neural coprocessor, and the thermographic modality by a lightweight binary CNN on 96 × 96 px frames. A preliminary intra-session late-fusion analysis suggests that a logistic-regression meta-learner over the three modality confidence scores can improve on single-modality baselines when no single modality already saturates, motivating multimodal sensing primarily for robustness and redundancy. An end-to-end latency experiment shows that the cloud-uplink leg dominates the budget (79–88%), establishing edge-first inference as a necessary condition for 6G URLLC gains to be observable at the application level. All experiments are conducted over Wi-Fi and MQTT with no 5G or 6G radio, so 6G compatibility is presented as a forward-looking roadmap rather than a tested capability. Full article
Show Figures

Figure 1

34 pages, 11421 KB  
Article
Algebraically Stabilizing Blocks for Quantized and Finite-Precision Neural Networks
by Kostadin Yotov, Emil Hadzhikolev and Stanka Hadzhikoleva
Axioms 2026, 15(7), 533; https://doi.org/10.3390/axioms15070533 - 16 Jul 2026
Viewed by 109
Abstract
This paper proposes a construction of algebraically stabilizing blocks for quantized and finite-precision neural networks. The approach is based on linear transformations defined by integer-valued matrices satisfying a condition of the form Wk=I+μD, which specifies algebraically [...] Read more.
This paper proposes a construction of algebraically stabilizing blocks for quantized and finite-precision neural networks. The approach is based on linear transformations defined by integer-valued matrices satisfying a condition of the form Wk=I+μD, which specifies algebraically controlled k-step behavior and modular periodicity in the integer-valued setting. The proposed property is invariant under conjugation, allowing the stabilizing construction to be transferred across equivalent linear representations. The resulting module can be integrated locally into existing neural architectures without requiring the algebraic structure to be imposed globally. The theoretical results concern algebraic and modular stabilization of the linear block and do not constitute a general guarantee of classical spectral or asymptotic stability in real-valued space. The approach is evaluated experimentally under multi-component harmonic, impulsive, and noisy inputs in both floating-point and INT8-quantized settings. Across several experimental configurations, the proposed block reduces output energy, component-wise variation, selected amplitude-related measures, and finite-precision deviation relative to floating-point reference trajectories. Norm-matched control experiments further suggest that the observed effects are not attributable solely to a reduction in operator magnitude, but may also reflect structural properties of the algebraically constructed operator. The proposed construction is particularly relevant to neural systems operating under limited numerical precision, including FPGA-, ASIC-, and edge-oriented implementations. It provides a structural approach for incorporating formally specified algebraic properties into the design of neural-network architectures. Full article
(This article belongs to the Special Issue Advances in Linear Algebra with Applications, 2nd Edition)
Show Figures

Figure 1

18 pages, 2692 KB  
Article
High-Throughput and Low-Latency DrDoS Detection Using Quantized ONNX Models
by Salih Eren, Alperen Gültekin, Ömer Özkan and İlker Özçelik
Symmetry 2026, 18(7), 1187; https://doi.org/10.3390/sym18071187 - 14 Jul 2026
Viewed by 246
Abstract
Distributed Reflection Denial-of-Service (DrDoS) attacks, such as DNS Amplification, exploit inherent asymmetries in standard network protocols to generate devastating traffic volumes. To restore defense-side balance against these asymmetric threats, we present a comprehensive, systematic comparison of ONNX compilation and quantization configurations across Multi-Layer [...] Read more.
Distributed Reflection Denial-of-Service (DrDoS) attacks, such as DNS Amplification, exploit inherent asymmetries in standard network protocols to generate devastating traffic volumes. To restore defense-side balance against these asymmetric threats, we present a comprehensive, systematic comparison of ONNX compilation and quantization configurations across Multi-Layer Perceptron (MLP), Convolutional Neural Network (CNN), and Gated Recurrent Unit (GRU) architectures evaluated on a unified reference platform. To achieve this, our evaluation analyzes training durations, throughput scalability, and sample latency across varying batch sizes to align with operational network environments. We demonstrate that while small batch sizes suffer from data transfer overhead, increasing batch configurations significantly accelerates GPU throughput, particularly for statically quantized ONNX models. Additionally, training durations exhibit an inverse scaling relationship with batch size, yielding massive temporal savings as workloads expand. Furthermore, larger batch sizes effectively amortize fixed execution costs across thousands of samples, reducing average per-sample latency. Ultimately, this systematic evaluation provides a deployment blueprint for highly efficient intrusion detection engines capable of neutralizing asymmetric network threats at line-rate. Full article
(This article belongs to the Section Computer)
Show Figures

Figure 1

33 pages, 5623 KB  
Article
Spiking Neural Network Based on Hierarchical Residual Quantization and Temporal Error Compensation for Remote Sensing Object Detection
by Yaming Cao, Yukai Xing, Liqun Kuang and Shichao Jiao
Appl. Sci. 2026, 16(14), 6908; https://doi.org/10.3390/app16146908 - 9 Jul 2026
Viewed by 338
Abstract
Compared with traditional artificial neural networks (ANNs), spiking neural networks (SNNs) have lower computational complexity, lower energy consumption, and faster inference speed, making them more promising for practical deployment on edge devices. However, in SNNs, since the output spikes of neurons are discrete, [...] Read more.
Compared with traditional artificial neural networks (ANNs), spiking neural networks (SNNs) have lower computational complexity, lower energy consumption, and faster inference speed, making them more promising for practical deployment on edge devices. However, in SNNs, since the output spikes of neurons are discrete, the model may face information loss, especially when the membrane potential is quantized into binary spikes, where quantization errors can lead to model precision loss and information loss. To address these challenges, this study proposes a spiking neural network based on hierarchical residual quantization and temporal error compensation (HRQ-TEC-SNN) and used for remote sensing object detection tasks. Through hierarchical residual quantization and temporal error compensation design, higher resolution quantization of membrane potential can be performed and the quantization threshold of membrane potential can be dynamically adjusted to compensate for errors introduced during the quantization process, allowing for fine-tuning at each time step and reducing information loss caused by coarse quantization. In terms of network structure, by introducing depthwise separable convolution modules, channel attention and spatial attention mechanisms, and integrating fast spatial pyramid pooling based on pulse neural networks, the model’s detection accuracy has been further improved while reducing the number of model parameters and computational costs. Experimental results show that the HRQ-TEC-SNN achieves significant advantages in both accuracy and energy consumption on the DOTAv1.0, DOTAv1.5 and DIOR datasets. Full article
Show Figures

Figure 1

17 pages, 807 KB  
Article
Adaptive A-Semilogarithmic Gradient Quantization for Efficient Deep Neural Network Training
by Stefan Panić, Milan Dubljanin, Milan Savić and Marko Smilić
Algorithms 2026, 19(7), 521; https://doi.org/10.3390/a19070521 - 29 Jun 2026
Viewed by 256
Abstract
This paper introduces an adaptive A-semilogarithmic gradient quantization framework aimed at reducing memory overhead and computational complexity during the training of deep neural networks. The approach employs a semilogarithmic companding function parameterized by a dynamically adjusted scaling factor A, which evolves [...] Read more.
This paper introduces an adaptive A-semilogarithmic gradient quantization framework aimed at reducing memory overhead and computational complexity during the training of deep neural networks. The approach employs a semilogarithmic companding function parameterized by a dynamically adjusted scaling factor A, which evolves in response to the statistical properties of gradients throughout the training process. Two distinct quantization strategies are proposed and evaluated: The switching piecewise A-quantizer, which adaptively toggles between low-bit uniform and high-bit semilogarithmic quantization according to an exponentially weighted moving-average (EMA) estimate of gradient variance; and the hybrid A-quantizer, which statically partitions the gradient domain, applying uniform quantization in low-magnitude regions and semilogarithmic companding in high-magnitude regions. The proposed methods are empirically evaluated on both multilayer perceptron (MLP) and convolutional neural network (CNN) architectures using tabular and image-classification benchmarks, including DCCC, CIFAR-10, CIFAR-100, and ImageNet. Quantitative results demonstrate that both models achieve comparable classification accuracy to full-precision (FP32) baselines while significantly reducing gradient reconstruction error. Notably, the hybrid A-quantizer consistently yields better validation accuracy, reduced RMSE, and improved convergence behavior relative to its switching counterpart. These findings underscore the effectiveness of hybrid semilogarithmic quantization as a robust and efficient solution for training deep models in resource-constrained or bandwidth-limited environments, with strong potential for scalable deployment across diverse hardware platforms. Full article
(This article belongs to the Special Issue Deep Neural Networks and Optimization Algorithms (2nd Edition))
Show Figures

Figure 1

21 pages, 5740 KB  
Article
A Low-Power Mixed-Signal Differential In-Memory Matrix–Vector Computing Circuit Architecture with RISC-V Control for Edge AI
by David Ng, King Hang Lam, Si Qi Bu, Wen Chin Lo, Chi Hong Chan, Roy Ng, Sunny Chan, Matt Mak, Hugo Wong, Steve Chim, Patrick Chang, Raymond Chik, Steven Wong and Wai Ming To
J. Low Power Electron. Appl. 2026, 16(3), 22; https://doi.org/10.3390/jlpea16030022 - 24 Jun 2026
Viewed by 644
Abstract
Analog in-memory computing (AIMC) has emerged as a promising approach to mitigate the Von Neumann bottleneck in matrix operations, which are common in deep learning applications. However, the practical implementation of resistive crossbar arrays is limited by challenges in signed weight representation, conductance [...] Read more.
Analog in-memory computing (AIMC) has emerged as a promising approach to mitigate the Von Neumann bottleneck in matrix operations, which are common in deep learning applications. However, the practical implementation of resistive crossbar arrays is limited by challenges in signed weight representation, conductance quantization, and device nonlinearity. This paper presents a differential mixed-signal architecture for accurate signed matrix–vector multiplication (MVM), integrated with a RISC-V microcontroller for edge inference applications. A structured digital-to-analog mapping framework encodes quantized neural network weights into programmable conductance values while preserving arithmetic correctness. The design employs voltage-mode input encoding, differential current summation, and transimpedance-based readout followed by analog-to-digital conversion, enabling single-cycle signed accumulation without duplicating crossbar resources. A 32 × 16 dual-layer prototype crossbar was fabricated and experimentally characterized. Measurements demonstrate a mean absolute percentage error (MAPE) below 1% within the linear operating region and below 4% over the full-scale conductance range. These results validate the robustness of the proposed mapping methodology and confirm the feasibility of hybrid analog–digital acceleration for edge AI systems. Consequently, this discrete prototype serves as a physical verification platform for the AIMC approach, providing valuable insights for more efficient mixed-signal computing integrated circuit (IC) designs. Full article
(This article belongs to the Topic Advanced Integrated Circuit Design and Application)
Show Figures

Graphical abstract

13 pages, 3658 KB  
Article
TR-ABFT: Tile-Resilient Fault Detection for Neural Processing Units
by Yang Hua, Yunhong Bai, Bo Wang, Wei Zhuang and Yuanfu Zhao
Electronics 2026, 15(12), 2715; https://doi.org/10.3390/electronics15122715 - 19 Jun 2026
Viewed by 319
Abstract
Spaceborne neural processing units (NPUs) increasingly support real-time deep-learning inference, but their dense multiply-accumulate arrays are vulnerable to radiation-induced soft errors. Conventional radiation-hardening methods improve reliability through hardware redundancy, but they incur substantial area, performance and compiler-mapping overheads. This paper proposes tile-resilient algorithm-based [...] Read more.
Spaceborne neural processing units (NPUs) increasingly support real-time deep-learning inference, but their dense multiply-accumulate arrays are vulnerable to radiation-induced soft errors. Conventional radiation-hardening methods improve reliability through hardware redundancy, but they incur substantial area, performance and compiler-mapping overheads. This paper proposes tile-resilient algorithm-based fault tolerance (TR-ABFT), a software-scheduled, detection-oriented scheme for quantized NPU inference. TR-ABFT generates checksum information at tile granularity and maps checking tasks onto the original processing element (PE) array without changing the hardware topology. To make ABFT compatible with INT8 datapaths, we design two checksum-coding strategies: checksum decomposition and modulo-239 checksum coding. The modulo-239 scheme removes structural missed detections for two-bit flips with bit-position spacings in (1, 31), while preserving compatibility with signed INT8 inputs. Evaluations on ResNet, YOLOv8, and RT-DETR show that, on a 16×16 array, TR-ABFT introduces only 6.37% to 24.61% additional computational overhead. By converting spatial redundancy into schedulable temporal redundancy, TR-ABFT preserves systolic-array regularity and provides a low-overhead reliability-enhancement mechanism for space-grade neural-network accelerators. Full article
(This article belongs to the Special Issue Artificial Intelligence and Microsystems)
Show Figures

Figure 1

26 pages, 3114 KB  
Article
Design and Evaluation of a Compact CNN for EMG-Based Wearable Systems Under Embedded Constraints
by Valentina Tirsu, Andrei Dorogan, Lilia Sava, Larisa Dunai, Alexandru Ilev and Nelea Manin
Sensors 2026, 26(12), 3862; https://doi.org/10.3390/s26123862 - 17 Jun 2026
Cited by 1 | Viewed by 329
Abstract
Electromyographic (EMG) signals are increasingly used in wearable cyber–physical systems (CPS), where reliable movement recognition must be achieved under limited computational resources. In this study, we present a compact EMG processing framework that integrates signal acquisition, preprocessing, segmentation, and movement classification within a [...] Read more.
Electromyographic (EMG) signals are increasingly used in wearable cyber–physical systems (CPS), where reliable movement recognition must be achieved under limited computational resources. In this study, we present a compact EMG processing framework that integrates signal acquisition, preprocessing, segmentation, and movement classification within a unified pipeline designed for embedded-oriented applications. The proposed approach combines a multi-channel EMG acquisition system with a lightweight one-dimensional convolutional neural network (1D CNN) developed according to TinyML principles, withprocessing input windows of size 32 × 3 and low computational complexity and memory requirements. Experimental evaluation was conducted on a dataset collected from 15 participants performing squat, walking, and running activities under realistic acquisition conditions. The proposed model achieved an accuracy of 0.9135, an F1-score of 0.9124, and a ROC AUC of approximately 0.96, demonstrating reliable classification performance. Following 8-bit quantization, the model size was reduced to approximately 2 KB, supporting deployment on resource-constrained embedded platforms. The results show that compact CNN architectures can effectively classify EMG-based movement patterns while maintaining a small computational footprint, providing a practical foundation for future wearable CPS and TinyML-enabled applications. Full article
(This article belongs to the Section Wearables)
Show Figures

Figure 1

32 pages, 22589 KB  
Article
Blood Typing at the Edge: A Hybrid Deep Learning Pipeline for Point-of-Care Blood Type Classification
by Bruno Silva, Enmanuel Abilheira, Ljiljana Dukanovic, Afonso Pinheiro and Vítor Carvalho
Appl. Sci. 2026, 16(12), 6089; https://doi.org/10.3390/app16126089 - 16 Jun 2026
Viewed by 227
Abstract
Blood typing remains a manual, subjective procedure when not reliant on centralized laboratory infrastructure. This study presents an automated blood typing system for point-of-care deployment, developed in collaboration with CRIAM, whose portable device captures reaction images for in vitro diagnostics. The system integrates [...] Read more.
Blood typing remains a manual, subjective procedure when not reliant on centralized laboratory infrastructure. This study presents an automated blood typing system for point-of-care deployment, developed in collaboration with CRIAM, whose portable device captures reaction images for in vitro diagnostics. The system integrates computer vision and artificial intelligence to classify these reactions automatically. Fourteen classification pipelines were trained and evaluated with a 3090-image dataset, encompassing fine-tuned convolutional neural networks, raw pixel-based classifiers, and hybrid architectures pairing pretrained embeddings from DINOv2 and EfficientNet-B4 with lightweight classifiers. Embedding-based approaches consistently outperformed alternatives in accuracy and cross-fold stability. The best pipeline, in terms of performance and suitability for low-power devices, combined DINOv2-small embeddings with logistic regression, achieving 99.87 ± 0.12% mean accuracy. After 8-bit integers (INT8) quantization and retraining with data augmentation, accuracy improved to 99.97 ± 0.03%, surpassing the uncompressed baseline. All misclassifications were traced to borderline weak-positive Rh/D reactions, confirming errors are localized and explainable. Held-out validation on 856 images yielded 99.53% accuracy, with the single error attributed to a lighting artifact. While deployment on a legacy 32-bit CPU prototype processes four images in approximately 4.7 min, hardware benchmarking confirmed feasibility, from a Raspberry Pi Zero 2W to high-end mobile processors. These results establish quantized embedding-driven architectures as a solution for automated blood typing in point-of-care and resource-limited settings. Full article
Show Figures

Figure 1

19 pages, 11225 KB  
Article
Accelerated Graph Neural Networks on an SoC FPGA for Onboard LEO Satellite Network Routing
by Jinhyung Park, Heoncheol Lee, Sungryul Kim, Bongsoo Roh and Myonghun Han
Electronics 2026, 15(12), 2664; https://doi.org/10.3390/electronics15122664 - 16 Jun 2026
Viewed by 359
Abstract
This paper presents a system-on-chip field-programmable gate array (SoC FPGA) acceleration architecture for graph-neural-network- and deep-reinforcement-learning (GNN–DRL)-based routing inference in low-Earth-orbit (LEO) satellite networks. Because LEO satellites move at high orbital speeds, the network topology changes continuously, and routing decisions must track the [...] Read more.
This paper presents a system-on-chip field-programmable gate array (SoC FPGA) acceleration architecture for graph-neural-network- and deep-reinforcement-learning (GNN–DRL)-based routing inference in low-Earth-orbit (LEO) satellite networks. Because LEO satellites move at high orbital speeds, the network topology changes continuously, and routing decisions must track the current link state rather than rely only on static rules. GNN-based DRL routing can represent the graph structure of the network when selecting paths, but its message-passing and readout stages are computationally expensive for resource-constrained onboard platforms. To address this limitation, the trained GNN routing model is ported to an SoC FPGA and implemented with a collaborative processing-system (PS) and programmable-logic (PL) architecture. The PS handles candidate-path generation, environment setup, path selection, and network-state updates, whereas the PL executes the computationally dominant message-passing neural network (MPNN) and readout layers. Post-training INT8 quantization, nonlinear-function approximation, vector-level parallelization, and a parallel multiply–accumulate structure are applied to reduce memory pressure and execution time. Experiments on a ZCU104 board using a PYNQ-controlled PS–PL implementation and an NSFNET-based routing environment show that the proposed PS–PL structure reduces the evaluation time from 94.08 s to 12.63 s compared with the PS-only implementation while maintaining an evaluation score close to that of the original model. Full article
(This article belongs to the Special Issue Recent Advances in AI Hardware Design)
Show Figures

Figure 1

56 pages, 6689 KB  
Review
AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence
by Mohamed M. Morsy
Electronics 2026, 15(12), 2645; https://doi.org/10.3390/electronics15122645 - 15 Jun 2026
Viewed by 2023
Abstract
The rapid growth of artificial intelligence (AI) workloads is reshaping semiconductor design across architecture, interconnect, memory hierarchy, packaging, timing, and design automation. Rather than converging on a single hardware solution, the field is expanding into a heterogeneous ecosystem that includes data-center graphics processing [...] Read more.
The rapid growth of artificial intelligence (AI) workloads is reshaping semiconductor design across architecture, interconnect, memory hierarchy, packaging, timing, and design automation. Rather than converging on a single hardware solution, the field is expanding into a heterogeneous ecosystem that includes data-center graphics processing units (GPUs), edge neural processing units (NPUs), and application-specific integrated circuits (ASICs), field-programmable gate array (FPGA)-based and hybrid AI system-on-chip (SoC) platforms, chiplet-enabled systems, and emerging beyond-conventional-silicon approaches such as photonic, neuromorphic, and analog in-memory processors. This paper presents a comprehensive review of AI-on-chip systems from a cross-layer perspective. It examines AI chip architectures and hardware platforms, network-on-chip (NoC) designs for AI communication patterns, and algorithm–hardware co-design methods for model acceleration, including compression, quantization, and sparsity-aware optimization. It also reviews clocking, synchronization, and clock-domain-crossing (CDC) challenges in large heterogeneous systems and chiplets, as well as manufacturing, advanced packaging, and reliability issues, including two-and-a-half-dimensional (2.5D) and three-dimensional (3D) integration, thermal and mechanical constraints, assembly quality, and long-term yield considerations. In parallel, the paper surveys the growing role of AI in chip design itself, covering machine-learning-assisted analysis, Bayesian and reinforcement-learning-based optimization, and the emerging use of large language models (LLMs) and AI agents for register-transfer level (RTL) generation, design-space exploration, and autonomous electronic design automation (EDA) workflows. Finally, it discusses beyond-silicon AI chip directions and the broader economic and industry context shaping cloud, on-premises, and edge deployment. By integrating these topics into a unified framework, this review highlights the key technological drivers, system-level tradeoffs, and future research directions that will define next-generation scalable, reliable, and energy-efficient AI-on-chip systems. Full article
(This article belongs to the Topic AI Agents: Progress, Architecture, and Applications)
Show Figures

Figure 1

20 pages, 4583 KB  
Article
Optimizing Convolutional Operation and Dataflow in FPGA Acceleration of Bayesian Convolutional Neural Network
by Shulei Wang, Yun Ling, Daolin Cai, Hao Zhang, Mingxin Liu, Cheng Cheng, Qihang Ding, Zhu Fu, Jiale Zhao, Haoyu Zhou and Junxin Zhang
Electronics 2026, 15(12), 2603; https://doi.org/10.3390/electronics15122603 - 12 Jun 2026
Viewed by 274
Abstract
A Bayesian convolutional neural network (BCNN) quantifies prediction uncertainty by introducing randomness into weights or activations, which is important for safety-critical applications such as medical diagnosis and autonomous driving. However, BCNN inference typically relies on Monte Carlo sampling requiring multiple forward passes, leading [...] Read more.
A Bayesian convolutional neural network (BCNN) quantifies prediction uncertainty by introducing randomness into weights or activations, which is important for safety-critical applications such as medical diagnosis and autonomous driving. However, BCNN inference typically relies on Monte Carlo sampling requiring multiple forward passes, leading to computation and energy consumption far beyond standard CNN hardware acceleration. FPGA, with its parallel processing, reconfigurability, and high-energy efficiency, are ideal platforms for dedicated BCNN accelerators. This paper designs and implements an FPGA acceleration method for BCNN-using high-level synthesis. First, convolution, pooling, and fully connected modules are individually optimized. Then, a mean/variance dual-path parallel expansion is adopted, combined with mixed-precision quantization and global scaling compensation, local reparameterization sampling, parameter reordering, and ping-pong buffering, achieving low resource usage and high-energy efficiency while enabling uncertainty evaluation. Experimental results on Bayes VGG16 show resource utilization of 24,776 LUT, 23,378 FF, 115 BRAM, and 129 DSP, with total power of 2.049 W. Compared with an unoptimized Bayesian implementation, the proposed design reduces inference latency to about one-third, and its latency is only 17% higher than that of the classical VGG16. Compared with PC-based floating-point models, the accuracy loss on four BCNN models (tested on CIFAR-10) is within 1%. The predictive entropy effectively distinguishes normal, noisy, and out-of-distribution (OOD) samples, validating the uncertainty quantification capability of the BCNN FPGA accelerator. Full article
Show Figures

Figure 1

33 pages, 981 KB  
Article
A Cascaded Quantized Spiking Neural Network for Real-Time ECG Arrhythmia Detection on Edge Hardware
by Olamilekan Banjo and Behnaz Ghoraani
Sensors 2026, 26(12), 3723; https://doi.org/10.3390/s26123723 - 11 Jun 2026
Viewed by 304
Abstract
Wearable ECG monitors enable continuous cardiac surveillance, but most still rely on cloud-based analysis with limited on-device support for multi-class arrhythmia detection. Spiking neural networks (SNNs) are promising for low-power edge inference, yet it remains unclear how class-imbalance loss design interacts with RR-interval [...] Read more.
Wearable ECG monitors enable continuous cardiac surveillance, but most still rely on cloud-based analysis with limited on-device support for multi-class arrhythmia detection. Spiking neural networks (SNNs) are promising for low-power edge inference, yet it remains unclear how class-imbalance loss design interacts with RR-interval features in directly trained quantized SNNs, and FPGA validation in this setting is largely unexplored. We propose a quantized convolutional spiking neural network (QCSNN) for real-time arrhythmia detection on resource-constrained hardware. The model uses a dual-head architecture that jointly trains binary and four-class classifiers, subsequently reorganized into a cascaded pipeline that routes only abnormal beats to the second stage. At inference, beats classified as Normal exit at Stage 1; only beats classified as Abnormal are routed to the four-class head, so the bulk of the inference cost is absorbed by Stage 1. We evaluate two loss functions, Cross-Entropy and Focal Loss, under four RR-feature routing strategies. Without RR features, Focal Loss improves macro F1 by 2.3–2.5% over Cross-Entropy (mean Δ = +0.013 in Stage-2 macro F1; Wilcoxon two-sided p = 0.031). With RR features, this advantage largely disappears (Wilcoxon two-sided p ≥ 0.219 at all RR routings); meanwhile, RR features at the strongest routing improve Stage-2 macro F1 by +0.028 to +0.034 depending on loss function—a gain that exceeds the entire Focal-Loss-over-Cross-Entropy advantage, suggesting that RR features provide discriminative information that compensates for class imbalance at the input level. Based on clinically prioritized sensitivity, the CE:RR→Both configuration was deployed on a PYNQ-Z2 FPGA, achieving 99.02% cascaded accuracy, 11.54 ms per-beat latency, and 0.33 W accelerator power—a 31.66× power reduction and 4.01× energy reduction versus GPU inference, within 1% macro F1. These results demonstrate quantized SNNs as a practical solution for real-time edge arrhythmia monitoring that operates independently of cloud connectivity—removing the network-dependent latency, connectivity-dropout failure modes, and continuous-transmission energy burden that constrain current wearable monitors and, to our knowledge, represent one of the first systematic studies of loss-function/RR-feature interactions in directly trained SNN arrhythmia classification and one of the first FPGA deployments of a fully quantized, directly trained SNN for multi-class ECG arrhythmia detection. All code generated and used in this study has been made publicly available. Full article
(This article belongs to the Section Biomedical Sensors)
Show Figures

Figure 1

9 pages, 1632 KB  
Proceeding Paper
Hardware Implementation of an Autoencoder on a Field Programmable Gate Array
by Minh-Hieu Vo, Thien-Van Nguyen, Trong-Nhan Huynh, Tan-Phat Dang and Huu-Thuan Huynh
Eng. Proc. 2026, 141(1), 11; https://doi.org/10.3390/engproc2026141011 - 10 Jun 2026
Viewed by 168
Abstract
An autoencoder is an unsupervised deep learning architecture designed to compress input data, extract meaningful features, and reconstruct the original input for applications such as anomaly detection and data compression. However, CPU-based implementations often suffer from limited performance and high power consumption. To [...] Read more.
An autoencoder is an unsupervised deep learning architecture designed to compress input data, extract meaningful features, and reconstruct the original input for applications such as anomaly detection and data compression. However, CPU-based implementations often suffer from limited performance and high power consumption. To address these challenges, this paper presents an FPGA-based autoencoder with a hardware-friendly neural network architecture optimized for both resource utilization and processing performance. In addition, optimization techniques such as network size reduction, quantization, and pipelining are applied to improve efficiency in real-time applications. The proposed autoencoder accelerator is integrated into a Nios II system to evaluate its effectiveness. Implemented on a Cyclone V 5CSXFC6D6F31C6 FPGA (Intel Corporation, San Jose, California, United States) at 50 MHz, the system occupies 81% of logic resources, 3% of memory blocks, and 3% of digital signal processing blocks. Experimental results show that, while an Intel Xeon CPU at 2.2 GHz requires more than 0.2 s to process a single handwritten digit from the Modified National Institute of Standards and Technology dataset, the proposed system performs the same task in approximately 4.5 milliseconds, providing a 44× speedup. This demonstrates the effectiveness of the proposed FPGA-based autoencoder accelerator. Full article
Show Figures

Figure 1

12 pages, 9413 KB  
Communication
Photosensing PUF from an Intrinsically Random SnTe Memristor for Image Encryption and Recognition
by Wendi Xu, Jia Zhang, Junjie Xie, Tianzhu Xu, Jia Wu and Hong Wang
Nanomaterials 2026, 16(12), 715; https://doi.org/10.3390/nano16120715 - 10 Jun 2026
Viewed by 408
Abstract
Physical unclonable function (PUF) based on intrinsic device randomness has emerged as promising hardware security primitives, yet combining secure encryption with neuromorphic recognition within a single device platform remains challenging. Here, we demonstrate a photosensing PUF based on an intrinsically random SnTe memristor [...] Read more.
Physical unclonable function (PUF) based on intrinsic device randomness has emerged as promising hardware security primitives, yet combining secure encryption with neuromorphic recognition within a single device platform remains challenging. Here, we demonstrate a photosensing PUF based on an intrinsically random SnTe memristor capable of both image encryption and memristive neural network recognition. The SnTe memristor, fabricated with an In2O3:SnO2/SnTe/Nb:SrTiO3 structure, exhibits stable resistive switching and stable retention exceeding 4000 s. Synaptic biomimetic behaviors including learning-experience emulation, short-term plasticity and long-term plasticity are also realized. Notably, the device displays pronounced optical sensitivity that produces stochastic photocurrent fluctuations originating from unavoidable device-to-device variations under illumination. By quantizing these random photocurrents, an encryption key stream is generated and utilized for image scrambling and diffusion. A memristive neural network is constructed to classify the encrypted images, achieving a recognition accuracy of 95.1% with a loss of 0.15 after 300 training epochs. This work establishes a viable pathway from intrinsic optical randomness to secure neuromorphic computing, highlighting the multifunctional potential of SnTe memristors in integrated hardware security and brain-inspired computation. Full article
Show Figures

Graphical abstract

Back to TopTop