Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (162)

Search Parameters:
Keywords = ASIC (Application Specific Integrated Circuit)

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
38 pages, 990 KB  
Article
Permutation-Equivariant Graph Reinforcement Learning for Thermal-Aware Microservice Scheduling in Co-Packaged Optics Data Centers
by Zhaoqi Qiu, Linya Peng, Fuming Fan, Haoran Zuo, Wenjie Qiu, Bo Xu and Tianping Deng
Symmetry 2026, 18(7), 1182; https://doi.org/10.3390/sym18071182 - 13 Jul 2026
Viewed by 553
Abstract
As data-center interconnects move to co-packaged optics (CPO), high-power application-specific integrated circuits (ASICs) and heat-sensitive optical engines share a single interposer, and the resulting intra-module thermal coupling overwhelms conventional schedulers. Thermal-aware microservice directed acyclic graph (DAG) scheduling on CPO modules is recast here [...] Read more.
As data-center interconnects move to co-packaged optics (CPO), high-power application-specific integrated circuits (ASICs) and heat-sensitive optical engines share a single interposer, and the resulting intra-module thermal coupling overwhelms conventional schedulers. Thermal-aware microservice directed acyclic graph (DAG) scheduling on CPO modules is recast here as a question of symmetry. Whereas a homogeneous-graph policy assumes the full node-permutation symmetry, that symmetry is broken twice: by the distinct task and processor node types, and by the asymmetric ASIC–engine coupling. We therefore propose a heterogeneous-graph Proximal Policy Optimization (PPO) scheduler in which the placement head is permutation-equivariant, the delay head is permutation-invariant, and the parameterization stays invariant to the processor count N. Because these symmetries hold by construction, the policy transfers zero-shot across module sizes. Heterogeneous edge typing and the resistor–capacitor (RC) coupling edge attribute are isolated by a six-test ablation chain. Evaluated on the Alibaba 2021 microservice trace across all module sizes and ambient regimes under the standard auto-cool budget, the proposed scheduler cuts the thermal-violation rate from roughly 98% under the Heterogeneous Earliest-Finish-Time (HEFT) heuristic to about 0.3%; at the hot operating point it lowers peak temperature by about 25% and raises DAG completion from about 26% to 100%, with the rare residual violations most frequent in the extreme-ambient band. With the env auto-cool budget disabled, a controlled single-axis comparison shows that removing the RC-coupling edge attribute raises the violation rate by over an order of magnitude, isolating its contribution. A single parameter set serves every N without retraining. Full article
(This article belongs to the Section A: Computer Science)
►▼ Show Figures

Figure 1

19 pages, 5237 KB  
Article
Distributed Wireless Neural Recording System for Multi-Region Brain Activity Monitoring
by Liu Yang, Changhua You, Gang Wang, Xuan Zhang, Canyang Wang, Bo Cheng, Zhengtuo Zhao, Ning Xue and Lei Yao
Biosensors 2026, 16(7), 370; https://doi.org/10.3390/bios16070370 - 7 Jul 2026
Viewed by 1329
Abstract
Distributed neural interfaces for multi-region implantation require both scalable interconnects and robust telemetry, yet conventional centralized or fully distributed architectures often trade-off wiring complexity, resource reuse, and transmission stability. This work presents a distributed wireless neural recording system based on a parallel-link architecture [...] Read more.
Distributed neural interfaces for multi-region implantation require both scalable interconnects and robust telemetry, yet conventional centralized or fully distributed architectures often trade-off wiring complexity, resource reuse, and transmission stability. This work presents a distributed wireless neural recording system based on a parallel-link architecture and a custom 12-channel neural recording Application-Specific Integrated Circuit (ASIC). Each remote module is connected to a central hub through an independent four-wire link (VDD/GND/LVDS±). The ASIC integrates modular digital pixels (MDPs), an on-chip oscillator, a Manchester encoding, and a Low-Voltage Differential Signaling (LVDS) output to reduce interconnect count while maintaining reliable serial transmission. Fabricated in SMIC 0.18 μm CMOS, the chip occupies 4.84 mm × 0.36 mm and consumes 10.13 mW in total, with 48.5 μW/channel consumed by the recording channels excluding the LVDS driver. It achieves 5.6 μVrms input-referred noise and a measured per-channel sampling rate of 28.93 kSps. A compact 20 mm2 recording module and an FPGA-based central hub with real-time decoding and compression were implemented for validation. In vivo mouse experiments demonstrate clear action-potential recordings across 12 channels, confirming the feasibility of stable and scalable multi-region neural signal acquisition. Full article
►▼ Show Figures

Figure 1

17 pages, 6434 KB  
Communication
Design of a SoC-Based Highly Integrated RF Transceiver Module
by Jianxi Wu, Hao Zhou, Linfeng Shang, Yawei Shao and Kan Wang
Sensors 2026, 26(13), 4173; https://doi.org/10.3390/s26134173 - 2 Jul 2026
Viewed by 1128
Abstract
To address the issues of high customization, long development cycles, and excessive power/volume in radio frequency (RF) transceiver modules for Synthetic Aperture Radar (SAR) and radar systems, this paper presents an ultra-compact universal RF transceiver module design based on a full application-specific integrated [...] Read more.
To address the issues of high customization, long development cycles, and excessive power/volume in radio frequency (RF) transceiver modules for Synthetic Aperture Radar (SAR) and radar systems, this paper presents an ultra-compact universal RF transceiver module design based on a full application-specific integrated circuit (ASIC) architecture. Centered on a wideband RF System-on-chip (SoC) and a reconfigurable digital SoC, the module integrates the complete RF transceiver chain—including filtering, amplification, mixing, Analog-to-Digital/Digital-to-Analog Converter (ADC/DAC) conversion, digital preprocessing, and high-speed data transmission. Test results demonstrate that the 8-channel module achieves a 53.1% area reduction and 55.1% lower power consumption (only 40.9 W) compared with conventional architectures, while all key RF specifications meet system requirements. The proposed solution improves upon existing limitations in high integration, low power, and generality, offering a low-cost, rapid-development technical route for transceiver modules in radar and communication applications. Full article
(This article belongs to the Section Radar Sensors)
►▼ Show Figures

Figure 1

28 pages, 872 KB  
Article
An Optimized Floating-Point Unit Set for FPGA-Based DSP: Improving Area, Energy, and Throughput Trade-Offs
by Fernando Flores, Juan Portela Queimaño, Jesús Manuel Costa Pazo, María Dolores Valdés-Peña, Camilo Quintáns Graña and José Manuel Villapún Sánchez
Electronics 2026, 15(13), 2850; https://doi.org/10.3390/electronics15132850 - 30 Jun 2026
Viewed by 914
Abstract
Floating-point arithmetic provides the dynamic range that fixed-point lacks for digital signal processing (DSP) algorithms with widely varying operand magnitudes. This work presents a parameterizable floating-point unit set for field programmable gate array (FPGA)-based DSP. The set consists of five units: adder/subtractor, multiplier, [...] Read more.
Floating-point arithmetic provides the dynamic range that fixed-point lacks for digital signal processing (DSP) algorithms with widely varying operand magnitudes. This work presents a parameterizable floating-point unit set for field programmable gate array (FPGA)-based DSP. The set consists of five units: adder/subtractor, multiplier, multiply–accumulate (MAC), fixed-to-float and float-to-fixed converters. Two architectural choices distinguish the proposed format from IEEE-754: configurable exponent and mantissa widths during synthesis and a 0.f significand encoding that reduces corner-case logic at the cost of one additional mantissa bit. The format is therefore IEEE-754-inspired rather than fully compliant: special values (NaN, ±∞) are not implemented, and overflow and underflow are handled through saturation to predefined constants. The design is implemented in standard VHDL-2008 without relying on high-level synthesis (HLS) tools or vendor-specific primitives, ensuring portability across different FPGA families and application-specific integrated circuits (ASICs). The multiplier and MAC are evaluated in two configurations: inferring DSP blocks or look-up table (LUT)-only, both close timing at 300MHz on Artix-7 and Kintex Ultrascale devices. The proposed blocks outperform vendor IP Cores and recent academic designs in terms of area-throughput-power (ATP), achieving improvements from 10% to 108%, except for the adder/subtractor, which does not outperform two optimized Xilinx IP cores (HS-R and HS-P) and is therefore included for design coherence rather than as a strict resource improvement over all vendor IPs. All these blocks meet the theoretical error bound, and a representative 200-tap finite impulse response (FIR) filter built from them closes timing at 300MHz with 76% LUT utilization. Full article
(This article belongs to the Special Issue Design and Application of Digital Circuit and Systems)
►▼ Show Figures

Figure 1

56 pages, 6689 KB  
Review
AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence
by Mohamed M. Morsy
Electronics 2026, 15(12), 2645; https://doi.org/10.3390/electronics15122645 - 15 Jun 2026
Viewed by 3847
Abstract
The rapid growth of artificial intelligence (AI) workloads is reshaping semiconductor design across architecture, interconnect, memory hierarchy, packaging, timing, and design automation. Rather than converging on a single hardware solution, the field is expanding into a heterogeneous ecosystem that includes data-center graphics processing [...] Read more.
The rapid growth of artificial intelligence (AI) workloads is reshaping semiconductor design across architecture, interconnect, memory hierarchy, packaging, timing, and design automation. Rather than converging on a single hardware solution, the field is expanding into a heterogeneous ecosystem that includes data-center graphics processing units (GPUs), edge neural processing units (NPUs), and application-specific integrated circuits (ASICs), field-programmable gate array (FPGA)-based and hybrid AI system-on-chip (SoC) platforms, chiplet-enabled systems, and emerging beyond-conventional-silicon approaches such as photonic, neuromorphic, and analog in-memory processors. This paper presents a comprehensive review of AI-on-chip systems from a cross-layer perspective. It examines AI chip architectures and hardware platforms, network-on-chip (NoC) designs for AI communication patterns, and algorithm–hardware co-design methods for model acceleration, including compression, quantization, and sparsity-aware optimization. It also reviews clocking, synchronization, and clock-domain-crossing (CDC) challenges in large heterogeneous systems and chiplets, as well as manufacturing, advanced packaging, and reliability issues, including two-and-a-half-dimensional (2.5D) and three-dimensional (3D) integration, thermal and mechanical constraints, assembly quality, and long-term yield considerations. In parallel, the paper surveys the growing role of AI in chip design itself, covering machine-learning-assisted analysis, Bayesian and reinforcement-learning-based optimization, and the emerging use of large language models (LLMs) and AI agents for register-transfer level (RTL) generation, design-space exploration, and autonomous electronic design automation (EDA) workflows. Finally, it discusses beyond-silicon AI chip directions and the broader economic and industry context shaping cloud, on-premises, and edge deployment. By integrating these topics into a unified framework, this review highlights the key technological drivers, system-level tradeoffs, and future research directions that will define next-generation scalable, reliable, and energy-efficient AI-on-chip systems. Full article
(This article belongs to the Topic AI Agents: Progress, Architecture, and Applications)
►▼ Show Figures

Figure 1

16 pages, 2379 KB  
Article
A Novel Standard Cell Structure and Physical Design Methodology to Enhance Routability
by Seongjun Lee and Changho Han
Electronics 2026, 15(8), 1690; https://doi.org/10.3390/electronics15081690 - 17 Apr 2026
Viewed by 1415
Abstract
In the era of highly integrated circuits, continuous miniaturization has significantly increased routing complexity, thereby directly impacting circuit performance. As process scaling advances and the number of on-chip metal layers increases, conventional standard cell libraries face limitations that cause severe routing bottlenecks. To [...] Read more.
In the era of highly integrated circuits, continuous miniaturization has significantly increased routing complexity, thereby directly impacting circuit performance. As process scaling advances and the number of on-chip metal layers increases, conventional standard cell libraries face limitations that cause severe routing bottlenecks. To overcome these limitations, this paper proposes a dual-component approach. First, we introduce a novel standard cell structure that improves routing flexibility by expanding the degrees of freedom for pin access, particularly in highly congested regions. Second, we present a physical design methodology specifically designed to ensure seamless integration with existing electronic design automation (EDA) tools, allowing new cells to be effectively placed and routed without major modifications to current flows. The proposed approach was validated using the open-source ASAP7 process design kit (PDK). Experimental results confirm significant reductions in via count and total wirelength, leading to improved routability, reduced power consumption, and enhanced performance. These findings demonstrate that combining the new cell architecture with a tailored design methodology provides a practical alternative to conventional solutions, enabling more efficient and scalable circuit designs for future technology nodes. Full article
(This article belongs to the Section Circuit and Signal Processing)
►▼ Show Figures

Figure 1

19 pages, 6307 KB  
Article
Design of a Compact Space Search Coil Magnetometer
by Yunho Jang, Ho Jin, Minjae Kim, Ik-Joon Chang, Ickhyun Song and Chae Kyung Sim
Sensors 2026, 26(8), 2415; https://doi.org/10.3390/s26082415 - 15 Apr 2026
Cited by 1 | Viewed by 1067
Abstract
Search coil magnetometers (SCMs) are widely used in space science missions to measure time-varying magnetic fields. However, conventional SCM designs often increase sensor mass and electronic power consumption in order to meet mission-specific sensitivity requirements. This study presents the design and ground-based test [...] Read more.
Search coil magnetometers (SCMs) are widely used in space science missions to measure time-varying magnetic fields. However, conventional SCM designs often increase sensor mass and electronic power consumption in order to meet mission-specific sensitivity requirements. This study presents the design and ground-based test results of a space search coil magnetometer (SSCM) concept aimed at reducing sensor mass and electronic power consumption while maintaining practical system operability for platform-constrained missions. Mass reduction was achieved by adopting a rolling-sheet core configuration. In addition, printed circuit board (PCB)-based interconnections between segmented windings were implemented to improve the reproducibility of assembly and mechanical robustness without additional structural complexity. Power reduction was achieved by employing an application-specific integrated circuit (ASIC)-based sensor amplifier and a compact control electronic unit implemented as a modular stack with a 1U CubeSat standard board form factor. Performance tests confirmed the stable operation of the integrated sensor–electronics chain over the target measurement band. The system-level noise-equivalent magnetic induction (NEMI) measured under laboratory conditions was 33 fT/√Hz at 1 kHz. Environmental tests including vibration and thermal cycling were performed to further verify the structural safety and functional stability of the sensor assembly under space-relevant conditions. The proposed SSCM architecture provides a practical approach for implementing low-mass and low-power magnetic field instruments for platform-constrained space missions. Full article
(This article belongs to the Special Issue Smart Magnetic Sensors and Application)
►▼ Show Figures

Figure 1

20 pages, 3159 KB  
Article
ROM-Less Co(Sine) Synthesizer
by Florentina-Giulia Stoica, Alex Calinescu and Marius Enachescu
Electronics 2026, 15(5), 1093; https://doi.org/10.3390/electronics15051093 - 5 Mar 2026
Viewed by 2527
Abstract
Sine and cosine wave synthesis is utilized for generating sinusoidal-like values in the digital domain. While this task is commonly handled through software, dedicated hardware like Direct Digital Synthesis (DDS) is also available. However, both methods rely on memory resources, such as look-up [...] Read more.
Sine and cosine wave synthesis is utilized for generating sinusoidal-like values in the digital domain. While this task is commonly handled through software, dedicated hardware like Direct Digital Synthesis (DDS) is also available. However, both methods rely on memory resources, such as look-up tables and Read-Only Memories (ROMs), which face latency limitations related to additional memory access times on top of additional Si area. With the advent of real-time arithmetic for sine wave approximation, this paper presents a digital module that employs iterative multiply-accumulate (MAC) operations for sine and cosine synthesis. To support the integration of this module into Systems-on-Chip (SoCs), Field-Programmable Gate Arrays (FPGAs), and standalone Application-Specific Integrated Circuits (ASICs), a comprehensive figure of merit (FoM) comparison against various ROM-less methods is provided. When implemented on a Xilinx (AMD) XC7A100T-3CSG324 FPGA, the proposed architecture compared to other ROM-less solutions like the Taylor approximation, achieves 80.80% lower resource utilization, 80.89% reduced propagation delay, and 36.66% higher accuracy in sine and cosine wave approximation, both operating as 32-bit systems with one sample per clock cycle. Furthermore, the proposed sine accelerator, accompanying control and communication IPs, and custom firmware were deployed on an FPGA-based function generator platform and experimentally validated. Full article
(This article belongs to the Section Circuit and Signal Processing)
►▼ Show Figures

Figure 1

17 pages, 10981 KB  
Article
NeuroGator: A Low-Power Gating System for Asynchronous BCI Based on LFP Brain State Estimation
by Benyuan He, Chunxiu Liu, Zhimei Qi, Ning Xue and Lei Yao
Brain Sci. 2026, 16(2), 141; https://doi.org/10.3390/brainsci16020141 - 28 Jan 2026
Cited by 2 | Viewed by 1092
Abstract
The continuous handling of the large amount of raw data generated by implantable brain–computer interface (BCI) devices requires a large amount of hardware resources and is becoming a bottleneck for implantable BCI systems, particularly for power-constrained wireless systems. To overcome this bottleneck, we [...] Read more.
The continuous handling of the large amount of raw data generated by implantable brain–computer interface (BCI) devices requires a large amount of hardware resources and is becoming a bottleneck for implantable BCI systems, particularly for power-constrained wireless systems. To overcome this bottleneck, we present NeuroGator, an asynchronous gating system using Local Field Potential (LFP) for the implantable BCI system. Unlike a conventional continuous data decoding approach, NeuroGator uses hierarchical state classification to efficiently allocate hardware resources to reduce the data size before handling or transmission. The proposed NeuroGator operates in two stages: Firstly, a low-power hardware silence detector filters out background noise and non-active signals, effectively reducing the data size by approximately 69.4%. Secondly, a Dual-Resolution Gate Recurrent Unit (GRU) model controls the main data processing procedure on the edge side, using a first-level model to scan low-precision LFP data for potential activity and a second-level model to analyze high-precision LFP data for confirmation of an active state. The experiment shows that NeuroGator reduces overall data throughput by 82% while maintaining an F1-Score of 0.95. This architecture allows the Implantable BCI system to stay in an ultra-low-power state for over 85% of its entire operation period. The proposed NeuroGator has been implemented in an Application-Specific Integrated Circuit (ASIC) with a standard 180 nm Complementary Metal Oxide Semiconductor (CMOS) process, occupying a silicon area of 0.006mm2 and consuming 51 nW power. NeuroGator effectively resolves the resource efficiency dilemma for implantable BCI devices, offering a robust paradigm for next-generation asynchronous implantable BCI systems. Full article
(This article belongs to the Special Issue Trends and Challenges in Neuroengineering)
►▼ Show Figures

Figure 1

54 pages, 3083 KB  
Review
A Survey on Green Wireless Sensing: Energy-Efficient Sensing via WiFi CSI and Lightweight Learning
by Rod Koo, Xihao Liang, Deepak Mishra and Aruna Seneviratne
Energies 2026, 19(2), 573; https://doi.org/10.3390/en19020573 - 22 Jan 2026
Cited by 4 | Viewed by 2369
Abstract
Conventional sensing expends energy at three stages: powering dedicated sensors, transmitting measurements, and executing computationally intensive inference. Wireless sensing re-purposes WiFi channel state information (CSI) inherent in every packet, eliminating extra sensors and uplink traffic, though reliance on deep neural networks (DNNs) often [...] Read more.
Conventional sensing expends energy at three stages: powering dedicated sensors, transmitting measurements, and executing computationally intensive inference. Wireless sensing re-purposes WiFi channel state information (CSI) inherent in every packet, eliminating extra sensors and uplink traffic, though reliance on deep neural networks (DNNs) often trained and run on graphics processing units (GPUs) can negate these gains. This review highlights two core energy efficiency levers in CSI-based wireless sensing. First ambient CSI harvesting cuts power use by an order of magnitude compared to radar and active Internet of Things (IoT) sensors. Second, integrated sensing and communication (ISAC) embeds sensing functionality into existing WiFi links, thereby reducing device count, battery waste, and carbon impact. We review conventional handcrafted and accuracy-first methods to set the stage for surveying green learning strategies and lightweight learning techniques, including compact hybrid neural architectures, pruning, knowledge distillation, quantisation, and semi-supervised training that preserve accuracy while reducing model size and memory footprint. We also discuss hardware co-design from low-power microcontrollers to edge application-specific integrated circuits (ASICs) and WiFi firmware extensions that align computation with platform constraints. Finally, we identify open challenges in domain-robust compression, multi-antenna calibration, energy-proportionate model scaling, and standardised joules per inference metrics. Our aim is a practical battery-friendly wireless sensing stack ready for smart home and 6G era deployments. Full article
►▼ Show Figures

Graphical abstract

15 pages, 2027 KB  
Article
Weight Standardization Fractional Binary Neural Network for Image Recognition in Edge Computing
by Chih-Lung Lin, Zi-Qing Liang, Jui-Han Lin, Chun-Chieh Lee and Kuo-Chin Fan
Electronics 2026, 15(2), 481; https://doi.org/10.3390/electronics15020481 - 22 Jan 2026
Viewed by 598
Abstract
In order to achieve better accuracy, modern models have become increasingly large, leading to an exponential increase in computational load, making it challenging to apply them to edge computing. Binary neural networks (BNNs) are models that quantize the filter weights and activations to [...] Read more.
In order to achieve better accuracy, modern models have become increasingly large, leading to an exponential increase in computational load, making it challenging to apply them to edge computing. Binary neural networks (BNNs) are models that quantize the filter weights and activations to 1-bit. These models are highly suitable for small chips like advanced RISC machines (ARMs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), system-on-chips (SoCs) and other edge computing devices. To design a model that is more friendly to edge computing devices, it is crucial to reduce the floating-point operations (FLOPs). Batch normalization (BN) is an essential tool for binary neural networks; however, when convolution layers are quantized to 1-bit, the floating-point computation cost of BN layers becomes significantly high. This paper aims to reduce the floating-point operations by removing the BN layers from the model and introducing the scaled weight standardization convolution (WS-Conv) method to avoid the significant accuracy drop caused by the absence of BN layers, and to enhance the model performance through a series of optimizations, adaptive gradient clipping (AGC) and knowledge distillation (KD). Specifically, our model maintains a competitive computational cost and accuracy, even without BN layers. Furthermore, by incorporating a series of training methods, the model’s accuracy on CIFAR-100 is 0.6% higher than the baseline model, fractional activation BNN (FracBNN), while the total computational load is only 46% of the baseline model. With unchanged binary operations (BOPs), the FLOPs are reduced to nearly zero, making it more suitable for embedded platforms like FPGAs or other edge computers. Full article
(This article belongs to the Special Issue Advances in Algorithm Optimization and Computational Intelligence)
►▼ Show Figures

Figure 1

37 pages, 483 KB  
Review
Lattice-Based Cryptographic Accelerators for the Post-Quantum Era: Architectures, Optimizations, and Implementation Challenges
by Hua Yan, Lei Wu, Qiming Sun and Pengzhou He
Electronics 2026, 15(2), 475; https://doi.org/10.3390/electronics15020475 - 22 Jan 2026
Cited by 1 | Viewed by 5059
Abstract
The imminent threat of large-scale quantum computers to modern public-key cryptographic devices has led to extensive research into post-quantum cryptography (PQC). Lattice-based schemes have proven to be the top candidate among existing PQC schemes due to their strong security guarantees, versatility, and relatively [...] Read more.
The imminent threat of large-scale quantum computers to modern public-key cryptographic devices has led to extensive research into post-quantum cryptography (PQC). Lattice-based schemes have proven to be the top candidate among existing PQC schemes due to their strong security guarantees, versatility, and relatively efficient operations. However, the computational cost of lattice-based algorithms—including various arithmetic operations such as Number Theoretic Transform (NTT), polynomial multiplication, and sampling—poses considerable performance challenges in practice. This survey offers a comprehensive review of hardware acceleration for lattice-based cryptographic schemes—specifically both the architectural and implementation details of the standardized algorithms in the category CRYSTALS-Kyber, CRYSTALS-Dilithium, and FALCON (Fast Fourier Lattice-Based Compact Signatures over NTRU). It examines optimization measures at various levels, such as algorithmic optimization, arithmetic unit design, memory hierarchy management, and system integration. The paper compares the various performance measures (throughput, latency, area, and power) of Field-Programmable Gate Array (FPGA) and Application-Specific Integrated Circuit (ASIC) implementations. We also address major issues related to implementation, side-channel resistance, resource constraints within IoT (Internet of Things) devices, and the trade-offs between performance and security. Finally, we point out new research opportunities and existing challenges, with implications for hardware accelerator design in the post-quantum cryptographic environment. Full article
16 pages, 998 KB  
Article
Architecture Design of a Convolutional Neural Network Accelerator for Heterogeneous Computing Based on a Fused Systolic Array
by Yang Zong, Zhenhao Ma, Jian Ren, Yu Cao, Meng Li and Bin Liu
Sensors 2026, 26(2), 628; https://doi.org/10.3390/s26020628 - 16 Jan 2026
Cited by 2 | Viewed by 1681
Abstract
Convolutional Neural Networks (CNNs) generally suffer from excessive computational overhead, high resource consumption, and complex network structures, which severely restrict the deployment on microprocessor chips. Existing related accelerators only have an energy efficiency ratio of 2.32–6.5925 GOPs/W, making it difficult to meet the [...] Read more.
Convolutional Neural Networks (CNNs) generally suffer from excessive computational overhead, high resource consumption, and complex network structures, which severely restrict the deployment on microprocessor chips. Existing related accelerators only have an energy efficiency ratio of 2.32–6.5925 GOPs/W, making it difficult to meet the low-power requirements of embedded application scenarios. To address these issues, this paper proposes a low-power and high-energy-efficiency CNN accelerator architecture based on a central processing unit (CPU) and an Application-Specific Integrated Circuit (ASIC) heterogeneous computing architecture, adopting an operator-fused systolic array algorithm with the YOLOv5n target detection network as the application benchmark. It integrates a 2D systolic array with Conv-BN fusion technology to achieve deep operator fusion of convolution, batch normalization and activation functions; optimizes the RISC-V core to reduce resource usage; and adopts a locking mechanism and a prefetching strategy for the asynchronous platform to ensure operational stability. Experiments on the Nexys Video development board show that the architecture achieves 20.6 GFLOPs of computational performance, 1.96 W of power consumption, and 10.46 GOPs/W of energy efficiency ratio, which is 58–350% higher than existing mainstream accelerators, thus demonstrating excellent potential for embedded deployment. Full article
(This article belongs to the Section Intelligent Sensors)
►▼ Show Figures

Figure 1

26 pages, 373 KB  
Perspective
Hardware Accelerators for Cardiovascular Signal Processing: A System-on-Chip Perspective
by Rami Hariri, Marcian Cirstea, Mahdi Maktab Dar Oghaz, Khaled Benkrid and Oliver Faust
Micromachines 2026, 17(1), 51; https://doi.org/10.3390/mi17010051 - 30 Dec 2025
Cited by 1 | Viewed by 1806
Abstract
This study presents a comprehensive systematic analysis, investigating hardware accelerators specifically designed for real-time cardiovascular signal processing, focusing mainly on Electrocardiogram (ECG), Photoplethysmogram (PPG), and blood pressure monitoring systems. Cardiovascular Diseases (CVDs) represent the world’s leading cause of morbidity and mortality, creating an [...] Read more.
This study presents a comprehensive systematic analysis, investigating hardware accelerators specifically designed for real-time cardiovascular signal processing, focusing mainly on Electrocardiogram (ECG), Photoplethysmogram (PPG), and blood pressure monitoring systems. Cardiovascular Diseases (CVDs) represent the world’s leading cause of morbidity and mortality, creating an urgent demand for efficient and accurate diagnostic technologies. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, we systematically analysed 59 research papers on this topic, published from 2014 to 2024, categorising them into three main categories: signal denoising, feature extraction, and decision support with Machine Learning (ML) or Deep Learning (DL). A comprehensive performance benchmarking across energy efficiency, processing speed, and clinical accuracy demonstrates that hybrid Field Programmable Gate Array (FPGA)-Application Specific Integrated Circuit (ASIC) architectures and specialised Artificial Intelligence (AI) on Edge accelerators represent the most promising solutions for next-generation CVD monitoring systems. The analysis identifies key technological gaps and proposes future research directions focused on developing ultra-low-power, clinically robust, and highly scalable physiological signal processing systems. The findings provide guidance for advancing hardware-accelerated cardiovascular diagnostics toward practical clinical deployment. Full article
(This article belongs to the Special Issue Advances in Field-Programmable Gate Arrays (FPGAs))
►▼ Show Figures

Figure 1

28 pages, 1661 KB  
Article
Fault Injection Tool for FPGA Reliability Testing: A Novel Method and Discovery of LUT-Specific Logical Redundancies
by Mariusz Węgrzyn, Orest Kochan and Ihor Maikiv
Electronics 2025, 14(23), 4600; https://doi.org/10.3390/electronics14234600 - 24 Nov 2025
Cited by 2 | Viewed by 1956
Abstract
FPGAs are well suited for prototyping complex digital systems for industrial and research purposes, as well as for the practical application of artificial intelligence (AI) methods in industrial autonomous control, automotives and space. FPGAs serve as platforms for inferring based on AI algorithms. [...] Read more.
FPGAs are well suited for prototyping complex digital systems for industrial and research purposes, as well as for the practical application of artificial intelligence (AI) methods in industrial autonomous control, automotives and space. FPGAs serve as platforms for inferring based on AI algorithms. In recent years, an increase in FPGA system applications with respect to advanced computing functions for physical and chemical research in space has been observed. Research on the reliability of applications operating in the above-mentioned areas exposed to radiation is of particular importance. Testing applications implemented on FPGAs requires the development of new methods that differ significantly from those intended for Application-Specific Integrated Circuits (ASICs). The FPGA logic is realized by SRAM-Based Look-Up Tables (LUTs). SRAM is relatively susceptible to single-event upsets (SEUs) generated by cosmic radiation. The existing fault injection (FI) tools do not model the faults generated by SEUs in SRAM-based FPGAs precisely enough. New FI tools are crucial for evaluating newly developed FPGA-specific tests. Thus, we developed a new tool that uses an accurate SEU model in LUTs. This new tool is written in Perl, and its tasks are to inject faults into the structural VHDL description and to control the CADENCE simulator. The novelty of this solution is that the tool models SEUs by modifying the logical functions generated by the LUTs. Furthermore, in this way, stuck-at faults at the LUT inputs and outputs can also be modeled. This method involves modifying the “INIT” parameters in the structural VHDL. Our tool was evaluated using several test programs, and a high fault coverage (FC) of 94.76% was achieved. This tool can be used to examine any LUT-based FPGA technology regardless of its implementation age. Moreover, during our research, a new mechanism of generating so-called logical redundancies caused by the injection of single faults in LUTs was discovered. This is a side effect of FI in LUTs, which makes it impossible to achieve 100% fault coverage of applications implemented on FPGAs. The mechanism of this phenomenon does not occur when injecting traditional stuck-at faults. Full article
►▼ Show Figures

Figure 1

Back to TopTop