FPGA-Based Reconfigurable SoCs for Safety-Critical AI Inference: A Systematic Literature Review
Abstract
1. Introduction
1.1. Research Questions
- RQ1: What design methodologies enable effective reconfigurable parallel SoCs for safety-critical applications?
- RQ2: How do reconfiguration mechanisms, particularly dynamic partial reconfiguration, facilitate runtime adaptability while maintaining verifiability?
- RQ3: What parallel architectures achieve optimal performance–efficiency trade-offs across CNN acceleration, signal processing, and control applications?
- RQ4: What performance metrics and benchmarks demonstrate the advantages of reconfigurable SoCs, and how do they relate to safety certification requirements?
1.2. Contributions
- 1.
- Convergence–Divergence Analysis (CDA) Framework: We present a systematic quantitative framework that maps how different research threads (DPR, NoC, HLS, and CNN acceleration) are evolving toward convergence or diverging into specialization. This framework provides a structured exploratory lens for identifying research trajectory trends and informing hypothesis generation, with confidence levels attached to each convergence classification.
- 2.
- Safety-Critical Gap Analysis: We systematically analyze the surveyed literature against a three-layer functional safety framework (ISO 26262, ISO 21448/SOTIF, and ISO/PAS 8800), identifying that the overwhelming majority of studies ignore safety certification entirely. While recent work has begun establishing WCET bounds for FPGA SoC platforms, none of the surveyed AI accelerator studies provide WCET bounds required for ASIL certification.
- 3.
- Uncertainty Quantification Opportunity: We propose a novel research direction connecting conformal prediction with FPGA-based CNN accelerators for safety-critical inference with calibrated confidence estimates, addressing a gap at the intersection of statistical learning theory and hardware implementation.
- 4.
- Design Space Taxonomy: We organize the surveyed architectures into a four-dimensional taxonomy spanning reconfigurability granularity, parallelism exploitation, design automation level, and safety criticality, providing a structured framework for design decisions and technology selection.
- 5.
- Quantified Research Agenda: We provide a prioritized research agenda with specific problems, proposed approaches, estimated complexity (person-months), and realistic timelines, transforming this survey from passive summary to active research guidance.
- 6.
- Application Case Study: We ground the analysis in a concrete application—driver drowsiness detection—demonstrating how the identified gaps and proposed research directions apply to a real-world safety-critical deployment scenario.
1.3. Paper Organization
2. Background: FPGAs for Safety-Critical Systems
2.1. Reconfigurable Architectures and the Safety Context
2.2. Design Automation and Verification Challenges
2.3. Connection to Driver Monitoring Systems
3. Systematic Review Protocol for FPGA-Based Reconfigurable SoC Literature
3.1. Search Strategy
Search Queries
- “FPGA reconfigurable parallel processing SoC design implementation dynamic partial reconfiguration AMD Xilinx Zynq”;
- “FPGA multiprocessor systems reconfigurable architecture parallel computing embedded systems”;
- “dynamic reconfiguration FPGA SoC parallel processing acceleration high-performance computing”;
- “FPGA-based reconfigurable CNN accelerator parallel processing edge computing”;
- “modular FPGA multiprocessor reconfigurable systems-on-chip heterogeneous computing”.
3.2. Inclusion Criteria
- Published in peer-reviewed venues (conferences, journals, or technical reports);
- Focus on FPGA-based reconfigurable parallel processing SoC designs;
- Include empirical evaluation or architectural contributions;
- Published between 1999 and 2024;
- Available in English.
3.3. Exclusion Criteria
- No FPGA implementation (simulation-only or purely theoretical work without hardware results).
- Focused exclusively on software or CPU/GPU implementations without reconfigurable hardware contribution.
- Insufficient architectural detail for data extraction (e.g., missing resource utilization, clock frequency, or target device).
- Duplicate or overlapping studies reporting identical results under different titles.
- Wrong publication type (editorials, tutorials, or non-peer-reviewed preprints).
- Non-English publications.
3.4. Study Selection
3.4.1. Identification Phase
3.4.2. Screening Phase
3.4.3. Eligibility Phase
3.4.4. Inclusion Phase

| Exclusion Reason | Count |
|---|---|
| No FPGA implementation/simulation only | 31 |
| Software or CPU/GPU only (no reconfigurable hardware) | 24 |
| Insufficient architectural detail | 22 |
| Duplicate or overlapping studies | 13 |
| Wrong publication type | 8 |
| Total excluded | 98 |
3.5. Data Extraction Framework
- Design Approach: Describe the specific FPGA-based design methodology, including tools, languages (e.g., VHDL and Verilog), and high-level synthesis used for the reconfigurable SoC.
- Reconfiguration Mechanism: Detail how reconfiguration is achieved, such as dynamic partial reconfiguration (DPR), runtime adaptability, or modular architectures in the parallel processing system.
- Parallel Processing Architecture: Outline the parallel processing elements, such as multiprocessor cores, accelerator units, or task partitioning, implemented on the FPGA SoC.
- Implementation Details: Extract key implementation aspects like target FPGA device (e.g., AMD Xilinx Zynq or Intel (Altera) Stratix), resource utilization (LUTs, BRAM, or DSP), clock frequency, and area efficiency.
- Performance Evaluation: Summarize performance metrics, including throughput, latency, and speedup, compared to non-reconfigurable systems and any benchmarks used.
- Applications and Challenges: Identify targeted applications (e.g., signal processing and AI acceleration) and discuss challenges addressed or limitations in the reconfigurable parallel SoC design.
- Key Findings: Highlight the main contributions, innovations, and conclusions from the paper regarding FPGA-based reconfigurable parallel processing SoCs.
3.6. Synthesis Approach
- Consistency of findings across multiple studies;
- Number of supporting studies for each theme;
- Quality of empirical validation;
- Relevance to current FPGA architectures and tools.
- Strong: Consistent findings across 5+ studies with high design quality;
- Moderate: Consistent findings across 2–4 studies with reasonable validation;
- Tentative: Limited evidence or inconsistent findings requiring further research.
3.7. Quality Assessment
- Research design clarity: Is the hardware architecture clearly described with implementation details?
- Empirical validation: Does the study include experimental results on physical hardware (not simulation only)?
- Reproducibility: Are sufficient implementation details provided (target device, resource utilization, and clock frequency)?
- Comparison baseline: Is performance evaluated against a meaningful baseline or alternative approach?
- Metric completeness: Are standard metrics reported (throughput, latency, resource utilization, and power)?
- Limitation acknowledgment: Does the study discuss limitations and assumptions?
- Safety relevance: Does the study address or acknowledge safety certification requirements?
3.8. Inter-Rater Reliability
4. Results
4.1. Overview of Study Characteristics
4.2. Thematic Findings
4.2.1. Design Approaches for Reconfigurable Parallel SoCs
4.2.2. Reconfiguration Mechanisms and Runtime Adaptability
4.2.3. Parallel Processing Architectures
4.2.4. Implementation Details and Resource Utilization
4.2.5. Performance Evaluation and Metrics
4.3. Cross-Study Comparative Analysis
4.3.1. Comparison of Design Approaches
4.3.2. Reconfiguration Granularity Analysis
4.4. Assessment of Evidence Quality
5. Thematic Analysis: Design Dimensions
5.1. Reconfigurability and Runtime Adaptation
5.1.1. Reconfiguration Latency and Infrastructure Overhead
5.1.2. Multi-Granularity Reconfiguration
5.1.3. DPR–NoC Convergence
5.1.4. Safety Certification Gap
5.2. Interconnect and Communication Architectures
5.2.1. Resource Costs and Topology Trade-Offs
5.2.2. Routing and Flow Control
5.2.3. Heterogeneous Platform Interconnect
5.2.4. Safety Implications of Shared Interconnect
5.3. Design Automation and the Productivity Frontier
5.3.1. Performance-Quality of Results
5.3.2. The Verification Challenge
5.3.3. Domain-Specific Frameworks
5.3.4. The Safety–Productivity Paradox
5.4. CNN Acceleration for Edge AI
5.4.1. Performance Evolution
5.4.2. Quantization: The Dominant Optimization
| Study | Network | Throughput | Efficiency | Platform |
|---|---|---|---|---|
| Zhang (2015) [20] | Various | 61.6 GOPS | – | Virtex-7 |
| Qiu [43] | VGG-16 | 137.2 GOPS | 8.3 GOPS/W | Zynq |
| Wei [44] | ResNet | 294 GOPS | – | UltraScale+ |
| Yan [39] | Custom | 21.1 GOPS | – | FPGA |
| Zhang (2020) [38] | YOLO | Competitive | Competitive | FPGA |
5.4.3. Platform Positioning
5.4.4. Edge Deployment Constraints
5.4.5. The Safety Certification Gap in CNN Acceleration
6. Performance Synthesis
6.1. Characteristics of Included Studies
6.2. Performance Metrics
| Study | Year | Key Focus | Target Device | Application |
|---|---|---|---|---|
| Patel et al. [19] | 2006 | Scalable multiprocessor | Off-the-shelf FPGAs | Molecular dynamics |
| Zamacola et al. [23] | 2020 | Multi-grained reconfiguration | Xilinx 7 Series | Image processing, NN |
| Gohringer & Becker [21] | 2010 | Runtime-adaptive MPSoC | General FPGAs | HPC |
| Vipin & Fahmy [11] | 2018 | DPR architecture review | Zynq UltraScale+, Stratix | Signal processing, AI |
| Patel et al. [25] | 2011 | CGRA modeling & simulation | Commercial FPGAs | CGRA design exploration |
| Mhadhbi et al. [50] | 2014 | MicroBlaze FPGA configurations | General FPGAs | HW/SW partitioning |
| Dorta et al. [17] | 2009 | MPSoC overview | General FPGAs | Embedded systems |
| Monmasson & Cirstea [18] | 2007 | FPGA design methodology | General FPGAs | Industrial control |
| Li & Hauck [70] | 2010 | Configuration prefetching | General FPGAs | Partial reconfiguration |
| Minhas et al. [51] | 2022 | Task-specific partitioning | General FPGAs | Cloud/edge multi-tasking |
| Boutros et al. [52] | 2022 | RAD co-design | Beyond-FPGA RADs | Datacenter workloads |
| Cozzi et al. [13] | 2009 | Reconfigurable NoC flow | General FPGAs | Embedded systems |
| Podobas et al. [24] | 2020 | CGRA performance survey | General FPGAs | CGRA evaluation |
| Zhang et al. (2020) [38] | 2020 | FPGA CNN accelerator | General FPGA | Object detection |
| Yan et al. [39] | 2022 | Resource-multiplexing CNN | FPGAs | Image classification |
| Irmak et al. [22] | 2021 | DPR for CNN flexibility | Xilinx FPGAs | Image classification |
| Johnson et al. [16] | 2023 | Distributed cluster | Zynq-7020, UltraScale+ MPSoC | Edge deep learning |
| Ma et al. [45] | 2017 | CNN dataflow optimization | Xilinx Zynq | CNN acceleration |
| Guo et al. [46] | 2018 | Complete CNN-to-FPGA flow | Xilinx Zynq | Image classification |
| Nguyen et al. [47] | 2021 | Mixed-precision FPGA design | Xilinx Zynq MPSoC | Object detection |
| Liu et al. [48] | 2017 | Throughput-optimized accelerator | Xilinx VC707 | CNN inference |
| Gong et al. [49] | 2021 | Dynamic/static co-reconfiguration | Xilinx Zynq MPSoC | CNN acceleration |
| Ijaz et al. [69] | 2023 | Dynamically scalable NoC | General FPGAs | Reconfigurable apps |
| Zhang et al. [30] | 2025 | WCET for CNN on FPGA SoC | Xilinx Zynq MPSoC | WCET estimation |
| Sestito et al. [40] | 2025 | TrIM systolic array for CNN | FPGA | CNN acceleration |
| Peccia et al. [41] | 2024 | Gemmini accelerator edge AI | Xilinx ZCU102 | Object detection |
| Rahoof et al. [42] | 2023 | Capsule network acceleration | Xilinx PYNQ-Z1 | Capsule networks |
| Restuccia & Biondi [37] | 2021 | Time-predictable DNN on SoC | Xilinx Zynq UltraScale+ | DNN timing analysis |
| Metric | Value | Reference |
|---|---|---|
| Throughput (YOLO) | Competitive GOPS | [38] |
| Energy Efficiency | Competitive GOPS/W | [38] |
| Speedup vs. Single-Processor | 3× | [19] |
| Multi-Tasking Throughput | 2.8× improvement | [51] |
| Inference Latency | Not reported | [38] |
| DPR Speedup vs. Static | 2–5× | [11] |
| Prediction Latency | 340.7 μs | [39] |
| Intra-Link Bandwidth | 10 Gbps | [52] |
6.3. Design Approach Comparison
| Approach | Studies | Strengths | Limitations |
|---|---|---|---|
| HLS-based | [51] | Rapid development | Lower peak efficiency |
| RTL-based | [19,50] | Maximum efficiency | Longer dev time |
| Hybrid (HLS+RTL) | [16] | Balanced trade-off | Integration complexity |
| CGRAs | [25] | Flexible computation | Area overhead |
| Granularity | Reconfig. Time | Use Case | Flexibility |
|---|---|---|---|
| Fine-grained (LUT-level) | 1–5 ms | Parameter tuning | Highest |
| Medium-grained (overlay) | 5–20 ms | Algorithm switching | Medium |
| Coarse-grained (module) | 20–100 ms | Application swapping | Lowest |
6.4. Evidence Quality Assessment
6.5. Safety-Relevant Performance Gaps
- No WCET Bounds in Accelerator Designs: All the surveyed CNN accelerator studies report average-case performance metrics. None provide worst-case execution time analysis, which is mandatory for ASIL-B and above under ISO 26262. While recent work has demonstrated WCET analysis feasibility for FPGA SoC platforms [29,30,31,37], with Restuccia & Biondi (2021) providing response-time analysis validated against hardware measurements on Zynq UltraScale+, these methods have not yet been adopted by the broader CNN accelerator design community. The reported inference latencies in surveyed CNN accelerators are average-case measurements without variance analysis or upper bounds [38].
- No Timing Determinism Analysis: Streaming architectures (e.g., Qiu et al. [43]) claim predictable latency, but none provide formal proof of timing determinism under all operating conditions, including memory contention and thermal throttling.
- Missing Confidence Calibration: CNN accelerator studies report classification accuracy (94–99%) but never calibration metrics—whether prediction confidence matches empirical correctness. This is essential for safety-critical decisions where low-confidence predictions must trigger failsafe behavior.
- No Fault Injection Results: None of the performance evaluations include fault injection testing to measure system behavior under hardware faults (SEUs in configuration memory, DSP errors, and routing failures).
- Incomparable Metrics: The heterogeneity of reported metrics prevents quantitative synthesis of safety-relevant parameters, such as timing margins, resource headroom, and error detection coverage.
7. Safety-Critical FPGA SoCs and Uncertainty Quantification
7.1. Functional Safety Requirements for FPGA SoCs
7.1.1. ISO 26262 Requirements
- Deterministic Latency: Bounded worst-case execution time (WCET) needs to be guaranteed. The surveyed studies report average-case latencies, and none provide formal WCET bounds meeting the analysis requirements of ISO 26262 Part 6, although recent analytical models demonstrate WCET estimation feasibility for multi-DPU FPGA SoC platforms (see Section 7).
- Hardware Fault Tolerance: Single-point fault metrics require detection mechanisms for configuration errors, transient faults, and hardware failures.
- Configuration Integrity: Dynamic partial reconfiguration must not introduce timing violations or data corruption in static regions performing safety functions.
7.1.2. DPR Challenges for Safety-Critical Systems
- Reconfiguration Timing: While studies report sub-10 ms reconfiguration latencies, safety-critical systems require deterministic guarantees that reconfiguration completes within bounded time regardless of system state.
- State Preservation: Critical system state must be preserved during reconfiguration. No surveyed study addresses formal methods for verifying state consistency.
- Fault Detection: Configuration memory CRC checks detect bit errors but cannot detect functional faults in the reconfigured logic.
- Temporal Separation: Safety-critical and non-safety-critical functions must be temporally isolated to prevent interference.
| Requirement | ASIL-A | ASIL-B | ASIL-C | ASIL-D |
|---|---|---|---|---|
| Random HW fault detection | 90% | 97% | 99% | 99.9% |
| WCET analysis | Recommended | Required | Required | Mandatory |
| DPR verification | Optional | Recommended | Required | Mandatory |
| Configuration integrity | CRC | CRC + ECC | CRC + ECC + monitor | Triple redundancy |
| Spurious actuation control | Monitor | Dual check | Dual + watchdog | Triple + watchdog |
7.2. Safety of the Intended Functionality (ISO 21448)
7.3. AI Safety Properties (ISO/PAS 8800)
- Neural Network Verification: Requirements for ensuring that learned models behave correctly across their operational design domain.
- Confidence Monitoring: Runtime monitoring of prediction confidence during inference—directly connecting to our conformal prediction proposal in Section 7.
- Safety Argumentation for ML Components: Guidance on constructing safety cases for ML components that supplement the traditional ISO 26262 V-model approach.
- Data Quality Requirements: Ensuring training and validation data representatively cover the operational design domain.
7.4. WCET Analysis for FPGA SoCs
- Lang, Kapre & Pellizzoni (2021) [29]: Provides worst-case latency analysis for the Versal NoC network packet switch, establishing tight bounds for on-chip communication latency. Directly relevant to multi-DPU FPGA SoC architectures where NoC contention causes timing variability. Limitation: covers NoC traversal, not end-to-end CNN inference WCET.
- Gu et al. (2014) [31]: Addresses WCET-aware partial control-flow checking for resource-constrained real-time embedded systems. Relevant to DPR timing verification. Limitation: focuses on control-flow checking rather than data-path timing of CNN accelerators.
- Zhang et al. (2025) [30]: Provides WCET estimation specifically for CNN inference on FPGA SoCs with multi-DPU engines using analytical models of DPU scheduling, memory hierarchy access patterns, and inter-DPU communication. This work demonstrates that end-to-end WCET bounds for CNN inference on AMD Xilinx DPU-based platforms are feasible. We have included this study in our expanded corpus.
- Restuccia & Biondi (2021) [37]: Presents time-predictable DNN acceleration on Zynq UltraScale+ FPGA SoC with the Xilinx DPU, proposing the DICTAT custom FPGA module to improve timing predictability and providing response-time analysis validated against hardware measurements. Limitation: focuses on the Xilinx DPU accelerator specifically rather than general FPGA CNN acceleration.
Safety Awareness Classification
| Study | WCET | Fault Tol. | Standard | Safety-Aware |
|---|---|---|---|---|
| Patel et al. [19] | – | – | – | No |
| Zamacola et al. [23] | – | – | – | No |
| Gohringer & Becker [21] | – | – | – | No |
| Vipin & Fahmy [11] | – | – | – | No |
| Patel et al. [25] | – | – | – | No |
| Mhadhbi et al. [50] | – | – | – | No |
| Dorta et al. [17] | – | – | – | No |
| Monmasson & Cirstea [18] | – | – | – | No |
| Li & Hauck [70] | – | – | – | No |
| Minhas et al. [51] | – | – | – | No |
| Boutros et al. [52] | – | – | – | No |
| Cozzi et al. [13] | – | – | – | No |
| Podobas et al. [24] | – | – | – | No |
| Zhang et al. (2020) [38] | – | – | – | No |
| Yan et al. [39] | – | – | – | No |
| Irmak et al. [22] | – | – | – | No |
| Johnson et al. [16] | – | – | – | No |
| Zhang et al. [20] | – | – | – | No |
| Qiu et al. [43] | – | – | – | No |
| Wei et al. [44] | – | – | – | No |
| Serrano et al. [26] | Partial | – | Mentioned | Yes |
| Wirthlin & Hutchings [4] | – | – | Mentioned | Partial |
| Hildebrandt & Timmermann [57] | – | – | – | No |
| Koch et al. [55] | – | – | – | No |
| Stensgaard [56] | – | – | – | No |
| Ma et al. [45] | – | – | – | No |
| Guo et al. [46] | – | – | – | No |
| Nguyen et al. [47] | – | – | – | No |
| Liu et al. [48] | – | – | – | No |
| Gong et al. [49] | – | – | – | No |
| Ijaz et al. [69] | – | – | – | No |
| Zhang et al. (2025) [30] | Partial | – | – | Yes |
| Sestito et al. [40] | – | – | – | No |
| Peccia et al. [41] | – | – | – | No |
| Rahoof et al. [42] | – | – | – | No |
| Restuccia & Biondi [37] | Partial | – | Mentioned | Yes |
7.5. Uncertainty Quantification in Hardware Accelerators
7.5.1. Motivation for Uncertainty-Aware Inference
- Calibrate Confidence: Prediction confidence should match empirical accuracy. A 90% confidence prediction should be correct 90% of the time.
- Detect Novelty: Inputs significantly different from training data should result in high uncertainty, triggering human intervention or failsafe behavior.
- Provide Guarantees: Uncertainty estimates should be distribution-free with formal coverage guarantees, not merely learned approximations.
7.5.2. Conformal Prediction for FPGA Acceleration
7.6. Research Gap Analysis
7.7. Connection to Driver Drowsiness Detection
- Real-Time Requirement: Detection must complete within 100 ms to enable timely intervention.
- Uncertainty Awareness: Low-confidence predictions should trigger warnings or transition to alternative sensors.
- Calibration: Confidence estimates must be reliable across diverse lighting conditions, driver positions, and partial occlusions.
- FPGA Deployment: Hardware acceleration enables edge deployment without cloud dependency.
- Coverage Guarantee: 90% prediction sets cover the true label with user-specified probability.
- Adaptive Sets: Set size increases for difficult inputs, providing natural uncertainty indication.
- Efficient Computation: Quantile regression is implementable on resource-constrained FPGA platforms.
8. Discussion
8.1. Principal Findings and Interpretation
8.1.1. Convergence on NoC Interconnects
8.1.2. Evolution of Reconfiguration Methods
8.1.3. Performance Mechanisms
8.2. Comparison with Existing Literature
8.2.1. Consistency with Prior Work
8.2.2. Resolution of Apparent Contradictions
- Some studies claim linear scaling with resources, while others note suboptimal utilization in fixed-slot partitioning [51].
- This heterogeneity likely reflects methodological differences—earlier conceptual overviews versus recent empirical benchmarks on larger devices.
- Static evaluations suit prototyping [17], but dynamic workloads expose bottlenecks absent in controlled tests.
8.3. Practical Implications
8.3.1. Edge AI Deployments
8.3.2. Datacenter Settings
8.3.3. Medical Imaging
8.3.4. Regulatory Considerations
8.4. Strengths and Limitations of the Review
8.4.1. Strengths
- Comprehensive search across vast databases yielding diverse temporal coverage (1999–2024).
- Systematic thematic synthesis integrating design, performance, and application insights.
- Transparent extraction of structured data for cross-study comparisons.
- Evidence confidence assessment based on consistency and study quality.
8.4.2. Limitations of Included Studies
- Predominant focus on AMD Xilinx devices, potentially biasing generalizability to other FPGA families.
- Inconsistent quantitative reporting of metrics like exact LUT utilization.
- Conceptual emphases in early works lacking empirical depth.
8.4.3. Review Limitations
- Reliance on abstracts and extracted data without full-text access for all studies.
- Abstract-based screening may miss nuances in methodology.
- Absence of formal risk-of-bias assessment, although thematic analysis mitigates this by prioritizing consistent patterns.
- Cross-platform performance comparisons (FPGA vs. GPU vs. CPU) may reflect unequal optimization effort rather than inherent platform advantages; system-level costs (development time, tool licensing, and memory bottlenecks) are underreported in the surveyed literature and in our synthesis.
8.5. Evidence Summary
9. Convergence–Divergence Analysis Framework
9.1. Framework Definition
- Convergence: When multiple research threads merge into unified approaches.
- Divergence: When research branches into specialized sub-fields.
- Citation Co-Occurrence: How frequently are two techniques cited together in later studies?
- Architecture Unification: Are separate components being integrated into single architectures?
- Tool Integration: Are design tools incorporating multiple techniques?
9.2. Convergence Score Calculation
- Citation Co-Occurrence (CC): The fraction of studies in a given period that cite both techniques together, indicating topical overlap in the community’s attention.
- Architecture Unification (AU): The fraction of studies that integrate both techniques within a single FPGA design, indicating practical convergence in implementations.
- Tool Integration (TI): The fraction of studies reporting use of shared design tools or toolchains that support both techniques, indicating infrastructure convergence.
| Thread Pair | Early Phase | Recent Phase | Conv. Score | Confidence |
|---|---|---|---|---|
| DPR + NoC | 0.15 | 0.70 | 0.72 | High |
| HLS + CNN | 0.10 | 0.85 | 0.91 | High |
| Edge + DPR | 0.20 | 0.60 | 0.68 | Moderate |
| CNN + Quantization | 0.05 | 0.82 | 0.89 | Moderate |
| NoC + Memory | 0.30 | 0.55 | 0.45 | Low |
Sensitivity Analysis
9.3. Convergence Analysis
9.3.1. DPR + NoC Convergence
- 2006–2012: Only 15% of DPR studies also addressed NoC integration;
- 2019–2024: 70% of DPR studies incorporate NoC communication;
- Convergence Score: 0.72 (high convergence).
9.3.2. HLS + CNN Acceleration Convergence
- Pre-2015: CNN accelerators primarily implemented in manual RTL;
- Post-2018: 85% of CNN accelerator papers use HLS;
- Convergence Score: 0.91 (very high convergence).
9.3.3. Edge + DPR Convergence
- Edge Applications with DPR: 60% in recent phase vs. 20% in early phase;
- Edge+DPR+DNN: 45% of edge AI papers address all three;
- Convergence Score: 0.68 (moderate–high convergence).
9.4. Qualitative Scenario Projections
- Convergence to Unified Frameworks (2025–2027): DPR + NoC + HLS will converge into integrated design flows with automated exploration.
- Divergence into Safety-Critical vs. Best-Effort (2026–2028): Clear separation between safety-certified and performance-optimized research tracks.
- Convergence of Uncertainty + Acceleration (2027–2030): Integration of uncertainty quantification with CNN acceleration will become standard.
- Divergence of Edge vs. Cloud (Ongoing): Distinct optimization targets and methodologies for edge and cloud deployments.
9.5. Implications for Researchers
- For new researchers: Enter at convergence points where unified frameworks are emerging; avoid investing in diverging sub-fields unless specific expertise exists.
- For funding agencies: Convergence areas (DPR+NoC+HLS integration and uncertainty quantification) represent high-impact opportunities; divergence areas may require specialized investment.
- For industry: Safety-critical specialization represents underexplored convergence opportunity combining DPR + uncertainty + certification.
- For tool developers: Unified design flows integrating current disparate tools represent significant market opportunity.
10. Design Space Taxonomy and Decision Framework
10.1. Multi-Dimensional Taxonomy
- Reconfigurability Granularity.
- Parallelism Exploitation.
- Design Automation Level.
- Safety Criticality.
10.1.1. Dimension 1: Reconfigurability Granularity
10.1.2. Dimension 2: Parallelism Exploitation
10.1.3. Dimension 3: Design Automation Level
10.1.4. Dimension 4: Safety Criticality
10.2. Decision Framework
Technology Selection Guidelines
- For latency-critical applications:
- Choose spatial parallelism over temporal (DPR has overhead);
- Use streaming architectures to minimize memory access;
- Target II = 1 in HLS for critical loops.
- For power-constrained applications:
- Aggressive quantization (INT8 minimum, consider binary);
- DPR for time-sharing expensive resources;
- Clock gating and power-aware HLS directives.
- For safety-critical applications:
- Uncertainty quantification for confidence estimation;
- WCET analysis for all timing paths;
- Consider manual RTL for critical sections.
- For productivity-focused projects:
- HLS with domain-specific libraries (Vitis AI);
- Template-based design for standard workloads;
- Accept 10–30% efficiency trade-off.
10.3. Design Space Coverage Analysis
11. Research Agenda
11.1. Evidence Gaps
- Non-Xilinx Implementations: Sparse reporting on Intel (Altera) Stratix, Lattice, and Microsemi platforms limits applicability to diverse FPGA ecosystems. The 66% concentration on AMD Xilinx devices creates potential bias in design conclusions and tool compatibility assumptions.
- Inconsistent Metrics: Heterogeneous performance reporting (peak vs. sustained throughput; batch vs. single-sample latency) prevents quantitative cross-study synthesis of safety-relevant parameters.
- Missing Mechanistic Details: Physical-level analysis of routing delays, power pathways during DPR, and configuration memory fault propagation is absent, leaving causal links between architecture and performance unproven.
- Scalability Contradictions: Contradictions in scaling behavior (linear growth vs. fixed-slot limits) remain unresolved due to varying benchmarks, with no replication in safety-critical domains.
11.2. Prioritized Research Gaps
11.3. Research Directions
11.4. Research Roadmap
- Phase 1 (Year 1): Foundation—Benchmark standardization proposal; uncertainty-aware CNN implementation; survey of formal DPR methods; SOTIF triggering event analysis for FPGA-based AI.
- Phase 2 (Year 2): Integration—Unified design flow prototype; WCET analysis for DPR extending Zhang et al. [30]; cross-domain transfer study; ISO/PAS 8800 confidence monitoring integration.
- Phase 3 (Year 3): Validation—Safety-certified prototype; benchmark suite validation; real-world deployment case study.
11.5. Opportunities for Researchers
- For practitioners: Benchmark standardization offers immediate contribution with moderate effort.
- For systems researchers: A unified design flow addresses critical productivity barriers.
- For safety researchers: DPR safety certification is a high-impact, high-effort opportunity.
- For ML/HW researchers: Uncertainty-aware acceleration connects expertise to safety-critical applications.
- For PhD students: The drowsiness detection application provides a concrete case study for validating techniques.
12. Conclusions
12.1. Key Findings
- Safety certification is universally neglected: 33 out of the 36 (92%) surveyed studies ignore safety certification requirements entirely (ISO 26262 and IEC 61508), focusing solely on average-case performance metrics without worst-case execution time analysis.
- Research threads are converging: The convergence–divergence analysis (Section 9) reveals convergence scores of 0.72–0.91 for the DPR, NoC, and HLS threads, indicating increasing research coherence.
- No ASIL-compliant implementations exist: Despite ISO 26262 mandating deterministic timing for ASIL-B and above, no surveyed FPGA implementation provides the formal verification or timing guarantees required for certification.
- Conformal prediction bridges the gap: CP offers a distribution-free framework for uncertainty quantification with finite-sample coverage guarantees that is computationally compatible with FPGA implementation, representing a promising research direction for certifiable AI inference.
12.2. Confidence Assessment
- Strong confidence: DPR effectiveness for latency reduction (<10 ms) and parallel speedup (3×) is supported by eight or more convergent implementations.
- Moderate confidence: CNN acceleration performance metrics show consistent directions but heterogeneous measurement methodologies.
- Tentative: Scalability claims and universal resource metrics require further validation due to limited replication across device generations.
12.3. Critical Uncertainty
12.4. Threats to Validity
- Internal validity: The safety-awareness classification (Table 13) relies on reported content in published papers; studies may have addressed safety considerations without explicitly reporting them. The binary safety-aware classification may undercount studies with partial or informal safety considerations.
- Selection bias: The inclusion criterion requiring FPGA hardware implementation excludes simulation-based safety research and formal verification studies that address safety certification without physical deployment. This scope decision focuses the review on implementation practices but may underestimate the broader community’s engagement with safety, particularly in the formal methods literature. Readers should interpret the 92% safety gap as specific to hardware-implemented FPGA accelerator studies, not the full spectrum of safety-related FPGA research.
- Sample size: With 36 included studies, quantitative generalizations are limited. The CDA convergence scores should be interpreted as indicative rather than statistically significant. Confidence intervals for the 92% safety gap estimate, assuming a binomial distribution, yield a 95% confidence interval of [77%, 98%], indicating that the finding is robust despite the small sample.
- Construct validity: The CDA framework uses heuristic metrics (CC, AU, and TI) with equal weighting. Alternative weighting schemes or different temporal splits may yield different convergence classifications, although our sensitivity analysis (Section 9) suggests stability under reasonable parameter variations.
- Evaluation criteria alignment: Assessing exploratory academic prototypes against ASIL-D certification requirements may appear misaligned as most studies were not designed with certification as a goal. We adopt this lens intentionally as a forward-looking gap analysis, not as a criticism of individual studies.
- Database coverage: The formal search was limited to Scopus, IEEE Xplore, Web of Science, and ACM Digital Library. Relevant studies published in venue-specific proceedings, institutional repositories, or non-English languages may be underrepresented.
- Publication bias: Studies with positive results are more likely to be published, potentially inflating the reported performance metrics (GOPS and GOPS/W) relative to typical implementations. Studies that failed to achieve safety certification may be underrepresented due to the “file drawer” effect.
- Temporal validity: The search was conducted in early 2025. Rapid developments in FPGA-based AI acceleration, particularly from major vendors (AMD/Xilinx and Intel/Altera), may shift the landscape significantly by the time of publication.
12.5. Key Takeaway
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| ASIL | Automotive Safety Integrity Level |
| CDA | Convergence–Divergence Analysis |
| CNN | Convolutional Neural Network |
| CP | Conformal Prediction |
| DPR | Dynamic Partial Reconfiguration |
| FPGA | Field-Programmable Gate Array |
| HLS | High-Level Synthesis |
| IEC | International Electrotechnical Commission |
| ISO | International Organization for Standardization |
| NoC | Network-on-Chip |
| SoC | System-on-Chip |
| WCET | Worst-Case Execution Time |
Appendix A. List of Included Studies
- 1.
- Boutros et al. (2022)—RAD co-design for datacenter workloads [52]
- 2.
- Cozzi et al. (2009)—Reconfigurable NoC flow for embedded systems [13]
- 3.
- Dorta et al. (2009)—MPSoC overview of general FPGAs [17]
- 4.
- Ferreira et al. (2011)—Virtual CGRA on commercial FPGAs [25]
- 5.
- Gohringer & Becker (2010)—Runtime-adaptive MPSoC for HPC [21]
- 6.
- Wirthlin & Hutchings (1998)—Early DPR methodology [4]
- 7.
- Irmak et al. (2021)—DPR for CNN flexibility on AMD Xilinx FPGAs [22]
- 8.
- Johnson et al. (2023)—Distributed cluster for edge deep learning [16]
- 9.
- Kalte et al. (2002)—Early DPR on Virtex-II [57]
- 10.
- Koch et al. (2013)—Recobus-X column-based reconfiguration [55]
- 11.
- Le Beux et al. (2007)—Iterative refactoring for parallel applications [18]
- 12.
- Minhas et al. (2022)—Task-specific partitioning for cloud/edge multi-tasking [51]
- 13.
- Muralikrishna et al. (2014)—MicroBlaze-based SoC on Spartan-3E [50]
- 14.
- Panel et al. (2010)—Multi-FPGA block reuse for supercomputing [70]
- 15.
- Patel et al. (2006)—Scalable multiprocessor for molecular dynamics [19]
- 16.
- Qiu et al. (2016)—Streaming CNN inference on Zynq [43]
- 17.
- Serrano et al. (2021)—Safety-aware DPR on Zynq MPSoC [26]
- 18.
- Stensgaard (2008)—ReNoC runtime topology switching [56]
- 19.
- Syed et al. (2023)—Multi-task CNN accelerator for multimodal AI [24]
- 20.
- Vipin & Fahmy (2018)—DPR architecture review [11]
- 21.
- Zhang et al. (2020)—FPGA-based CNN accelerator for YOLO [38]
- 22.
- Wei et al. (2017)—Aggressive parallelization on UltraScale+ [44]
- 23.
- Yan et al. (2022)—Resource-multiplexed CNN for image classification [39]
- 24.
- Zamacola et al. (2020)—Multi-grained reconfiguration on AMD Xilinx 7 Series [23]
- 25.
- Zhang et al. (2015)—Systolic array CNN on Virtex-7 [20]
- 26.
- Ma et al. (2017)—FPGA loop optimization for CNN dataflow [45]
- 27.
- Guo et al. (2018)—Angel-Eye CNN-to-FPGA design flow [46]
- 28.
- Nguyen et al. (2021)—Mixed-precision FPGA for CNN object detectors [47]
- 29.
- Liu et al. (2017)—Throughput-optimized FPGA accelerator [48]
- 30.
- Gong et al. (2021)—Dynamic/static co-reconfiguration for CNN [49]
- 31.
- Ijaz et al. (2023)—Dynamically scalable NoC [69]
- 32.
- Zhang et al. (2025)—WCET estimation for CNN on FPGA SoC [30]
- 33.
- Sestito et al. (2025)—TrIM systolic array for CNN acceleration [40]
- 34.
- Peccia et al. (2024)—Gemmini accelerator for edge AI on ZCU102 [41]
- 35.
- Rahoof et al. (2023)—Capsule network acceleration on FPGA [42]
- 36.
- Restuccia & Biondi (2021)—Time-predictable DNN on FPGA SoC [37]
References
- ISO 26262; Road Vehicles—Functional Safety. International Organization for Standardization: Geneva, Switzerland, 2018.
- ISO 21448; Road Vehicles—Safety of the Intended Functionality. International Organization for Standardization: Geneva, Switzerland, 2022.
- ISO/PAS 8800; Safety Properties of Artificial Intelligence in Road Vehicles. International Organization for Standardization: Geneva, Switzerland, 2024.
- Wirthlin, M.J.; Hutchings, B.L. Improving Functional Density Using Run-Time Circuit Reconfiguration [FPGAs]. IEEE Trans. Very Large Scale Integr. (VLSI) Syst. 1998, 6, 247–256. [Google Scholar] [CrossRef] [Scilit]
- Compton, K.; Hauck, S. Reconfigurable Computing: A Survey of Systems and Software. ACM Comput. Surv. 2002, 34, 171–210. [Google Scholar] [CrossRef] [Scilit]
- Kuon, I.; Tessier, R.; Rose, J. FPGA Architecture: Survey and Challenges. Found. Trends Electron. Des. Autom. 2008, 2, 135–253. [Google Scholar] [CrossRef] [Scilit]
- Dong, Y.; Hu, Z.; Uchimura, K.; Murayama, N. Driver Inattention Monitoring System for Intelligent Vehicles: A Review. IEEE Trans. Intell. Transp. Syst. 2011, 12, 596–614. [Google Scholar] [CrossRef] [Scilit]
- Ramzan, M.; Khan, H.U.; Awan, S.M.; Ismail, A.; Ilyas, M.; Mahmood, A. A Survey on State-of-the-Art Drowsiness Detection Techniques. IEEE Access 2019, 7, 61904–61919. [Google Scholar] [CrossRef] [Scilit]
- Vovk, V.; Gammerman, A.; Shafer, G. Algorithmic Learning in a Random World; Springer Science & Business Media: New York, NY, USA, 2005. [Google Scholar] [CrossRef] [Scilit]
- Shafer, G.; Vovk, V. A Tutorial on Conformal Prediction. J. Mach. Learn. Res. 2008, 9, 371–421. [Google Scholar]
- Vipin, K.; Fahmy, S.A. FPGA Dynamic and Partial Reconfiguration: A Survey of Architectures, Methods, and Applications. ACM Comput. Surv. 2018, 51, 72. [Google Scholar] [CrossRef] [Scilit]
- Capra, M.; Bussolino, B.; Marchisio, A.; Masera, G.; Martina, M.; Shafique, M. Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead. IEEE Access 2020, 8, 225134–225180. [Google Scholar] [CrossRef] [Scilit]
- Cozzi, D.; Farè, C.; Meroni, A.; Rana, V.; Santambrogio, M.D.; Sciuto, D. Reconfigurable NoC Design Flow for Multiple Applications Run-Time Mapping on FPGA Devices. In Proceedings of the ACM Great Lakes Symposium on VLSI (GLSVLSI); ACM: New York, NY, USA, 2009; pp. 421–424. [Google Scholar]
- Nechi, A.; Groth, L.; Mulhem, S.; Merchant, F.; Buchty, R.; Berekovic, M. FPGA-Based Deep Learning Inference Accelerators: Where Are We Standing? ACM Trans. Reconfigurable Technol. Syst. 2023, 16, 60. [Google Scholar] [CrossRef] [Scilit]
- Münch, D.; Paulitsch, M.; Honold, M.; Schlecker, W.; Herkersdorf, A. Iterative FPGA Implementation Easing Safety Certification for Mixed-Criticality Embedded Real-Time Systems. In Proceedings of the 2014 Euromicro Conference on Digital System Design (DSD), Verona, Italy, 27–29 August 2014; pp. 303–311. [Google Scholar] [CrossRef] [Scilit]
- Johnson, H.; Fang, T.; Perez-Vicente, A.; Saniie, J. Reconfigurable Distributed FPGA Cluster Design for Deep Learning Accelerators. In Proceedings of the IEEE International Conference on Electro Information Technology (EIT); IEEE: New York, NY, USA, 2023. [Google Scholar]
- Dorta, T.; Jiménez, J.; Martín, J.L.; Bidarte, U.; Astarloa, A. Overview of FPGA-Based Multiprocessor Systems. In Proceedings of the International Conference on Reconfigurable Computing and FPGAs (ReConFig); IEEE: New York, NY, USA, 2009; pp. 273–278. [Google Scholar] [CrossRef] [Scilit]
- Monmasson, E.; Cirstea, M.N. FPGA Design Methodology for Industrial Control Systems—A Review. IEEE Trans. Ind. Electron. 2007, 54, 1824–1842. [Google Scholar] [CrossRef] [Scilit]
- Patel, A.; Madill, C.; Saldana, M.; Comis, C.; Pomes, R.; Chow, P. A Scalable FPGA-based Multiprocessor. In Proceedings of the IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM), Napa Valley, CA, USA, 24–26 April 2006; pp. 111–120. [Google Scholar]
- Zhang, C.; Li, P.; Sun, G.; Guan, Y.; Xiao, B.; Cong, J. Optimizing FPGA-Based Accelerator Design for Deep Convolutional Neural Networks. In Proceedings of the ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA); ACM: New York, NY, USA, 2015; pp. 161–170. [Google Scholar]
- Göhringer, D.; Becker, J. New Dimensions in Design Space and Runtime Adaptivity for Multiprocessor Systems Through Dynamic and Partial Reconfiguration: The RAMPSoC Approach. In Proceedings of the IEEE Computer Society Annual Symposium on VLSI (ISVLSI); Selected Papers, Lecture Notes in Electrical Engineering; Springer: Dordrecht, The Netherlands, 2010; Volume 105, pp. 335–346. [Google Scholar]
- Irmak, H.; Ziener, D.; Alachiotis, N. Increasing Flexibility of FPGA-Based CNN Accelerators with Dynamic Partial Reconfiguration. In Proceedings of the International Conference on Field-Programmable Logic and Applications (FPL); IEEE: New York, NY, USA, 2021; pp. 306–311. [Google Scholar] [CrossRef] [Scilit]
- Zamacola, R.; Otero, A.; García Ortiz, A.; de la Torre, E. An Integrated Approach and Tool Support for the Design of FPGA-Based Multi-Grain Reconfigurable Systems. IEEE Access 2020, 8, 202133–202152. [Google Scholar] [CrossRef] [Scilit]
- Podobas, A.; Sano, K.; Matsuoka, S. A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective. IEEE Access 2020, 8, 123695–123717. [Google Scholar] [CrossRef] [Scilit]
- Patel, K.; McGettrick, S.; Bleakley, C.J. Rapid Functional Modelling and Simulation of Coarse Grained Reconfigurable Array Architectures. J. Syst. Archit. 2011, 57, 383–391. [Google Scholar] [CrossRef] [Scilit]
- Serrano-Cases, A.; Reina, J.M.; Abella, J. Leveraging Hardware QoS to Control Contention in the Xilinx Zynq UltraScale+ MPSoC. In Proceedings of the Euromicro Conference on Real-Time Systems (ECRTS); Schloss Dagstuhl: Wadern, Germany, 2021; pp. 3:1–3:24. [Google Scholar]
- Nane, R.; Sima, V.M.; Pilato, C.; Choi, J.; Fort, B.; Canis, A.; Chen, Y.T.; Hsiao, H.; Brown, S.; Ferrandi, F.; et al. A Survey and Evaluation of FPGA High-Level Synthesis Tools. IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst. 2016, 35, 1591–1604. [Google Scholar] [CrossRef] [Scilit]
- Cong, J.; Liu, B.; Neuendorffer, S.; Noguera, J.; Vissers, K.; Zhang, Z. High-Level Synthesis for FPGAs: From Prototyping to Deployment. IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst. 2011, 30, 473–491. [Google Scholar] [CrossRef] [Scilit]
- Lang, I.; Kapre, N.; Pellizzoni, R. Worst-Case Latency Analysis for the Versal NoC Network Packet Switch. In Proceedings of the IEEE/ACM International Symposium on Networks-on-Chip (NOCS), Madison, WI, USA, 14–15 October 2021; pp. 55–60. [Google Scholar]
- Zhang, W.; Yu, Y.; Jiang, X.; Guan, N.; Zhan, N.; Ju, L. WCET Estimation for CNN Inference on FPGA SoC With Multi-DPU Engines. IEEE Trans. Parallel Distrib. Syst. 2025, 36, 1146–1160. [Google Scholar] [CrossRef] [Scilit]
- Gu, Z.; Wang, C.; Zhang, M.; Wu, Z. WCET-Aware Partial Control-Flow Checking for Resource-Constrained Real-Time Embedded Systems. IEEE Trans. Ind. Electron. 2014, 61, 5652–5661. [Google Scholar] [CrossRef] [Scilit]
- Gannous, A.; Andrews, A.; Gallina, B. Bridging the Gap Between Testing and Safety Certification. In Proceedings of the 2018 IEEE Aerospace Conference; IEEE: New York, NY, USA, 2018; pp. 1–18. [Google Scholar] [CrossRef] [Scilit]
- Iturbe, X.; Ebrahim, A.; Benkrid, K.; Hong, C.; Arslan, T.; Perez, J.; Keymeulen, D.; Santambrogio, M.D. R3TOS-Based Autonomous Fault-Tolerant Systems. IEEE Micro 2014, 34, 20–30. [Google Scholar] [CrossRef] [Scilit]
- Elderhalli, Y.; El-Araby, N.; Hasan, O.; Jantsch, A.; Tahar, S. Dynamic Fault Tree Models for FPGA Fault Tolerance and Reliability. In Proceedings of the 2021 IEEE Computer Society Annual Symposium on VLSI (ISVLSI); IEEE: New York, NY, USA, 2021; pp. 194–199. [Google Scholar] [CrossRef] [Scilit]
- Deng, W.; Wu, R. Real-Time Driver-Drowsiness Detection System Using Facial Features. IEEE Access 2019, 7, 118727–118738. [Google Scholar] [CrossRef] [Scilit]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Restuccia, F.; Biondi, A. Time-Predictable Acceleration of Deep Neural Networks on FPGA SoC Platforms. In Proceedings of the 42nd IEEE Real-Time Systems Symposium (RTSS); IEEE: New York, NY, USA, 2021; pp. 441–454. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Cao, J.; Zhang, Q.; Zhang, Q.; Zhang, Y.; Wang, Y. An FPGA-Based Reconfigurable CNN Accelerator for YOLO. In Proceedings of the 2020 IEEE 3rd International Conference on Electronics Technology (ICET), Chengdu, China, 8–12 May 2020; pp. 74–78. [Google Scholar] [CrossRef] [Scilit]
- Yan, F.; Zhang, Z.; Liu, Y.; Liu, J. Design of Convolutional Neural Network Processor Based on FPGA Resource Multiplexing Architecture. Sensors 2022, 22, 5967. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sestito, C.; Agwa, S.; Prodromakis, T. TrIM, Triangular Input Movement Systolic Array for Convolutional Neural Networks: Architecture and Hardware Implementation. IEEE Trans. Circuits Syst. I Regul. Pap. 2025, 72, 2263–2273. [Google Scholar] [CrossRef] [Scilit]
- Peccia, F.N.; Pavlitska, S.; Fleck, T.; Bringmann, O. Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator. In Proceedings of the 27th Euromicro Conference on Digital System Design (DSD); IEEE: New York, NY, USA, 2024; pp. 418–426. [Google Scholar] [CrossRef] [Scilit]
- Rahoof, A.; Chaturvedi, S.; Shafique, M. FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays. In Proceedings of the International Joint Conference on Neural Networks (IJCNN); IEEE: New York, NY, USA, 2023; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
- Qiu, J.; Wang, J.; Yao, S.; Guo, K.; Li, B.; Zhou, E.; Yu, J.; Tang, T.; Xu, N.; Song, S.; et al. Going Deeper with Embedded FPGA Platform for Convolutional Neural Network. In Proceedings of the ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA); ACM: New York, NY, USA, 2016; pp. 26–35. [Google Scholar]
- Wei, X.; Yu, C.H.; Zhang, P.; Chen, Y.; Wang, Y.; Hu, H.; Liang, Y.; Cong, J. Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAs. In Proceedings of the ACM/IEEE Design Automation Conference (DAC); IEEE: New York, NY, USA, 2017; pp. 1–6. [Google Scholar]
- Ma, Y.; Cao, Y.; Vrudhula, S.; Seo, J.s. Optimizing Loop Operation and Dataflow in FPGA Acceleration of Deep Convolutional Neural Networks. In Proceedings of the ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA); ACM: New York, NY, USA, 2017; pp. 45–54. [Google Scholar]
- Guo, K.; Sui, L.; Qiu, J.; Yu, J.; Wang, J.; Yao, S.; Han, S.; Wang, Y.; Yang, H. Angel-Eye: A Complete Design Flow for Mapping CNN onto Embedded FPGA. IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst. 2018, 37, 35–47. [Google Scholar] [CrossRef] [Scilit]
- Nguyen, D.T.; Kim, H.; Lee, H.J. Layer-Specific Optimization for Mixed Data Flow With Mixed Precision in FPGA Design for CNN-Based Object Detectors. IEEE Trans. Circuits Syst. Video Technol. 2021, 31, 2450–2464. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Dou, Y.; Jiang, J.; Xu, J.; Li, S.; Zhou, Y.; Xu, Y. Throughput-Optimized FPGA Accelerator for Deep Convolutional Neural Networks. ACM Trans. Reconfigurable Technol. Syst. 2017, 10, 17. [Google Scholar] [CrossRef] [Scilit]
- Gong, L.; Wang, C.; Li, X.; Zhou, X. Improving HW/SW Adaptability for Accelerating CNNs on FPGAs Through A Dynamic/Static Co-Reconfiguration Approach. IEEE Trans. Parallel Distrib. Syst. 2021, 32, 1854–1865. [Google Scholar] [CrossRef] [Scilit]
- Mhadhbi, I.; Litayem, N.; Ben Othman, S.; Ben Saoud, S. Impact of Hardware/Software Partitioning and MicroBlaze FPGA Configurations on the Embedded Systems Performances. In Complex System Modelling and Control Through Intelligent Soft Computations; Springer International Publishing: Cham, Switzerland, 2015; pp. 711–744. [Google Scholar] [CrossRef] [Scilit]
- Minhas, U.I.; Woods, R.; Nikolopoulos, D.S.; Karakonstantis, G. Efficient, Dynamic Multi-Task Execution on FPGA-Based Computing Systems. IEEE Trans. Parallel Distrib. Syst. 2022, 33, 710–722. [Google Scholar] [CrossRef] [Scilit]
- Boutros, A.; Nurvitadhi, E.; Betz, V. Architecture and Application Co-Design for Beyond-FPGA Reconfigurable Acceleration Devices. IEEE Access 2022, 10, 95067–95082. [Google Scholar] [CrossRef] [Scilit]
- Xilinx Inc. Vivado Design Suite User Guide: Dynamic Function Exchange (UG909); Xilinx Inc.: San Jose, CA, USA, 2023. [Google Scholar]
- Lysaght, P.; Blodget, B.; Mason, J.; Young, J.; Bridgford, B. Invited Paper: Enhanced Architectures, Design Methodologies and CAD Tools for Dynamic Reconfiguration of Xilinx FPGAs. In Proceedings of the International Conference on Field Programmable Logic and Applications (FPL); IEEE: New York, NY, USA, 2006; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Koch, D.; Beckhoff, C.; Teich, J. ReCoBus-Builder—A Novel Tool and Technique to Build Statically and Dynamically Reconfigurable Systems for FPGAs. In Proceedings of the International Conference on Field-Programmable Logic and Applications (FPL); IEEE: New York, NY, USA, 2008; pp. 119–124. [Google Scholar] [CrossRef] [Scilit]
- Stensgaard, M.B.; Sparsø, J. ReNoC: A Network-on-Chip Architecture with Reconfigurable Topology. In Proceedings of the IEEE/ACM International Symposium on Networks-on-Chip (NoCS); IEEE: New York, NY, USA, 2008; pp. 55–64. [Google Scholar]
- Hildebrandt, J.; Timmermann, D. An FPGA Based Scheduling Coprocessor for Dynamic Priority Scheduling in Hard Real-Time Systems. In Field-Programmable Logic and Applications: The Roadmap to Reconfigurable Computing (FPL 2000); Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2000; Volume 1896, pp. 777–780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Horta, E.; Lockwood, J.W.; Parlour, D. Dynamic Hardware Plugins in an FPGA with Partial Run-Time Reconfiguration. In Proceedings of the ACM/IEEE Design Automation Conference (DAC); IEEE: New York, NY, USA, 2002; pp. 343–348. [Google Scholar] [CrossRef] [Scilit]
- Benini, L.; De Micheli, G. Networks on Chips: A New SoC Paradigm. Computer 2002, 35, 70–78. [Google Scholar] [CrossRef] [Scilit]
- Dally, W.J.; Towles, B. Route Packets, Not Wires: On-Chip Interconnection Networks. In Proceedings of the ACM/IEEE Design Automation Conference (DAC); IEEE: New York, NY, USA, 2001; pp. 684–689. [Google Scholar] [CrossRef] [Scilit]
- Molina, R.S.; Gil-Costa, V.; Crespo, M.L.; Ramponi, G. High-Level Synthesis Hardware Design for FPGA-Based Accelerators: Models, Methodologies, and Frameworks. IEEE Access 2022, 10, 90429–90455. [Google Scholar] [CrossRef] [Scilit]
- Del Sozzo, E.; Conficconi, D.; Zeni, A.; Salaris, M.; Sciuto, D.; Santambrogio, M.D. Pushing the Level of Abstraction of Digital System Design: A Survey on How to Program FPGAs. ACM Comput. Surv. 2022, 55, 106. [Google Scholar] [CrossRef] [Scilit]
- Duarte, J.; Han, S.; Harris, P.; Jindariani, S.; Kreinar, E.; Kreis, B.; Ngadiuba, J.; Pierini, M.; Rivera, R.; Tran, N.; et al. Fast Inference of Deep Neural Networks in FPGAs for Particle Physics. J. Instrum. 2018, 13, P07027. [Google Scholar] [CrossRef] [Scilit]
- Han, S.; Mao, H.; Dally, W.J. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. In Proceedings of the International Conference on Learning Representations (ICLR), San Juan, PR, USA, 2–4 May 2016; pp. 1–14. [Google Scholar]
- Courbariaux, M.; Bengio, Y.; David, J.P. BinaryConnect: Training Deep Neural Networks with Binary Weights During Propagations. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); MIT Press: Cambridge, MA, USA, 2015; pp. 3123–3131. [Google Scholar]
- Rastegari, M.; Ordonez, V.; Redmon, J.; Farhadi, A. XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2016; pp. 525–542. [Google Scholar]
- Mittal, S. A Survey of FPGA-Based Accelerators for Convolutional Neural Networks. Neural Comput. Appl. 2020, 32, 1109–1139. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.H.; Krishna, T.; Emer, J.S.; Sze, V. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. IEEE J. Solid-State Circuits 2017, 52, 127–138. [Google Scholar] [CrossRef] [Scilit]
- Ijaz, Q.; Kidane, H.L.; Bourennane, E.B.; Ochoa-Ruiz, G. Dynamically Scalable NoC Architecture for Implementing Run-Time Reconfigurable Applications. Micromachines 2023, 14, 1913. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, Z.; Hauck, S. Configuration Prefetching Techniques for Partial Reconfigurable Coprocessor with Relocation and Defragmentation. In Proceedings of the ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA); ACM: New York, NY, USA, 2002; pp. 187–195. [Google Scholar] [CrossRef] [Scilit]
- Angelopoulos, A.N.; Bates, S. A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. arXiv 2021, arXiv:2107.07511. [Google Scholar]
- Romano, Y.; Patterson, E.; Candès, E. Conformalized Quantile Regression. Adv. Neural Inf. Process. Syst. (NeurIPS) 2019, 32, 3543–3553. [Google Scholar]







| Survey | SoC | DPR | AI/ML | Safety | Studies |
|---|---|---|---|---|---|
| Kuon & Rose [6] | ✓ | – | – | – | N/A |
| Compton & Hauck [5] | ✓ | ✓ | – | – | N/A |
| Nechi et al. [14] | ✓ | ✓ | ✓ | – | N/A |
| Capra et al. [12] | – | – | ✓ | – | N/A |
| This review | ✓ | ✓ | ✓ | ✓ | 36 |
| Platform | Representative Studies | Count | Percentage |
|---|---|---|---|
| AMD Xilinx Virtex family | [19,25] | 5 | 14% |
| AMD Xilinx Zynq family | [16,22,30,37,41,42,43,44,45,46,47,48,49] | 21 | 58% |
| Intel (Altera) FPGA | [11] | 4 | 11% |
| General/unspecified | [17,18,40] | 6 | 17% |
| Dimension | High | Medium | Low |
|---|---|---|---|
| Implementation completeness | Working prototype | Simulation only | Conceptual |
| Performance evaluation | Real measurements | Simulated results | No evaluation |
| Reproducibility | Open source | Sufficient detail | Insufficient detail |
| Comparison baseline | Multiple platforms | Single baseline | No comparison |
| Theme | Confidence | Supporting Studies |
|---|---|---|
| DPR effectiveness | Strong | 8+ |
| Parallel speedup | Strong | 10+ |
| CNN acceleration | Moderate | 6+ |
| Energy efficiency | Moderate | 5+ |
| Scalability | Tentative | 4+ |
| Failure Mode | Detection Mechanism | ASIL Coverage |
|---|---|---|
| Configuration SEU | CRC scrubbing | Partial |
| DPR timing violation | STA per configuration | Not addressed |
| Power supply glitch | Brownout detection | Vendor-specific |
| Clock domain crossing | CDC verification | Tool-dependent |
| Routing congestion | Post-route STA | Design-specific |
| Partial bitstream corruption | HMAC authentication | Research only |
| Gap Area | Current State | Required for Safety | Priority |
|---|---|---|---|
| Timing Analysis | Average case reported | WCET bounds (methods emerging [29,30,37]) | Critical |
| Uncertainty Quantification | Not addressed | Calibrated confidence | Critical |
| DPR Verification | Simulation-based | Formal methods | High |
| Fault Tolerance | Not addressed | Detection/recovery | High |
| Certification Tools | Not available | Automated compliance | Medium |
| SOTIF Compliance | Not addressed | Triggering event analysis | Critical |
| AI Safety (ISO/PAS 8800) | Not addressed | Confidence monitoring | Critical |
| Theme | Key Finding | Direction | Confidence |
|---|---|---|---|
| Design Approaches | Multi-grained tools, design flows | Positive | Moderate |
| Reconfiguration | DPR latency < 10 ms, 90% reduction | Positive | Strong |
| Parallel Arch. | 3× speedup with NoC | Positive | Strong |
| Implementation | 65% LUTs, 40% BRAM at 100 MHz | Positive | Moderate |
| Performance | Competitive throughput & efficiency | Positive | Moderate |
| Applications | Resource-multiplexing, DPR for CNNs | Positive | Strong |
| Level | Description | Example |
|---|---|---|
| Static | No runtime reconfiguration | Fixed CNN accelerator |
| Partial (Coarse) | Module-level swapping | Accelerator substitution |
| Partial (Medium) | Overlay reconfiguration | DSP functional updates |
| Partial (Fine) | LUT-level tuning | Parameter adjustment |
| Full Adaptive | Hierarchical multi-grained | Combined approaches |
| Type | Description | Implementation |
|---|---|---|
| Spatial Only | Fixed parallel resources | Unrolled loops, PEs |
| Temporal Only | Time-multiplexed resources | DPR-based sharing |
| Pipeline | Staged computation | Layer-by-layer streaming |
| Heterogeneous | Mixed PE types | CPU + GPU + FPGA SoC |
| Hybrid | Combined spatial/temporal | Adaptive resource allocation |
| Level | Tools | Trade-Off |
|---|---|---|
| Manual RTL | Verilog/VHDL | Maximum control, high effort |
| HLS-Guided | Vitis HLS, Intel HLS | Balanced productivity/control |
| Auto-ML | Vitis AI, NAS | High automation, limited control |
| Template-Based | DNNweaver, FINN | Domain optimization |
| End-to-End | Full stack | Minimum effort, opaque |
| Level | Requirements | Standards |
|---|---|---|
| Best-Effort | None | None |
| Quality-Assured | Testing coverage | Internal QA |
| Deterministic | WCET bounds | IEC 61508 SIL-2 |
| Safety-Certified | Formal verification | ISO 26262 ASIL-B/C |
| Mission-Critical | Fault tolerance | DO-254 DAL-A |
| Application | Latency | Power | Recommended Architecture | Confidence |
|---|---|---|---|---|
| CNN Inference (Edge) | <50 ms | <5 W | DPR + Quantization + Streaming | High |
| CNN Inference (Cloud) | Batch | >50 W | Spatial + HLS + Full Precision | High |
| Signal Processing | Real-Time | <10 W | NoC + Pipelined + DPR | Medium |
| Scientific HPC | Variable | >100 W | Multi-FPGA + NoC + Manual RTL | Medium |
| Safety-Critical | Deterministic | <10 W | DPR + Uncertainty + Formal | Low |
| Edge AI | Cloud HPC | Signal Proc. | Safety-Critical | |
|---|---|---|---|---|
| Studies | 8 | 5 | 7 | 3 |
| Coverage | High | High | Medium | Low |
| Tools Available | Mature | Mature | Maturing | Emerging |
| Confidence | High | High | Medium | Low |
| Gap | Current Limitation | Proposed Approach | Complexity | Timeline |
|---|---|---|---|---|
| Uncertainty in HW | No confidence estimates | Conformal prediction | Medium | 1–2 years |
| DPR Safety | No ASIL-compliant DPR | Formal verification | High | 2–3 years |
| WCET Bounds | Average case only; methods emerging [29,30] | Extend multi-DPU WCET to DPR | High | 1–2 years |
| SOTIF Compliance | Not addressed | Triggering event + prediction sets | High | 1–2 years |
| ISO/PAS 8800 | Not addressed | Confidence monitoring | Medium | 1–2 years |
| Benchmark Std. | Incomparable metrics | Open suite proposal | Low | 6 months |
| Cross-Domain | Not validated | Transfer learning | Medium | 1–2 years |
| Tool Integration | Fragmented flows | Unified framework | High | 2–3 years |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hussein, Y.M.; Hassan, R.F.; Chisab, R.F. FPGA-Based Reconfigurable SoCs for Safety-Critical AI Inference: A Systematic Literature Review. Electronics 2026, 15, 2695. https://doi.org/10.3390/electronics15122695
Hussein YM, Hassan RF, Chisab RF. FPGA-Based Reconfigurable SoCs for Safety-Critical AI Inference: A Systematic Literature Review. Electronics. 2026; 15(12):2695. https://doi.org/10.3390/electronics15122695
Chicago/Turabian StyleHussein, Yasmeen M., Raaed F. Hassan, and Raad Farhood Chisab. 2026. "FPGA-Based Reconfigurable SoCs for Safety-Critical AI Inference: A Systematic Literature Review" Electronics 15, no. 12: 2695. https://doi.org/10.3390/electronics15122695
APA StyleHussein, Y. M., Hassan, R. F., & Chisab, R. F. (2026). FPGA-Based Reconfigurable SoCs for Safety-Critical AI Inference: A Systematic Literature Review. Electronics, 15(12), 2695. https://doi.org/10.3390/electronics15122695

