AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence
Abstract
1. Introduction
2. Taxonomy of AI-on-Chip Paradigms
2.1. Functional Taxonomy
2.1.1. AI for Chip Design (EDA Automation and Optimization)
- Floorplanning optimization
- Routing congestion prediction
- Thermal and power estimation
- Manufacturing defect prediction
- Yield optimization in advanced packaging technologies
2.1.2. AI on Chips (Inference, Perception, and Control)
- CPU
- Neural Processing Units (NPUs)
- FPGA fabrics
- Digital Signal Processors (DSPs)
- Low-latency inference
- Energy-efficient computation
- Localized data processing
- Reduced cloud dependence
- Neural Network Inference Engine (NNIE)
- CPU/DSP subsystems
- Image signal processors
- Heterogeneous acceleration pipelines
- Smart surveillance
- Industrial monitoring
- Autonomous inspection systems
2.1.3. AI with Chips (Co-Designed Intelligent Systems)
2.2. Structural Taxonomy
2.2.1. Monolithic AI SoCs
2.2.2. NoC-Enabled Multi-Core AI Chips
2.2.3. Chiplet-Based 2.5D/3D AI Systems
2.2.4. Directions of Structural Taxonomy
3. AI Chip Architectures and Hardware Platforms
3.1. GPU and Data-Center AI Chips
3.1.1. Training-Centric Architectures
3.1.2. Energy and Scalability Challenges
3.2. Edge AI Accelerators
3.2.1. NPUs and ASIC Accelerators
3.2.2. SDK-Driven Deployment Ecosystems
3.3. FPGA and Hybrid AI SoCs
3.3.1. ARM-FPGA-AI Integration
3.3.2. Reconfigurability Versus Performance
3.4. Heterogeneous and Chiplet-Based AI Architectures
3.4.1. Chiplet Communication and Memory Hierarchies
3.4.2. Neuromorphic and Mixed-Signal AI Chips
4. Network-on-Chip (NoC) Architectures for AI Workloads
4.1. Communication Patterns in AI Chips
4.1.1. Multicast-Heavy Neural Workloads
4.1.2. Global Buffer-to-Compute-Array Traffic
4.2. Multicast-Optimized NoC Designs
4.2.1. Tree-Based Versus Mesh-Based NoCs
4.2.2. Energy and Area Tradeoffs
4.3. NoC-SoC Co-Integration for AI Systems
4.3.1. Scalability Limits of Bus-Based SoCs
4.3.2. NoC-Enabled Many-Core AI Chips
4.4. NoC-Enabled Edge and IIoT AI Systems
4.4.1. Predictive Maintenance and Fault Diagnosis
4.4.2. Real-Time Industrial AI Platforms
5. AI Model Acceleration and Algorithm–Hardware Co-Design
5.1. Model Compression and Quantization
Pruning, Quantization, and Operator Fusion
5.2. CNN and Object Detection Acceleration on AI Chips
5.2.1. YOLO-Family Optimizations
5.2.2. Heterogeneous Acceleration Strategies
5.3. Algorithm-Architecture Co-Design Methodologies
5.3.1. Hardware-Aware Neural Network Design
5.3.2. Precision and Sparsity Co-Design
6. Clocking, Synchronization, and Timing in AI Chips
6.1. Clock Distribution Challenges
6.1.1. Large Die Sizes and Skew Management
6.1.2. Multi-Clock-Domain AI Systems
6.2. Clock Trees for Heterogeneous AI Chiplets
6.2.1. Energy-Aware Clocking Techniques
6.2.2. Modular and Hierarchical Clock Networks
7. AI for Chip Design and Design Automation
7.1. AI-Assisted Analysis in Chip Design
Timing, IR-Drop, and Congestion Prediction
7.2. AI-Driven Optimization
Bayesian Optimization and Reinforcement Learning
7.3. LLMs and AI Agents for Chip Design
7.3.1. RTL Generation and Design-Space Exploration
7.3.2. Autonomous EDA Workflows
8. Manufacturing, Packaging, and Reliability of AI Chips
8.1. Advanced Packaging for AI Chips
8.1.1. 2.5D and 3D Integration
8.1.2. Thermal and Mechanical Considerations
8.2. Assembly and Reliability Challenges
8.2.1. High-Density Soldering
8.2.2. Yield and Long-Term Reliability
9. Beyond Conventional Silicon: Advanced AI Chip Design Trends
9.1. Photonic AI Chips
9.2. Neuromorphic and Brain-Inspired Computing
9.3. In-Memory and Non-Von-Neumann Architectures
9.4. Multimodal and Reconfigurable AI Chips
10. Economic and Industry Perspectives
10.1. Edge Versus Cloud AI Economics
10.2. Cost, Scalability, and Sustainability Challenges
11. Conclusions and Future Research Directions
11.1. Conclusions
11.2. Future Outlook for AI-on-Chip Systems
Funding
Data Availability Statement
Conflicts of Interest
References
- Jiang, M.; Xu, Y.; Li, Z.; Li, C. Current Opinions on Memristor-Accelerated Machine Learning Hardware. Curr. Opin. Solid State Mater. Sci. 2025, 37, 101226. [Google Scholar] [CrossRef] [Scilit]
- Wang, T.; Guo, J.; Zhang, B.; Yang, G.; Li, D. Deploying AI on Edge: Advancement and Challenges in Edge Intelligence. Mathematics 2025, 13, 1878. [Google Scholar] [CrossRef] [Scilit]
- Mohith, V.; Sakthivel, R. A Review on Selective In-Memory Computing Processors: Potential Alternative to AI-Driven Applications. Results Eng. 2026, 29, 108460. [Google Scholar] [CrossRef] [Scilit]
- Bhowmik, B.; Hazarika, P.; Kale, P.; Jain, S. AI Technology for NoC Performance Evaluation. IEEE Trans. Circuits Syst. II Express Briefs 2021, 68, 3483–3487. [Google Scholar] [CrossRef] [Scilit]
- Alja’afreh, M.; Obaidat, M.; Karime, A.; Alouneh, S. Optimizing System-on-Chip Performance Using AI and SDN: Approaches and Challenges. In Proceedings of the 2022 Ninth International Conference on Software Defined Systems (SDS); IEEE: New York, NY, USA, 2022; pp. 1–8. [Google Scholar]
- Chen, H.-C.; Shen, P.-C.; Yen, Y.-C.; Wang, Y.-Y.; Chen, K.-C.J. NoC AI Chip Integration for Industrial IoT Fault Diagnosis and Notification System. In Proceedings of the 2024 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS); IEEE: New York, NY, USA, 2024; p. 1. [Google Scholar]
- Cao, Y.; Chen, Y.; Fan, X.; Fu, H.; Xu, B. Advanced Design for High-Performance and AI Chips. Nanomicro Lett. 2026, 18, 13. [Google Scholar] [CrossRef] [Scilit]
- Boutros, A.; Arora, A.; Betz, V. Field-Programmable Gate Array Architecture for Deep Learning: Survey and Future Directions. Proc. IEEE 2025, 113, 613–639. [Google Scholar] [CrossRef] [Scilit]
- Fang, T.; Perez-Vicente, A.; Johnson, H.; Saniie, J. Deep Learning Scheduling on a Field-Programmable Gate Array Cluster Using Configurable Deep Learning Accelerators. Information 2025, 16, 298. [Google Scholar] [CrossRef] [Scilit]
- Cordova-Cardenas, R.; Amor, D.; Gutiérrez, Á. Edge AI in Practice: A Survey and Deployment Framework for Neural Networks on Embedded Systems. Electronics 2025, 14, 4877. [Google Scholar] [CrossRef] [Scilit]
- Bamberg, L.; Minnella, F.; Bosio, R.; Ottati, F.; Wang, Y.; Lee, J.; Lavagno, L.; Fuks, A. EIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations. arXiv 2025, arXiv:2509.14388. [Google Scholar]
- Xudong, Z.; Meng, Y.; Xinxin, X.; Jianwen, Z.; Changling, W.; Fang, W. Research of YOLOv5s Model Acceleration Strategy in AI Chip. In Proceedings of the 2023 8th International Conference on Computer and Communication Systems (ICCCS); IEEE: New York, NY, USA, 2023; pp. 791–794. [Google Scholar]
- Qian, W.; Zhu, Z.; Zhu, C.; Zhu, Y. FPGA-Based Accelerator for YOLOv5 Object Detection with Optimized Computation and Data Access for Edge Deployment. Parallel Comput. 2025, 124, 103138. [Google Scholar] [CrossRef] [Scilit]
- Lamichhane, B.R.; Srijuntongsiri, G.; Horanont, T. CNN Based 2D Object Detection Techniques: A Review. Front. Comput. Sci. 2025, 7, 1437664. [Google Scholar] [CrossRef] [Scilit]
- Carrillo, S.; Harkin, J.; McDaid, L.J.; Morgan, F.; Pande, S.; Cawley, S.; McGinley, B. Scalable Hierarchical Network-on-Chip Architecture for Spiking Neural Network Hardware Implementations. IEEE Trans. Parallel Distrib. Syst. 2013, 24, 2451–2461. [Google Scholar] [CrossRef] [Scilit]
- Soliman, K.; Li, C.; Shi, F. Reactive Deadlock Avoidance Based on Focus Routing Graph Classification for Triplet-Based Architecture Network-on-Chip. IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst. 2026, 45, 190–203. [Google Scholar] [CrossRef] [Scilit]
- Figliolia, T.; Andreou, A.G. The Conical-Fishbone Clock Tree: A Clock-Distribution Network for a Heterogeneous Chip Multiprocessor AI Chiplet. In Proceedings of the 2019 22nd Euromicro Conference on Digital System Design (DSD); IEEE: New York, NY, USA, 2019; pp. 160–165. [Google Scholar]
- Martins, R.M.F. A Survey of Machine and Deep Learning Techniques in Analog Integrated Circuit Layout Synthesis. Microelectronics 2025, 1, 2. [Google Scholar] [CrossRef] [Scilit]
- Biscontini, A.; Popovici, E.; Temko, A. Machine Learning for FPGA Electronic Design Automation. IEEE Access 2024, 12, 182640–182662. [Google Scholar] [CrossRef] [Scilit]
- Huang, G.; Hu, J.; He, Y.; Liu, J.; Ma, M.; Shen, Z.; Wu, J.; Xu, Y.; Zhang, H.; Zhong, K.; et al. Machine Learning for Electronic Design Automation: A Survey. ACM Trans. Des. Autom. Electron. Syst. 2021, 26, 1–46. [Google Scholar] [CrossRef] [Scilit]
- Hollstein, K.; Weide-Zaage, K. Advances in Packaging for Emerging Technologies. In Proceedings of the 2020 Pan Pacific Microelectronics Symposium (Pan Pacific); IEEE: New York, NY, USA, 2020; pp. 1–11. [Google Scholar]
- Wang, G.; Che, J.; Gao, C.; Han, Z.; Shen, J.; Cheng, Z.; Zhou, P. Integrated Neuromorphic Photonic Computing for AI Acceleration: Emerging Devices, Network Architectures, and Future Paradigms. Adv. Mater. 2025, e08029. [Google Scholar] [CrossRef] [Scilit]
- Cao, T.; Shen, C.; Wang, D.; Zhao, H.; Tan, R.; Zhou, Z.; Li, H.; Li, B.; Zhao, M.; Huang, H.-W. Emerging Memory Devices for Neuromorphic Computing in the Internet of Medical Things. Cell Rep. Phys. Sci. 2025, 6, 102735. [Google Scholar] [CrossRef] [Scilit]
- Raghuwanshi, P. Effects of Artificial Intelligence on Semiconductor Manufacturing: AI-Driven Innovations in Chip Fabrication and Electronic Design Automation. IEEE Electron. Devices Mag. 2025, 3, 15–17. [Google Scholar] [CrossRef] [Scilit]
- Masola, A.; Capodieci, N. Optimization Strategies for GPUs: An Overview of Architectural Approaches. Int. J. Parallel Emergent Distrib. Syst. 2023, 38, 140–154. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.R.; Liu, L.; Liu, S.; Liew, B.; Guan, D.; Chen, J.; Jones, S.; Dally, W.J. Co-Optimization of GPU AI Chip from Technology, Design, System and Algorithms. In Proceedings of the 2024 IEEE International Electron Devices Meeting (IEDM); IEEE: New York, NY, USA, 2024; pp. 1–4. [Google Scholar]
- Liu, Y.; Li, X.; Yin, S. Review of Chiplet-Based Design: System Architecture and Interconnection. Sci. China Inf. Sci. 2024, 67, 200401. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Du, Y.; Pu, B.; Yuan, G.; Liu, Y.; Zheng, L.; Wang, P.; Yang, A.; Li, Y.; Yu, C.; et al. Survey of Chiplet Technology: SoC Architecture, Interconnect, EDA, and Advanced Packaging. IEEE J. Emerg. Sel. Top. Circuits Syst. 2025, 15, 514–536. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Pan, S.; Wang, Y.; Huang, Z.; Rao, G.; Peng, W. Design and Implementation of Fully Programmable Heterogeneous FPSoC for Edge AI Applications. In Proceedings of the 2025 4th International Conference on Electronic Information Technology (EIT); IEEE: New York, NY, USA, 2025; pp. 717–721. [Google Scholar]
- Shen, F.-J.; Chen, J.-H.; Wang, W.-Y.; Tsai, D.-L.; Shen, L.-C.; Tseng, C.-T. A CNN-Based Human Head Detection Algorithm Implemented on Edge AI Chip. In Proceedings of the 2020 International Conference on System Science and Engineering (ICSSE); IEEE: New York, NY, USA, 2020; pp. 1–5. [Google Scholar]
- Chen, Y.-H.; Krishna, T.; Emer, J.S.; Sze, V. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. IEEE J. Solid-State Circuits 2017, 52, 127–138. [Google Scholar] [CrossRef] [Scilit]
- Jouppi, N.P.; Young, C.; Patil, N.; Patterson, D.; Agrawal, G.; Bajwa, R.; Bates, S.; Bhatia, S.; Boden, N.; Borchers, A.; et al. In-Datacenter Performance Analysis of a Tensor Processing Unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture; ACM: New York, NY, USA, 2017; pp. 1–12. [Google Scholar]
- Chen, Y.-H.; Emer, J.; Sze, V. Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks. ACM Sigarch Comput. Archit. News 2016, 44, 367–379. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Y.; Yang, H.; Shu, Y.; Jia, Y.; Huang, Z. MTREE: A Customized Multicast-Enabled Tree-Based Network on Chip for AI Chips. IEEE Embed. Syst. Lett. 2022, 14, 143–146. [Google Scholar] [CrossRef] [Scilit]
- Shao, Y.S.; Clemons, J.; Venkatesan, R.; Zimmer, B.; Fojtik, M.; Jiang, N.; Keller, B.; Klinefelter, A.; Pinckney, N.; Raina, P.; et al. Simba: Scaling Deep-Learning Inference with Multi-Chip-Module-Based Architecture. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture; ACM: New York, NY, USA, 2019; pp. 14–27. [Google Scholar]
- Kannan, R.; Gulhane, M.; Mirza, K.; Maurya, S.; Kumar, S.; Rakesh, N. Architectural Innovations for AI Chips, Testing, and High Performance Computing. In Proceedings of the 2024 IEEE 6th International Conference on Cybernetics, Cognition and Machine Learning Applications (ICCCMLA); IEEE: New York, NY, USA, 2024; pp. 365–369. [Google Scholar]
- NVIDIA. Corporation NVIDIA Blackwell Architecture. Available online: https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/ (accessed on 22 May 2026).
- Cerebras Systems Cerebras Systems Unveils World’s Fastest AI Chip with Whopping 4 Trillion Transistors. 2024. Available online: https://www.cerebras.ai/press-release/cerebras-announces-third-generation-wafer-scale-engine (accessed on 1 June 2026).
- Intel Corporation. Architecture Day 2021 Presentation; Intel Corporation: Santa Clara, CA, USA, 2021. [Google Scholar]
- Advanced Micro Devices Inc. AMD Instinct MI300X Accelerator; Advanced Micro Devices Inc.: Santa Clara, CA, USA, 2025. [Google Scholar]
- Advanced Micro Devices Inc. CDNA 3 Architecture; Advanced Micro Devices Inc.: Santa Clara, CA, USA, 2025. [Google Scholar]
- Tan, X.; Jing, L.; Kudriavtsev, V.; Laaksonen, T.; Chen, A.; Lin, H.; Karim, Z.; Chadda, S. Improving the Quality and Yield Performance of Vacuum Fluxless Reflow Soldering for High Density AI Chips. In Proceedings of the 2025 IEEE 75th Electronic Components and Technology Conference (ECTC); IEEE: New York, NY, USA, 2025; pp. 629–633. [Google Scholar]
- NVIDIA. NVIDIA H100 GPU Datasheet; NVIDIA: Santa Clara, CA, USA, 2026. [Google Scholar]
- Andersch, M.; Palmar, G.; Krashinsky, R.; Stam, N.; Mehta, V.; Brito, G.; Ramaswamy, S. NVIDIA Hopper Architecture In-Depth. 2022. Available online: https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/ (accessed on 1 June 2026).
- NVIDIA. NVIDIA NVLink and NVLink Switch. Available online: https://www.nvidia.com/en-us/data-center/nvlink (accessed on 9 March 2026).
- Jouppi, N.; Kurian, G.; Li, S.; Ma, P.; Nagarajan, R.; Nai, L.; Patil, N.; Subramanian, S.; Swing, A.; Towles, B.; et al. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings. In Proceedings of the 50th Annual International Symposium on Computer Architecture; ACM: New York, NY, USA, 2023; pp. 1–14. [Google Scholar]
- AMD. Instinct MI300 Series Accelerators. Available online: https://www.amd.com/en/products/accelerators/instinct/mi300.html (accessed on 9 March 2026).
- Oh, Y.; Jeon, H.; Kim, J.; Han, S.; Chung, Y.; Asanbekov, K.; Song, G.; Han, K.; Kim, J.; Lee, J.; et al. Mobilint’s ARIES: Chip for Edge AI. In Proceedings of the 2023 IEEE International Conference on Consumer Electronics-Asia (ICCE-Asia); IEEE: New York, NY, USA, 2023; pp. 1–4. [Google Scholar]
- Coral Accelerator Module. Available online: https://www.coral.ai/products/accelerator-module/ (accessed on 22 May 2026).
- SEMIFIVE. SEMIFIVE Starts Mass Production of Its 14nm AI Inference SoC Platform Based Product; SEMIFIVE: Seongnam, Republic of Korea, 2024. [Google Scholar]
- Coral Retrain an Image Classification Model. Available online: https://www.coral.ai/docs/edgetpu/retrain-classification/ (accessed on 22 May 2026).
- Coral Get Started with the Dev Board Micro. Available online: https://www.coral.ai/docs/dev-board-micro/get-started/ (accessed on 22 May 2026).
- NVIDIA. Installation Guide Overview—NVIDIA TensorRT. Available online: https://docs.nvidia.com/deeplearning/tensorrt/latest/installing-tensorrt/overview.html (accessed on 22 May 2026).
- AMD. Introduction to Versal Adaptive SoCs—AM009; AMD: Santa Clara, CA, USA, 2026. [Google Scholar]
- UCIe Consortium Specifications. Available online: https://www.uciexpress.org/specifications (accessed on 1 June 2026).
- Schuman, C.D.; Kulkarni, S.R.; Parsa, M.; Mitchell, J.P.; Date, P.; Kay, B. Opportunities for Neuromorphic Computing Algorithms and Applications. Nat. Comput. Sci. 2022, 2, 10–19. [Google Scholar] [CrossRef] [Scilit]
- Wan, W.; Kubendran, R.; Schaefer, C.; Eryilmaz, S.B.; Zhang, W.; Wu, D.; Deiss, S.; Raina, P.; Qian, H.; Gao, B.; et al. A Compute-in-Memory Chip Based on Resistive Random-Access Memory. Nature 2022, 608, 504–512. [Google Scholar] [CrossRef] [Scilit]
- Ambrogio, S.; Narayanan, P.; Okazaki, A.; Fasoli, A.; Mackin, C.; Hosokawa, K.; Nomura, A.; Yasuda, T.; Chen, A.; Friz, A.; et al. An Analog-AI Chip for Energy-Efficient Speech Recognition and Transcription. Nature 2023, 620, 768–775. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Zhou, T.; Ma, M.; Deng, C.; Dai, Q.; Fang, L. Large-Scale Photonic Chiplet Taichi Empowers 160-TOPS/W Artificial General Intelligence. Science 2024, 384, 202–209. [Google Scholar] [CrossRef] [Scilit]
- Muir, D.R.; Sheik, S. The Road to Commercial Success for Neuromorphic Technologies. Nat. Commun. 2025, 16, 3586. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, Y.-H.; Yang, T.-J.; Emer, J.S.; Sze, V. Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices. IEEE J. Emerg. Sel. Top. Circuits Syst. 2019, 9, 292–308. [Google Scholar] [CrossRef] [Scilit]
- Biglari, S.; Hosseini, F.; Upadhyay, A.; Zhao, H. Survey of Network-on-Chip (NoC) for Heterogeneous Multicore Systems. In Proceedings of the 2024 IEEE 17th International Symposium on Embedded Multicore/Many-Core Systems-on-Chip (MCSoC); IEEE: New York, NY, USA, 2024; pp. 155–162. [Google Scholar]
- Zhou, X.; Hao, P.; Liu, D. PCCNoC: Packet Connected Circuit as Network on Chip for High Throughput and Low Latency SoCs. Micromachines 2023, 14, 501. [Google Scholar] [CrossRef] [Scilit]
- Cai, H.; Gan, C.; Wang, T.; Zhang, Z.; Han, S. Once-for-All: Train One Network and Specialize It for Efficient Deployment. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Cai, H.; Zhu, L.; Han, S. ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware. In Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Yang, T.-J.; Howard, A.; Chen, B.; Zhang, X.; Go, A.; Sandler, M.; Sze, V.; Adam, H. NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications. In Proceedings of the European Conference on Computer Vision (ECCV); Springer International Publishing: Cham, Switzerland, 2018. [Google Scholar]
- Wang, K.; Liu, Z.; Lin, Y.; Lin, J.; Han, S. HAQ: Hardware-Aware Automated Quantization with Mixed Precision. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE Computer Society: Washington, DC, USA, 2019. [Google Scholar]
- He, Y.; Lin, J.; Liu, Z.; Wang, H.; Li, L.-J.; Han, S. AMC: AutoML for Model Compression and Acceleration on Mobile Devices. In Proceedings of the European Conference on Computer Vision (ECCV); Springer International Publishing: Cham, Switzerland, 2018. [Google Scholar]
- Snider, D.; Liang, R. Operator Fusion in XLA: Analysis and Evaluation. arXiv 2023, arXiv:2301.13062. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Moreau, T.; Jiang, Z.; Zheng, L.; Yan, E.; Shen, H.; Wang, M.; Zhu, Y.; Muench, A.; Yu, C.; et al. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18); USENIX Association: Berkeley, CA, USA, 2018; pp. 578–594. [Google Scholar]
- Jacob, B.; Kligys, S.; Chen, B.; Zhu, M.; Tang, M.; Howard, A.; Adam, H.; Kalenichenko, D. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 2704–2713. [Google Scholar]
- Han, S.; Mao, H.; Dally, W.J. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. arXiv 2016, arXiv:1510.00149. [Google Scholar] [CrossRef] [Scilit]
- Bochkovskiy, A.; Wang, C.-Y.; Liao, H.-Y.M. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. Scaled-YOLOv4: Scaling Cross Stage Partial Network. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2021; pp. 13024–13033. [Google Scholar]
- Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 7464–7475. [Google Scholar]
- Montgomerie-Corcoran, A.; Toupas, P.; Yu, Z.; Bouganis, C.-S. SATAY: A Streaming Architecture Toolflow for Accelerating YOLO Models on FPGA Devices. In Proceedings of the 2023 International Conference on Field Programmable Technology (ICFPT); IEEE: New York, NY, USA, 2023; pp. 179–187. [Google Scholar]
- Shao, C.; Tang, K.; Cheng, H.; Li, H.; Tang, Z. Implementation and Optimization of Object Detection on FPGA-Based CPU+NPU Heterogeneous System. In Proceedings of the 2025 3rd International Conference on Communication, Security, and Artificial Intelligence (ICCSAI); IEEE: New York, NY, USA, 2025; pp. 1489–1494. [Google Scholar]
- Anupreetham, A.; Ibrahim, M.; Hall, M.; Boutros, A.; Kuzhively, A.; Mohanty, A.; Nurvitadhi, E.; Betz, V.; Cao, Y.; Seo, J.-S. High Throughput FPGA-Based Object Detection via Algorithm-Hardware Co-Design. ACM Trans. Reconfigurable Technol. Syst. 2024, 17, 1–20. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Pool, J.; Han, S.; Dally, W.J. Efficient Sparse-Winograd Convolutional Neural Networks. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Restle, P.J. Processor Clock Generation, Distribution, and Clock Sensor/Management Loops. In Proceedings of the 2021 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 13–22 February 2021. [Google Scholar]
- Accellera Systems Initiative Clock Domain Crossing Standard Version 0.3 Draft for Public Review. Available online: https://www.accellera.org/images/downloads/drafts-review/CDC_0.3_Public_Review_Draft_2024.07.15.pdf (accessed on 7 June 2026).
- Zimmer, B.; Venkatesan, R.; Shao, Y.S.; Clemons, J.; Fojtik, M.; Jiang, N.; Keller, B.; Klinefelter, A.; Pinckney, N.; Raina, P.; et al. A 0.32–128 TOPS, Scalable Multi-Chip-Module-Based Deep Neural Network Inference Accelerator with Ground-Referenced Signaling in 16 Nm. IEEE J. Solid-State Circuits 2020, 55, 920–932. [Google Scholar] [CrossRef] [Scilit]
- Yu, C.-H.; Bae, J.; Kim, J.; Kim, H.; Shin, W.; Yoon, J.-S.; Jin, Y.-J.; Oh, J.; Lee, J.; Kim, E.; et al. A Quad-Chiplet AI SoC with Full-Chip Scalable Mesh Over 16Gb/s UCIe-Advanced Die-to-Die Interface for Large-Scale AI Inferencing. In Proceedings of the 2026 IEEE International Solid-State Circuits Conference (ISSCC); IEEE: New York, NY, USA, 2026; pp. 44–46. [Google Scholar]
- Murali, G.; Park, H.; Qin, E.; Torun, H.M.; Dolatsara, M.A.; Swaminathan, M.; Krishna, T.; Lim, S.K. Clock Delivery Network Design and Analysis for Interposer-Based 2.5-D Heterogeneous Systems. IEEE Trans. Very Large Scale Integr. VLSI Syst. 2021, 29, 605–616. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Hong, M.-S.; Do, K.; Choi, J.Y.; Park, J.; Kumar, M.; Kumar, M.; Tripathi, N.; Ranjan, A. Clock Domain Crossing Aware Sequential Clock Gating. In Proceedings of the 2015 Design, Automation & Test in Europe Conference & Exhibition (DATE), Grenoble, France, 9–13 March 2015; pp. 1–6. [Google Scholar]
- Rotem, E.; Mendelson, A.; Ginosar, R.; Weiser, U. Multiple Clock and Voltage Domains for Chip Multi Processors. In Proceedings of the 42nd Annual IEEE/ACM International Symposium on Microarchitecture; ACM: New York, NY, USA, 2009; pp. 459–468. [Google Scholar]
- Semeraro, G.; Magklis, G.; Balasubramonian, R.; Albonesi, D.H.; Dwarkadas, S.; Scott, M.L. Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency Scaling. In Proceedings of the Eighth International Symposium on High Performance Computer Architecture; IEEE Computer Society: Washington, DC, USA, 2002; pp. 29–40. [Google Scholar]
- Barboza, E.C.; Shukla, N.; Chen, Y.; Hu, J. Machine Learning-Based Pre-Routing Timing Prediction with Reduced Pessimism. In Proceedings of the 56th Annual Design Automation Conference 2019; ACM: New York, NY, USA, 2019; pp. 1–6. [Google Scholar]
- Xie, Z.; Li, H.; Xu, X.; Hu, J.; Chen, Y. Fast IR Drop Estimation with Machine Learning. In Proceedings of the 39th International Conference on Computer-Aided Design; ACM: New York, NY, USA, 2020; pp. 1–8. [Google Scholar]
- Jiang, X.; Chai, Z.; Zhao, Y.; Lin, Y.; Wang, R.; Huang, R. CircuitNet 2.0: An Advanced Dataset for Promoting Machine Learning Innovations in Realistic Chip Design Environment. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Jiang, X.; Guo, Z.; Chai, Z.; Zhao, Y.; Lin, Y.; Wang, R.; Huang, R. Invited Paper: Accelerating Routability and Timing Optimization with Open-Source AI4EDA Dataset CircuitNet and Heterogeneous Platforms. In Proceedings of the 2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD); IEEE: New York, NY, USA, 2023; pp. 1–9. [Google Scholar]
- Gao, X.; Jiang, Y.-M.; Shao, L.; Raspopovic, P.; Verbeek, M.E.; Sharma, M.; Rashingkar, V.; Jalota, A. Congestion and Timing Aware Macro Placement Using Machine Learning Predictions from Different Data Sources. In Proceedings of the 2022 International Symposium on Physical Design; ACM: New York, NY, USA, 2022; pp. 195–202. [Google Scholar]
- Oh, C.; Bondesan, R.; Kianfar, D.; Ahmed, R.; Khurana, R.; Agarwal, P.; Lepert, R.; Sriram, M.; Welling, M. Bayesian Optimization for Macro Placement. arXiv 2022, arXiv:2207.08398. [Google Scholar] [CrossRef] [Scilit]
- Mirhoseini, A.; Goldie, A.; Yazgan, M.; Jiang, J.W.; Songhori, E.; Wang, S.; Lee, Y.-J.; Johnson, E.; Pathak, O.; Nova, A.; et al. A Graph Placement Methodology for Fast Chip Design. Nature 2021, 594, 207–212. [Google Scholar] [CrossRef] [Scilit]
- Pan, J.; Zhou, G.; Chang, C.-C.; Jacobson, I.; Hu, J.; Chen, Y. A Survey of Research in Large Language Models for Electronic Design Automation. ACM Trans. Des. Autom. Electron. Syst. 2025, 30, 1–21. [Google Scholar] [CrossRef] [Scilit]
- Chen, D.; Ganesh, V.; Li, W.; Lin, Y.C.; Liu, Y.; Mitra, S.; Pan, D.Z.; Puri, R.; Cong, J.; Sun, Y. Report for NSF Workshop on AI for Electronic Design Automation. IEEE Circuits Syst. Mag. 2026, 26, 67–81. [Google Scholar] [CrossRef] [Scilit]
- Xu, K.; Schwachhofer, D.; Blocklove, J.; Polian, I.; Domanski, P.; Pflüger, D.; Garg, S.; Karri, R.; Sinanoglu, O.; Knechtel, J.; et al. Large Language Models (LLMs) for Electronic Design Automation (EDA): Special Session Paper. In Proceedings of the 2025 IEEE 38th International System-on-Chip Conference (SOCC); IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar]
- Ghose, A.; Kahng, A.B.; Kundu, S.; Wang, Z. ORFS-Agent: Tool-Using Agents for Chip Design Optimization. In Proceedings of the 2025 ACM/IEEE 7th Symposium on Machine Learning for CAD (MLCAD); IEEE: New York, NY, USA, 2025; pp. 1–13. [Google Scholar]
- Markov, I.L. Reevaluating Google’s Reinforcement Learning for IC Macro Placement. Commun. ACM 2024, 67, 60–71. [Google Scholar] [CrossRef] [Scilit]
- Islam, M.U.; Sami, H.; Gaillardon, P.-E.; Tenace, V. EDA-Aware RTL Generation with Large Language Models. In 2025 Design, Automation & Test in Europe Conference (DATE); IEEE: New York, NY, USA, 2024. [Google Scholar]
- Ahsan, S.M.M.; Shahriar, M.S.; Chowdhury, M.; Hossain, T.; Hasan, M.S.; Hoque, T. Accurate, Yet Scalable: A SPICE-Based Design and Optimization Framework for ENVM Based Analog In-Memory Computing. In Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design; ACM: New York, NY, USA, 2024; pp. 1–9. [Google Scholar]
- Liu, S.; Lu, Y.; Fang, W.; Li, M.; Xie, Z. OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation. In Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design; IEEE: New York, NY, USA, 2025. [Google Scholar]
- Zang, Z.; Song, Y.; Wang, A.; Ling, B.W.-K.; Sun, Q.; Lei, Z.; Yang, F.; Zhuo, C.; Luo, J. The Dawn of Agentic EDA: A Survey of Autonomous Digital Chip Design. arXiv 2026, arXiv:2512.23189. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Dong, R.; Ye, R.; Singh, S.S.K.; Wu, S.; Chen, C. A Review of Thermal Performance of 3D Stacked Chips. Int. J. Heat Mass Transf. 2024, 235, 126212. [Google Scholar] [CrossRef] [Scilit]
- Song, R.; Zhang, J.; Zhu, Z.; Shan, G.; Yang, Y. Fault and Self-Repair for High Reliability in Die-to-Die Interconnection of 2.5D/3D IC. Microelectron. Reliab. 2024, 158, 115429. [Google Scholar] [CrossRef] [Scilit]
- Zhou, A.; Zhang, Y.; Ding, F.; Lian, Z.; Jin, R.; Yang, Y.; Wang, Q.; Cao, L. Research Progress of Hybrid Bonding Technology for Three-Dimensional Integration. Microelectron. Reliab. 2024, 155, 115372. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Wang, X.; Zheng, R.; Zhao, X.; Zheng, H.; Zhao, Z.; Cheng, M.; Jiang, Y.; Jia, Y. Finite Element Analysis of 2.5D Packaging Processes Based on Multi-Physics Field Coupling for Predicting the Reliability of IC Components. Microelectron. Reliab. 2024, 163, 115530. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Wang, Y.; Wang, X.; Xu, H.; Song, Y.; Wang, Z.; Yang, Y.; Wu, Q.; Xu, J. Photonic Edge Intelligence Chip for Multi-Modal Sensing, Inference and Learning. Nat. Commun. 2025, 16, 5897. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Xu, H.; Wu, Q.; Wang, Y.; Yang, Z.; Song, Y.; Yang, H.; Wang, Z.; Jiang, T.; Chen, L.; et al. All-Analog Photoelectronic Chip for High-Speed Vision Tasks. Nature 2023, 623, 48–57. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cassidy, A.S.; Sawada, J.; Kahng, A.B.; Lu, J.; Le, M.P.; Rabaey, J.M.; Modha, D.S.; Esser, S.; Otero, C.O.; Sawada, J.; et al. IBM NorthPole: An Architecture for Neural Network Inference with a 12 Nm Chip. Science 2024, 385, eadk5952. [Google Scholar]
- Intel Intel Builds World’s Largest Neuromorphic System to Enable More Sustainable AI. Available online: https://newsroom.intel.com/artificial-intelligence/intel-builds-worlds-largest-neuromorphic-system-to-enable-more-sustainable-ai (accessed on 22 May 2026).
- Pedersen, J.E.; Abreu, S.; Jobst, M.; Lenz, G.; Fra, V.; Bauer, F.C.; Muir, D.R.; Zhou, P.; Vogginger, B.; Heckel, K.; et al. Neuromorphic Intermediate Representation: A Unified Instruction Set for Interoperable Brain-Inspired Computing. Nat. Commun. 2024, 15, 8067. [Google Scholar] [CrossRef] [Scilit]
- Yik, J.; Berghe, K.V.D.; Blanken, D.D.; Bouhadjar, Y.; Fabre, M.; Hueber, P.; Ke, W.; Khoei, M.A.; Kleyko, D.; Pacik-Nelson, N.; et al. The NeuroBench Framework for Benchmarking Neuromorphic Computing Algorithms and Systems. Nat. Commun. 2025, 16, 1332. [Google Scholar] [CrossRef] [Scilit]
- Sandia National Laboratories. Advancing Neuromorphic Computing at the Neural Exploration and Research Lab. 2024. Available online: https://www.sandia.gov/news/publications/hpc-annual-reports/article/advancing-neuromorphic-computing-at-the-neural-exploration-and-research-lab/ (accessed on 7 June 2026).
- Yang, Z.-Z.; Wang, C.; Zhao, Y.; Ruan, G.-J.; Yangdong, X.-J.; Yang, Y.; Pan, C.; Cheng, B.; Liang, S.-J.; Miao, F. Communication-Aware in-Memory Wireless Neural Networks. Nat. Electron. 2026, 9, 414–425. [Google Scholar] [CrossRef] [Scilit]
- Raha, A.; Mathaikutty, D.A.; Kundu, S.; Ghosh, S.K. FlexNPU: A Dataflow-Aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices. Front. High Perform. Comput. 2025, 3, 1570210. [Google Scholar] [CrossRef] [Scilit]
- Elbtity, M.; Chandarana, P.; Zand, R. Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture. arXiv 2024, arXiv:2407.08700. [Google Scholar] [CrossRef] [Scilit]
- AMD. AMD Instinct MI300X Accelerators. Available online: https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html (accessed on 22 May 2026).
- IDC. AI Infrastructure Spending Caps Historic Year at $89.9 Billion in Q4 2025; 2029 Spending to Eclipse $1 Trillion. Available online: https://www.idc.com/resource-center/blog/ai-infrastructure-spending-caps-historic-year-at-90-billion-in-q4-2025-2029-spending-to-eclipse-1-trillion/ (accessed on 22 May 2026).
- NVIDIA. NVIDIA Announces Financial Results for Fourth Quarter and Fiscal. 2026. Available online: https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026 (accessed on 22 May 2026).
- Synergy Research Group Hyperscale Operators to Account for 67% of All Data Center Capacity by 2031. Available online: https://www.srgresearch.com/articles/hyperscale-operators-to-account-for-67-of-all-data-center-capacity-by-2031 (accessed on 22 May 2026).
- Chang, Y.-H. AI Chips for Edge Applications 2026–2036: Technologies, Markets, Forecasts. 2026. Available online: https://www.idtechex.com/en/research-report/ai-chips-for-edge-applications/1148 (accessed on 7 June 2026).
- Merizzi, N.; Thomas, C.; Burns, E. The AI Infrastructure Reckoning: Optimizing Compute Strategy in the Age of Inference Economics. 2025. Available online: https://www.deloitte.com/ro/en/Industries/technology/perspectives/the-ai-infrastructure-reckoning-optimizing-compute-strategy-in-the-age-of-inference-economics.html (accessed on 7 June 2026).
- Who’s Funding the AI Data-Center Boom? Available online: https://www.mckinsey.com/featured-insights/themes/whos-funding-the-ai-data-center-boom (accessed on 22 May 2026).
- IEA. Energy and AI; IEA: Paris, France, 2025. [Google Scholar]























| Chip Type | Architecture | Application |
|---|---|---|
| Edge AI ASIC | CNN accelerator | Vision detection |
| FPSoC | CPU + FPGA + NPU | Adaptive edge AI |
| GPU AI chip | GPU AI chip | GPU AI chip |
| Optimization Method | Purpose | Hardware Impact |
|---|---|---|
| Model pruning | Reduce parameters | Lower memory usage |
| Quantization | Reduce precision | Energy savings |
| FFT/Winograd convolution | Faster inference | Reduced computation |
| Feature pyramid fusion | Improve accuracy | Enhanced parallelism |
| Feature | Lab-on-Chip Systems | IIoT AI-with-Chip Systems |
|---|---|---|
| Scale | Microscale | System-level/Industrial |
| Sensing Type | Biological/Chemical | Mechanical/Environmental |
| Actuation | Microfluidic/Optical | Control signals/Alerts |
| AI Integration | Embedded ML/DL models | FPGA + NoC AI accelerators |
| Real-Time Capability | High (localized processing) | High (edge + cloud hybrid) |
| Applications | Diagnostics, lab automation | Predictive maintenance, monitoring |
| Communication | Minimal/local | Wireless + cloud-connected |
| System Complexity | Compact, highly integrated | Distributed, multi-layered |
| Class | Integration Unit | Typical Interconnect | Memory Organization | Key Strengths | Main Limitations |
|---|---|---|---|---|---|
| Monolithic AI SoC [28] | Single die | Bus, crossbar, local fabric | On-die SRAM + off-chip DRAM/LPDDR | Low latency, compact integration, lower board complexity | Reticle/yield limits, lower maximum local memory |
| NoC-enabled multi-core chip [4,6] | PE cluster/tile | Mesh, tree, ring, multicast NoC | Tiled buffers, shared global buffers, scratchpads | Scales parallel compute; explicit reuse and multicast | NoC congestion, mapping complexity, sensitivity to workload pattern |
| Chiplet 2.5D/3D AI system [27,28] | Chiplet/package | Die-to-die links, interposer routes, vertical TSVs | HBM near-package, stacked memory, distributed caches | Scalability, modularity, heterogeneous process integration | Packaging cost, thermals, validation, coherency/test complexity |
| Platform | Hardware Class | Key Architectural Emphasis | Deployment Ecosystem | Main Implication | Ref. |
|---|---|---|---|---|---|
| Mobilint ARIES | Edge AI ASIC/NPU | Dedicated inference acceleration for edge deployment | qb SDK; ONNX, TensorFlow, PyTorch, TVM support | Strong example of hardware–SDK co-design for edge inference | [30,48,50] |
| Edge AI chip for CNN head detection | Application-specific edge AI chip | On-device CNN vision inference for embedded detection | Vendor evaluation-board workflow | Illustrates near-sensor inference for constrained vision tasks | [30] |
| Heterogeneous FPSoC | Hybrid edge AI SoC | CPU + NPU + FPGA logic + media support | Model conversion, quantization, heterogeneous mapping flow | Shows that flexible edge systems often combine fixed and programmable resources | [29] |
| Coral Edge TPU | Edge inference accelerator | Quantized low-power inference | TensorFlow Lite + Edge TPU Compiler + Coral APIs | Demonstrates SDK-driven deployment pipeline as part of the platform | [49,51,52] |
| NVIDIA TensorRT ecosystem | Embedded/edge GPU inference stack | Optimized runtime generation, graph fusion, mixed precision | TensorRT SDK with ONNX parser and runtime APIs | Highlights the importance of optimization software in deployment efficiency | [53] |
| Design theme | Typical Traffic Emphasis | Main Architectural Idea | Strength for AI Workloads | Main Limitation | Ref. |
|---|---|---|---|---|---|
| Bus/crossbar SoC | Shared arbitration, increasing contention | Centralized or semi-centralized interconnect | Simpler for small systems | Poor scalability with many heterogeneous clients | [62,63] |
| Spatial array with reuse-aware dataflow | Global-buffer-to-PE delivery, local reuse | Multi-level memory hierarchy + dataflow optimization | Reduces expensive data movement | Sensitive to mapping and layer shape | [31,33] |
| Hierarchical mesh NoC | Mixed multicast + high-bandwidth staged delivery | Clustered mesh linking global buffer and PE groups | Good balance of scalability and reuse support | More complex than simple mesh or bus | [61] |
| Tree-based multicast NoC | One-to-many fan-out | Router and topology specialized for multicast | Strong fit for broadcast-heavy AI traffic | Less general than regular mesh for arbitrary traffic | [34] |
| NoC-enabled many-core/chiplet AI | Distributed compute-memory communication | Interconnect co-designed with memory and compute scaling | Supports larger AI fabrics and modular scaling | Communication becomes a dominant bottleneck | [7,35] |
| NoC-enabled edge/IIoT AI platform | Sensor-to-buffer-to-inference-to-notification flow | Real-time heterogeneous integration | Good for predictive maintenance and industrial streaming AI | Design complexity across sensing, control, and inference | [6] |
| Technique | Main Idea | Typical Hardware Benefit | Representative Reported Statistic | Main Caveat | Ref. |
|---|---|---|---|---|---|
| Pruning + quantization + coding | Remove redundant weights, lower precision, entropy coding | Lower memory footprint and less data movement | AlexNet 35× compression; VGG-16 49× compression; 3×–4× layerwise speedup; 3×–7× energy improvement | Irregular sparsity may not map efficiently to dense engines | [72] |
| Integer-only quantization | Convert inference to integer arithmetic with co-designed scaling | Better fit to integer datapaths and low-power hardware | Preserved near-floating-point accuracy while enabling integer-only inference | Calibration and retraining sensitivity | [71] |
| AutoML compression | Search compression policies automatically for target hardware | Better measured speed/accuracy tradeoff than handcrafted rules | 1.81× speedup on Android phone; 1.43× on Titan XP at minimal accuracy loss | Search complexity and platform dependence | [68] |
| Mixed-precision quantization | Layer-wise bit-width assignment with hardware feedback | Lower latency and energy than uniform precision | 1.4×–1.95× lower latency and 1.9× lower energy than fixed 8-bit quantization | Requires hardware-in-the-loop evaluation | [67] |
| Platform-aware adaptation | Adapt model iteratively using measured latency/energy | Better real-device deployment tradeoff | Up to 1.7× speedup on mobile CPU/GPU for MobileNets | Device-specific adaptation flow | [66] |
| Operator fusion | Fuse operators to reduce intermediate memory transfers and launches | Higher arithmetic intensity and lower memory overhead | XLA fusion strategies reported up to 10.56× speedup in the evaluated case | Fusion opportunities depend on graph structure and compiler support | [67,70] |
| Strategy | Targeted Stage | Main Optimization Idea | Representative Reported Statistic | Hardware Implication | Ref. |
|---|---|---|---|---|---|
| YOLOv4 “bag-of-freebies” + efficient backbone design | End-to-end detector | Improve speed/accuracy without exotic hardware assumptions | 43.5% AP at ~65 FPS on Tesla V100 | Good baseline for real-time GPU deployment | [73] |
| Scaled-YOLOv4 family scaling | Backbone/head scaling across sizes | Move across the speed–accuracy frontier by coordinated scaling | YOLOv4-large: 55.5% AP at ~16 FPS; YOLOv4-tiny: 22.0% AP at 443 FPS; TensorRT FP16 tiny: 1774 FPS | Supports both high-accuracy and ultra-fast deployment modes | [74] |
| YOLOv7 trainable bag-of-freebies | Detector architecture + training design | Improve real-time AP without abandoning one-stage detection | 56.8% AP overall real-time regime; YOLOv7-E6: 55.9% AP at 56 FPS | Strong fit for high-throughput accelerator deployment | [75] |
| Streaming FPGA YOLO toolflow | Full detector on FPGA | Deeply pipelined streaming execution with automated generation | Competitive performance and energy characteristics vs. GPU devices | Suitable for low-latency edge inference | [76] |
| CPU + NPU heterogeneous object detection | Whole pipeline | Partition pipeline across compute styles | Demonstrates practical CPU + NPU collaboration on FPGA-based heterogeneous system | Useful when preprocessing/post-processing do not map well to one engine | [77] |
| Pipelined NMS and end-to-end FPGA object detection | Post-processing bottleneck | Remove serial NMS overhead | 2167 FPS, 2.13 ms latency, 5.3× throughput gain, 5× lower latency | Shows that end-to-end detector optimization must include NMS | [78] |
| Methodology | Core Design Loop | Hardware Signal Used | Representative Reported Outcome | Main Value | Ref. |
|---|---|---|---|---|---|
| NetAdapt | Progressive simplification of pre-trained model | Measured latency/energy | Up to 1.7× speedup on mobile CPU/GPU | Directly optimizes deployment metrics | [66] |
| ProxylessNAS | Direct architecture search on target hardware | Measured latency on target platform | 1.2× faster than MobileNetV2 with better top-1 accuracy in reported ImageNet case | Removes proxy mismatch between search and deployment | [65] |
| FBNet | Differentiable hardware-aware NAS | Operator latency table on target device | Device-aware ConvNet design via differentiable search | Couples block selection to measured operator cost | [64] |
| Once-for-All | Train supernetwork, then specialize subnetworks | Device/resource constraint during specialization | Same accuracy but 1.5× faster than MobileNetV3; 2.6× faster than EfficientNet in reported comparisons | Scales hardware-aware design to many devices | [64] |
| HAQ | Mixed-precision search with hardware feedback | Latency/energy simulator | 1.4×–1.95× lower latency and 1.9× lower energy than fixed 8-bit | Makes bit-width assignment platform-specific | [67] |
| Sparse-Winograd co-design | Reformulate, transform, and sparsity jointly | Multiplication count/mapped efficiency | 10.4×, 6.8×, and 10.8× multiplication reduction with <0.1% accuracy loss | Shows sparsity must align with execution transform | [79] |
| Challenge | Why It Arises in AI Chips | Main Timing Effect | Typical Design Response | Ref. |
|---|---|---|---|---|
| Long global clock paths | Large die area and distributed compute/memory macros | More skew, latency, and sensitivity to variation | Spines, meshes, hierarchical trees, local deskew | [80,87] |
| High sink count | Large arrays and dense sequential logic | Buffer growth and higher clock power | Hierarchical CTS and localized clock regions | [80,87] |
| Multiple generated clocks | Control, memory, I/O, and compute blocks run at different rates | CDC verification and metastability risk | Explicit clock grouping, synchronizers, async FIFOs | [81,86] |
| Timing/noise coupling | Local voltage droop and loading imbalance near active arrays | Reduced timing margin and jitter robustness | Local clock tuning, useful skew, guard-banding | [80,87] |
| Reused heterogeneous IP | Mix of accelerators, PHYs, CPU clusters, and controllers | Partial synchronicity and reset-crossing issues | CDC- and RDC-aware signoff methodology | [81,85] |
| Strategy | Main Idea | Benefit | Main Challenge | Ref. |
|---|---|---|---|---|
| Flat global clock tree | One dominant tree spans most logic | Simpler global timing model | Harder skew closure on large or heterogeneous systems | [80,87] |
| Hierarchical tree/spine/mesh | Global distribution with regional specialization | Better skew robustness and design closure | Higher implementation complexity and clock-network cost | [80] |
| Multi-clock-domain partitioning | Local PLLs/dividers generate domain clocks | Simpler local closure and domain-specific DVFS | CDC latency and verification burden | [81,87] |
| Interposer + on-chiplet clocking | Package-level reference plus local chiplet clocking | Better fit to 2.5D/3D packaging and chiplet modularity | Package-aware modeling and inter-chiplet synchronization | [84] |
| Chiplet-local gating and deskew | Each chiplet adapts clock locally | Lower dynamic clock power and more specialization | Requires stronger coordination across die-to-die boundaries | [81,84,85] |
| Analysis Target | Design-Stage Role | Why AI Helps | Representative Evidence | Ref. |
|---|---|---|---|---|
| Pre-routing timing prediction | Estimate timing before routing/sign-off | Reduces pessimism and avoids over-design earlier in flow | Accuracy reported near post-routing timing in ML-based pre-routing timing prediction | [88] |
| IR-drop prediction | Identify power-integrity risk earlier | Avoids repeated expensive analysis/fix cycles | It highlights IR-drop evaluation cost and the need for fast prediction | [89] |
| Congestion/routability prediction | Screen floorplans and placements before full routing | Predictive pruning of poor layouts reduces turnaround time | CircuitNet 2.0 provides >10,000 samples; additional studies report large congestion datasets from industrial-style designs | [90,92] |
| Timing + routability joint acceleration | Use AI surrogates and heterogeneous compute to speed physical design | Speeds optimization loops that no longer scale well on CPU alone | It targets routability and timing optimization using open-source AI4EDA datasets and heterogeneous platforms | [91] |
| Method Class | Typical Task | Main Strength | Main Limitation | Representative Reference |
|---|---|---|---|---|
| Supervised/surrogate prediction | Timing, IR-drop, congestion prediction | Speeds early screening and reduces expensive sign-off iterations | Requires representative training data and good transfer across designs | [88,89,90,91,92] |
| Bayesian optimization | Flow tuning, analog sizing, macro placement | Sample-efficient for expensive black-box objectives | Scales poorly in very high-dimensional spaces; needs explicit objectives | [93,98] |
| Reinforcement learning | Floorplanning/macro placement | Handles sequential placement decisions and learns reusable policies | Reproducibility and benchmark generalization remain debated | [94,99] |
| LLM-assisted RTL generation | RTL, scripts, assertions, code repair | Flexible interface from natural language to design artifacts | Functional correctness still fragile without tool feedback | [95,100,102] |
| Agentic EDA workflows | End-to-end or multi-stage orchestration | Uses tool logs, APIs, and verification feedback to iteratively improve designs | Needs robust infrastructure, metrics, and safety/verification controls | [95,96,97,98] |
| Packaging Approach | Main Architectural Benefit | Main Thermal Issue | Main Mechanical/Reliability Issue | Representative References |
|---|---|---|---|---|
| Monolithic die | Short on-die interconnect and simpler assembly | Local hotspots within one die | Large die yield loss and package-level warpage | [7,26] |
| 2.5D interposer integration | High-bandwidth chiplet/HBM connectivity | Interposer/package heat spreading and edge-to-center gradients | Underfill, interposer, and microbump stress during assembly | [26,107] |
| 3D stacking with TSVs or hybrid bonding | Maximum bandwidth density and shortest memory-logic paths | Vertical heat accumulation and buried-tier cooling difficulty | Thermomechanical stress, interconnect fatigue, bonding defects | [104,105,106] |
| Design Direction | Core Advantage | Principal Weakness | Most Natural Workloads | Near-Term Outlook |
|---|---|---|---|---|
| Photonic/photoelectronic | Extreme bandwidth, very low latency, strong energy efficiency for selected kernels | Calibration, analog noise, packaging immaturity, limited generality | Sensor-near vision, RF/spectral analytics, fast linear transforms | Strong for niche acceleration; not yet a universal replacement |
| Neuromorphic | Event-driven efficiency, adaptation, sparse/time-based sensing | Software fragmentation, benchmark variability, limited mainstream deployment | Continuous sensing, robotics, adaptive edge AI | Advancing, but ecosystem maturity remains decisive |
| Analog in-memory | Attacks the memory wall directly for matrix-heavy inference | Device nonidealities, precision management, retraining burden | Edge/cloud inference, speech, compact transformers | Most credible medium-term non-von-Neumann path |
| Multimodal photonic edge | Native analog-signal fusion on chip | Narrow application fit, optical integration complexity | Images, spectra, and RF at the edge | Promising for specialized embedded intelligence |
| Reconfigurable dataflow accelerators | Better utilization across heterogeneous layers and models | Area/control overhead, compiler complexity | Edge NPUs, mixed CNN/Transformer pipelines | Likely to expand as workloads diversify |
| Segment | Demand Driver | Economic Logic | Representative Hardware Profile | Main Bottleneck |
|---|---|---|---|---|
| Cloud training | Frontier-model scale, parallelism, HBM bandwidth | Highest absolute ASP and capex intensity; utilization must stay high | Very large 2.5D/3D accelerator packages with dense networking | Power, packaging, HBM, facility build-out |
| Cloud inference | Continuous token serving, agentic workloads | Strong revenue visibility; economics depend on utilization and latency | Large-memory chiplets, inference-optimized accelerators, hybrid fleets | Power, memory cost, and scheduling efficiency |
| Enterprise on-prem | Data sovereignty, predictable recurring inference | Capex is justified when cloud bills become persistent and high | Smaller clusters, often mixed CPU/GPU/NPU | Expertise, cooling, orchestration complexity |
| Edge consumer | Smartphones, PCs, wearables | Volume and bill-of-materials sensitivity dominate; energy efficiency is critical | Integrated NPUs, reconfigurable dataflow, and smaller form factors | Thermal envelope, software fragmentation |
| Edge industrial/sovereign | Real-time control, robotics, defense, and local compliance | Latency and resilience outweigh pure cloud efficiency | Ruggedized edge servers, localized accelerators, private AI stacks | Certification, reliability, lifecycle cost |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Morsy, M.M. AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence. Electronics 2026, 15, 2645. https://doi.org/10.3390/electronics15122645
Morsy MM. AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence. Electronics. 2026; 15(12):2645. https://doi.org/10.3390/electronics15122645
Chicago/Turabian StyleMorsy, Mohamed M. 2026. "AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence" Electronics 15, no. 12: 2645. https://doi.org/10.3390/electronics15122645
APA StyleMorsy, M. M. (2026). AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence. Electronics, 15(12), 2645. https://doi.org/10.3390/electronics15122645
