Joint Multiplier–Adder Approximation with Flag-Based Error Recovery for BF16 Digital Compute-in-Memory
Abstract
1. Introduction
2. Related Work
2.1. Multiplier-Centric Approximation
2.2. Accumulation-Centric Approximation
2.3. Joint Multiplier–Adder-Tree Approximation
3. Proposed Design
3.1. Joint Approximation
3.2. Design-Space Exploration
4. Evaluation and Analysis
4.1. Evaluation Setup and Baseline
4.2. PPA Evaluation
4.3. Numerical Error and Inference Accuracy
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Horowitz, M. 1.1 Computing’s Energy Problem (and What We Can Do About It). In Proceedings of the 2014 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 9–13 February 2014; pp. 10–14. [Google Scholar] [CrossRef] [Scilit]
- Sze, V.; Chen, Y.-H.; Yang, T.-J.; Emer, J.S. Efficient Processing of Deep Neural Networks: A Tutorial and Survey. Proc. IEEE 2017, 105, 2295–2329. [Google Scholar] [CrossRef] [Scilit]
- Verma, N.; Jia, H.; Valavi, H.; Tang, Y.; Ozatay, M.; Chen, L.-Y.; Zhang, B.; Deaville, P. In-Memory Computing: Advances and Prospects. IEEE Solid-State Circuits Mag. 2019, 11, 43–55. [Google Scholar] [CrossRef] [Scilit]
- Si, X.; Tu, Y.-N.; Huang, W.-H.; Su, J.-W.; Lu, P.-J.; Wang, J.-H.; Liu, T.-W.; Wu, S.-Y.; Liu, R.; Chou, Y.-C.; et al. A 28 nm 64 Kb 6T SRAM Computing-in-Memory Macro with 8b MAC Operation for AI Edge Chips. In Proceedings of the 2020 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 16–20 February 2020; pp. 246–248. [Google Scholar] [CrossRef] [Scilit]
- Chih, Y.-D.; Lee, P.-H.; Fujiwara, H.; Shih, Y.-C.; Lee, C.-F.; Naous, R.; Chen, Y.-L.; Lo, C.-P.; Lu, C.-H.; Mori, H.; et al. An 89 TOPS/W and 16.3 TOPS/mm2 All-Digital SRAM-Based Full-Precision Compute-In-Memory Macro in 22 nm for Machine-Learning Edge Applications. In Proceedings of the 2021 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 13–22 February 2021; pp. 252–254. [Google Scholar] [CrossRef] [Scilit]
- Henry, G.; Tang, P.T.P.; Heinecke, A. Leveraging the bfloat16 Artificial Intelligence Datatype for Higher-Precision Computations. In Proceedings of the 26th IEEE Symposium on Computer Arithmetic (ARITH), Kyoto, Japan, 10–12 June 2019; pp. 69–76. [Google Scholar] [CrossRef] [Scilit]
- Tu, F.; Wang, Y.; Wu, Z.; Liang, L.; Ding, Y.; Kim, B.; Liu, L.; Wei, S.; Xie, Y.; Yin, S. ReDCIM: Reconfigurable Digital Computing-In-Memory Processor with Unified FP/INT Pipeline for Cloud AI Acceleration. IEEE J. Solid-State Circuits 2023, 58, 243–255. [Google Scholar] [CrossRef] [Scilit]
- Guo, A.; Si, X.; Chen, X.; Dong, F.; Pu, X.; Li, D.; Zhou, Y.; Ren, L.; Xue, Y.; Dong, X.; et al. A 28 nm 64-kb 31.6-TFLOPS/W Digital-Domain Floating-Point-Computing-Unit and Double-Bit 6T-SRAM Computing-in-Memory Macro for Floating-Point CNNs. In Proceedings of the 2023 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 19–23 February 2023; pp. 128–130. [Google Scholar] [CrossRef] [Scilit]
- Wu, P.-C.; Su, J.-W.; Hong, L.-Y.; Ren, J.-S.; Chien, C.-H.; Chen, H.-Y.; Ke, C.-E.; Hsiao, H.-M.; Li, S.-H.; Sheu, S.-S.; et al. A 22 nm 832 Kb Hybrid-Domain Floating-Point SRAM In-Memory-Compute Macro with 16.2–70.2 TFLOPS/W for High-Accuracy AI-Edge Devices. In Proceedings of the 2023 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 19–23 February 2023; pp. 126–128. [Google Scholar] [CrossRef] [Scilit]
- Lin, C.-T.; Oh, J.; Seok, M. STAR-SRAM: 16-bit Floating-Point SRAM-Based Digital Computing-in-Memory Macro in a 28 nm. IEEE J. Solid-State Circuits 2026, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Kim, J.; Kim, H.; Lee, Y. A 28 nm 77.2 TFLOPS/W Digital Floating-Point Compute-In-Memory Macro Employing Dynamic Find-Max and Reduced-Cycle Bit-Serial Architecture with Approximation. In Proceedings of the 2025 IEEE Asian Solid-State Circuits Conference (A-SSCC), Daejeon, Republic of Korea, 2–5 November 2025; pp. 142–144. [Google Scholar] [CrossRef] [Scilit]
- Xu, Q.; Mytkowicz, T.; Kim, N.S. Approximate Computing: A Survey. IEEE Des. Test 2016, 33, 8–22. [Google Scholar] [CrossRef] [Scilit]
- Momeni, A.; Han, J.; Montuschi, P.; Lombardi, F. Design and Analysis of Approximate Compressors for Multiplication. IEEE Trans. Comput. 2015, 64, 984–994. [Google Scholar] [CrossRef] [Scilit]
- Ha, M.; Lee, S. Multipliers with Approximate 4–2 Compressors and Error Recovery Modules. IEEE Embed. Syst. Lett. 2018, 10, 6–9. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; Han, J.; Lombardi, F. A Low-Power, High-Performance Approximate Multiplier with Configurable Partial Error Recovery. In Proceedings of the 2014 Design, Automation and Test in Europe Conference and Exhibition (DATE), Dresden, Germany, 24–28 March 2014; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
- Jiang, H.; Liu, C.; Lombardi, F.; Han, J. Low-Power Approximate Unsigned Multipliers with Configurable Error Recovery. IEEE Trans. Circuits Syst. I Regul. Pap. 2019, 66, 189–202. [Google Scholar] [CrossRef] [Scilit]
- Strollo, A.G.M.; Napoli, E.; De Caro, D.; Petra, N.; Di Meo, G. Comparison and Extension of Approximate 4–2 Compressors for Low-Power Approximate Multipliers. IEEE Trans. Circuits Syst. I Regul. Pap. 2020, 67, 3021–3034. [Google Scholar] [CrossRef] [Scilit]
- Hashemi, S.; Bahar, R.I.; Reda, S. DRUM: A Dynamic Range Unbiased Multiplier for Approximate Applications. In Proceedings of the 2015 IEEE/ACM International Conference on Computer-Aided Design (ICCAD), Austin, TX, USA, 2–6 November 2015; pp. 418–425. [Google Scholar] [CrossRef] [Scilit]
- Strollo, A.G.M.; Napoli, E.; De Caro, D.; Petra, N.; Saggese, G.; Di Meo, G. Approximate Multipliers Using Static Segmentation: Error Analysis and Improvements. IEEE Trans. Circuits Syst. I Regul. Pap. 2022, 69, 2449–2462. [Google Scholar] [CrossRef] [Scilit]
- Sabetzadeh, F.; Moaiyeri, M.H.; Ahmadinejad, M. An Ultra-Efficient Approximate Multiplier with Error Compensation for Error-Resilient Applications. IEEE Trans. Circuits Syst. II Express Briefs 2023, 70, 776–780. [Google Scholar] [CrossRef] [Scilit]
- Kumari, A.; Palathinkal, R.P. Design and Analysis of Energy Efficient Approximate Multipliers for Image Processing and Deep Neural Network. IEEE Trans. Circuits Syst. I Regul. Pap. 2025, 72, 854–867. [Google Scholar] [CrossRef] [Scilit]
- Guo, R.; Chen, X.; Wang, L.; Tu, F.; Wei, S.; Hu, Y.; Yin, S. A 28 nm 4170-TFLOPS/W/b and 195-TFLOPS/mm2/b Multiply-Free Fully-Digital Floating-Point Compute-In-Memory Macro with Mitchell’s Approximation. In Proceedings of the 2024 IEEE Symposium on VLSI Technology and Circuits, Honolulu, HI, USA, 16–20 June 2024; pp. 1–2. [Google Scholar] [CrossRef] [Scilit]
- Wang, D.; Lin, C.-T.; Chen, G.K.; Knag, P.; Krishnamurthy, R.K.; Seok, M. DIMC: 2219 TOPS/W 2569F2/b Digital In-Memory Computing Macro in 28 nm Based on Approximate Arithmetic Hardware. In Proceedings of the 2022 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 20–24 February 2022; pp. 266–268. [Google Scholar] [CrossRef] [Scilit]
- He, C.; Wang, Z.; Xiang, F.; Dai, Z.; He, Y.; Yue, J.; Liu, Y. LSAC: A Low-Power Adder Tree for Digital Computing-in-Memory by Sparsity and Approximate Circuits Co-Design. IEEE Trans. Circuits Syst. II Express Briefs 2024, 71, 852–856. [Google Scholar] [CrossRef] [Scilit]
- Napoli, E.; Zacharelos, E.; Strollo, A.G.M.; Di Meo, G. Approximate Full-Adders: A Comprehensive Analysis. IEEE Access 2024, 12, 136054–136072. [Google Scholar] [CrossRef] [Scilit]





| A\B | 0000 | 0001 | 0010 | 0011 | 0100 | 0110 | 0111 |
|---|---|---|---|---|---|---|---|
| 0000 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0001 | 0 | 0 | 0 | 0 | 0 | 0 | −2 |
| 0010 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0011 | 0 | 0 | 0 | 0 | 0 | 0 | −2 |
| 0100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0110 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0111 | 0 | −2 | 0 | −2 | 0 | 0 | −4 |
| Placement | Total Transistors | Normalized RMSE (%) | Residual Error Rate (%) | Absolute Bias | Maximum Error | Full Recovery | OR Saturation |
|---|---|---|---|---|---|---|---|
| NONE | 74 | 4.0493 | 12.109% | 0.250000 | 4 | 0 | 0 |
| BIT0 | 94 | 2.1684 | 12.109% | 0.128906 | 3 | 0 | 0 |
| OR_BIT1 (Proposed) | 92 | 2.1960 | 2.734% | 0.062500 | 4 | 24 | 7 |
| OR_BIT2 | 92 | 4.0493 | 12.109% | 0.000000 | 4 | 0 | 15 |
| OR_BIT3 | 92 | 8.3910 | 12.109% | 0.187500 | 6 | 0 | 17 |
| OR_OUTPUT4 | 92 | 26.9495 | 12.109% | 1.687500 | 14 | 0 | 0 |
| Architecture | Total Transistors | RMSE (%) | Absolute Bias | Maximum Error | Role |
|---|---|---|---|---|---|
| Exact FULL4FA | 156 | 0.00 | 0.00 | 0 | Exact endpoint |
| Baseline 3FA | 108 | 2.17 | 0.13 | 3 | Baseline |
| Code 7 NONE | 74 | 4.05 | 0.25 | 4 | Minimum-cost endpoint |
| Code 7 BIT0 | 94 | 2.17 | 0.13 | 3 | Minimum-RMSE endpoint |
| Proposed: Code 7 + OR_BIT1 | 92 | 2.20 | 0.06 | 4 | Selected design point |
| Configuration | Error ε | Analytical Cost (T) * | Normalized RMSE (%) | Residual Error Rate | Absolute Bias | Maximum Error | Full Recovery | OR Saturation |
|---|---|---|---|---|---|---|---|---|
| Code 6 + OR_BIT1 | −3 | 104 | 2.3550 | 12.109% | 0.132812 | 4 | 0 | 0 |
| Code 7 + NONE | −2 | 74 | 4.0493 | 12.109% | 0.250000 | 4 | 0 | 0 |
| Code 7 + BIT0 | −2 | 94 | 2.1684 | 12.109% | 0.128906 | 3 | 0 | 0 |
| Code 7 + OR_BIT1 (Proposed) | −2 | 92 | 2.1960 | 2.734% | 0.062500 | 4 | 24 | 7 |
| Code 8 + NONE | −1 | 110 | 2.0246 | 12.109% | 0.125000 | 2 | 0 | 0 |
| Code 8 + BIT0 | −1 | 122 | 0.3472 | 0.391% | 0.003906 | 1 | 30 | 0 |
| Code 8 + OR_BIT1 | −1 | 120 | 1.9018 | 11.719% | 0.117188 | 1 | 1 | 0 |
| Design | Transistor Count | Normalized Area | Average Power (µW) | Normalized Power | Energy (fJ) | Worst-Case Propagation Delay (ps) |
|---|---|---|---|---|---|---|
| Exact | 156 | 1.000 | 1.71 | 1.00 | 546.40 | 75.457 |
| Baseline | 108 | 0.692 | 1.40 | 0.82 | 449.33 | 46.589 |
| Proposed | 92 | 0.590 | 1.00 | 0.59 | 321.00 | 54.602 |
| PVT Condition | Baseline Worst Delay (ps) | Proposed Worst Delay (ps) | Proposed vs. Baseline |
|---|---|---|---|
| SSF/0.9 V/125 °C | 64.135 | 76.108 | +18.67% |
| SSF/0.9 V/−40 °C | 84.754 | 98.568 | +16.30% |
| TT/1.0 V/25 °C | 46.589 | 54.602 | +17.20% |
| FFF/1.1 V/125 °C | 33.279 | 38.729 | +16.38% |
| FFF/1.1 V/−40 °C | 34.781 | 40.443 | +16.28% |
| Dataset | Model | Baseline | DIMC-S-Derived | LSAC OR+SXAFA-Derived | Proposed |
|---|---|---|---|---|---|
| CIFAR-10 | ResNet18 | 3.321/4.787 | 10.632/9.205 | 6.896/6.165 | 3.230/4.634 |
| CIFAR-10 | VGG16-BN | 2.940/5.333 | 14.684/13.529 | 6.906/5.648 | 2.856/5.086 |
| CIFAR-10 | AlexNet | 3.106/5.144 | 7.913/13.054 | 4.218/7.066 | 3.019/5.018 |
| CIFAR-100 | ResNet18 | 3.135/9.299 | 9.672/14.045 | 6.318/9.892 | 3.050/9.087 |
| CIFAR-100 | VGG16-BN | 3.185/7.591 | 13.782/19.854 | 7.466/11.050 | 3.094/7.321 |
| CIFAR-100 | AlexNet | 5.750/9.589 | 14.870/18.587 | 10.293/11.376 | 5.636/9.225 |
| Dataset | Model | Exceptional MUL (%) | Flagged Pair (%) | Full Recovery/Flag (%) | OR Saturation/Flag (%) | Residual/Pair (%) |
|---|---|---|---|---|---|---|
| CIFAR-10 | ResNet18 | 0.427 | 0.842 | 87.324 | 12.676 | 0.107 |
| CIFAR-10 | VGG16-BN | 0.368 | 0.726 | 88.528 | 11.472 | 0.083 |
| CIFAR-10 | AlexNet | 0.638 | 1.254 | 85.715 | 14.285 | 0.179 |
| CIFAR-100 | ResNet18 | 0.427 | 0.841 | 87.087 | 12.913 | 0.109 |
| CIFAR-100 | VGG16-BN | 0.405 | 0.799 | 87.860 | 12.140 | 0.097 |
| CIFAR-100 | AlexNet | 0.584 | 1.128 | 80.503 | 19.497 | 0.220 |
| Dataset | Model | Baseline (%) | DIMC-S-Derived (%) | LSAC OR+SXAFA-Derived (%) | Proposed (%) |
|---|---|---|---|---|---|
| CIFAR-10 | ResNet18 | 94.72 | 94.58 | 94.72 | 94.74 |
| CIFAR-10 | VGG16-BN | 93.22 | 92.57 | 93.12 | 93.26 |
| CIFAR-10 | AlexNet | 88.55 | 88.16 | 88.29 | 88.59 |
| CIFAR-100 | ResNet18 | 78.95 | 78.48 | 78.77 | 78.96 |
| CIFAR-100 | VGG16-BN | 74.67 | 73.70 | 74.37 | 74.79 |
| CIFAR-100 | AlexNet | 68.58 | 64.20 | 66.86 | 68.56 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jin, Y.; Kim, M. Joint Multiplier–Adder Approximation with Flag-Based Error Recovery for BF16 Digital Compute-in-Memory. Electronics 2026, 15, 3861. https://doi.org/10.3390/electronics15173861
Jin Y, Kim M. Joint Multiplier–Adder Approximation with Flag-Based Error Recovery for BF16 Digital Compute-in-Memory. Electronics. 2026; 15(17):3861. https://doi.org/10.3390/electronics15173861
Chicago/Turabian StyleJin, Yuhyeon, and Munhyeon Kim. 2026. "Joint Multiplier–Adder Approximation with Flag-Based Error Recovery for BF16 Digital Compute-in-Memory" Electronics 15, no. 17: 3861. https://doi.org/10.3390/electronics15173861
APA StyleJin, Y., & Kim, M. (2026). Joint Multiplier–Adder Approximation with Flag-Based Error Recovery for BF16 Digital Compute-in-Memory. Electronics, 15(17), 3861. https://doi.org/10.3390/electronics15173861
