Next Article in Journal
Design of an X-Band CMOS VCO with a Transformer-Coupled and Transconductance-Boosted Stacked Topology
Previous Article in Journal
A Low-Power 68.4 dB Signal-to-Noise-and-Distortion Ratio Noise-Shaping SAR ADC for Biomedical Applications
 
 
Article
Peer-Review Record

Thermal Nonuniformity-Aware Reliability Screening for Systolic AI Accelerators

J. Low Power Electron. Appl. 2026, 16(2), 18; https://doi.org/10.3390/jlpea16020018
by Larisa Goffman-Vinopal
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Reviewer 3: Anonymous
J. Low Power Electron. Appl. 2026, 16(2), 18; https://doi.org/10.3390/jlpea16020018
Submission received: 19 April 2026 / Revised: 24 May 2026 / Accepted: 29 May 2026 / Published: 31 May 2026
(This article belongs to the Topic Advanced Integrated Circuit Design and Application)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

This article focuses on a practical application issue: silent data corruption caused by localized thermal hot spots due to workload in systolic AI accelerators. The authors also propose a cross-layer screening methodology that connects workload, dataflow, power concentration, thermal imbalance, path vulnerability, and final computational errors to identify high-risk scenarios in early design stages. The viewpoints presented in this article are constructive and worth referencing in the development process of systolic AI accelerators. However, I have the following comments:

  1. The authors proposed a five-stage processing flow for the cross-layer screening methodology  in Section 3, conducted experiments on a 16 × 16 systolic array in Section 4, and presented the analysis results in Section 5. However, the authors’ discussion of the methodology is mainly conceptual and lacks technical explanations (e.g., whether computational methods or manual processing are used, whether any EDA tools are employed, etc.). I suggest the authors make arrangements in Sections 3/4/5 to provide such explanations.
  2. Following up on the previous comment, especially regarding the path-aware vulnerability modeling approach, please clarify how the authors' solution distinguishes between different hardware blocks and PC1~PC4 (algorithmic or user-defined?).
  3. Although this article mentions the insufficiencies of past literature as well as the practical perspectives and contributions proposed by this work, the content is mainly descriptive. I suggest that the authors provide a comparison table between their proposed solution and previous literature in the later sections of the article. In addition to qualitative explanations, it would be even better if the author could include quantitative data from the experiments and analyses in Sections 4 and 5 and compare them with the solutions in the previous literature.

Author Response

  1. Clarification of Methodology and Tools
    Reviewer Comment: The methodology is mainly conceptual and lacks technical explanations (e.g., whether computational methods or manual processing are used, whether EDA tools are employed).
    Response: We thank the reviewer for pointing this out. We have entirely revised Section 3 to clarify that the screening flow operates as a computational pipeline. To make the methodology technically explicit, we have replaced the prose descriptions with exact mathematical formulas (Equations 1–8), detailing the activity extraction, power proxy, iterative thermal diffusion, and stress-to-probability mappings. We also clarified that this is a lightweight screening proxy, explicitly not requiring full EDA tools at this early stage.
  2. Path-Aware Vulnerability Modeling (PC1–PC4)
    Reviewer Comment: Clarify how the solution distinguishes between different hardware blocks and PC1~PC4.
    Response: We have expanded Section 3.6 to explicitly define the path-class taxonomy. We clarify that the classes (MAC datapath, accumulator update, forwarding path, and control logic) are derived from the functional structure of a synthesizable systolic processing element. We also added Table 2 to explicitly map these hardware logic roles to their corruption interpretations and vulnerability settings.
  3. Comparison Table and Quantitative Data
    Reviewer Comment: Provide a comparison table between the proposed solution and previous literature, including quantitative data.
    Response: We have added a comprehensive comparison table (Table 1) in Section 2.5, which contrasts our approach against compact thermal modeling (e.g., HotSpot), RTL optimization, and DNN fault-injection studies. Furthermore, to provide quantitative alignment data, we added Section 5.6, which contains a preliminary alignment study (Table 5) comparing our thermal proxy's spatial hotspot correlation directly against a compact reference-style thermal model.

Reviewer 2 Report

Comments and Suggestions for Authors

Dear Author,

This manuscript addresses the connection between workload-dependent thermal nonuniformity and reliability behavior in systolic AI accelerators operating under tight design margins. The cross-layer perspective that links activity, thermal concentration, and path-aware corruption is conceptually interesting and well aligned with current concerns in low-power AI hardware design. The paper is well structured and clearly written, and the positioning of the work as an early-stage screening methodology is appropriate for the level of abstraction adopted.

However, several aspects of the manuscript would benefit from 
substantial strengthening before the work can be accepted for 
publication. My main suggestions are summarized below.

1. Findings appear to follow directly from modeling assumptions

A major concern is that the main reported findings may follow as 
direct consequences of the modeling choices rather than as empirical 
discoveries. For example, sparse workloads are constructed to suppress portions of the activity map (Section 3.1), output-stationary execution is defined to create stronger residency concentration (Section 3.1),  and accumulator paths are explicitly assigned the highest vulnerability weight (Section 3.5). The subsequent findings that sparse and output-stationary cases exhibit the strongest thermal nonuniformity, and that accumulator behavior dominates corruption propagation, therefore risk being read as restatements of these assumptions. The author is encouraged to clarify which observations represent genuine emergent behavior of the framework, as opposed to direct consequences of its construction.

2. Mathematical formalization of the proxy models

While the proxy nature of the framework is appropriate for early-stage screening, Section 3 currently relies entirely on prose. The inclusion of explicit equations for the power proxy (Section 3.2), the 
diffusion-based thermal proxy (Section 3.3), and the stress-to-probability mapping in the vulnerability model (Section 3.5) would substantially improve clarity and reproducibility without altering the screening-level character of the work.

3. Missing physical link between thermal stress and corruption

The framework treats thermal stress as a direct cause of corruption, but heat does not flip bits on its own. In practice, temperature affects circuits indirectly: it slows transistors, increases propagation delay, and reduces the available timing margin, which is what ultimately produces corruption.
The current model skips this chain and jumps from "stress above a threshold" to "corruption probability." Because the timing or voltage margin step is missing, the input could be labeled as any spatial stress quantity, which weakens the "thermal-aware" claim of the framework. References [6] and [7] describe exactly this thermal-to-timing pathway but are not integrated into the model.
A brief, even simplified, treatment of how temperature is assumed to influence timing margins would strengthen the methodological foundation.

4. Validation or calibration against established tools

The work would be considerably strengthened by even a limited 
cross-comparison between the proposed thermal proxy and an established 
compact thermal model such as HotSpot, or between the path-aware 
vulnerability assumptions and RTL-level fault injection on a 
representative tile. The author acknowledges this gap in Section 8, 
but at least a preliminary alignment study would increase confidence 
in the reported trends.

5. Statistical treatment of the experimental results

The reported corruption rates in Figure 5 all fall on multiples of 
12.5 percent, suggesting a small number of trials per condition. The 
author is encouraged to report the number of trials, the handling of 
random seeds, and confidence intervals. This is particularly relevant 
for the sparse workload under weight-stationary execution, where the 
observed 75 percent silent-corruption rate appears to deviate from 
the broader trend that output-stationary execution produces stronger 
thermal concentration.

6. Transparency of the calibrated stress conditions

The notion of "calibrated stress conditions" introduced in Section 4 
is open to the interpretation that parameters were tuned to produce 
specific outcomes. The author is encouraged to describe the 
calibration procedure explicitly and to include a brief sensitivity 
analysis showing how results change under variation of the calibration 
parameters.

7. Array size and dataflow coverage

The experiments are conducted on a 16×16 array, which is considerably 
smaller than production-scale systolic arrays (for example, the 
256×256 array of TPU v1). Since thermal coupling and edge effects 
scale nonlinearly with array size, the generality of the reported 
trends would benefit from at least one larger-scale comparison. 
Similarly, the evaluation is restricted to weight-stationary and 
output-stationary dataflows, while the cited Eyeriss work [3] uses 
row-stationary execution. A short discussion of how the framework is 
expected to extend to other dataflows would broaden the contribution.

8. Application-level impact of silent corruption

The reported silent corruption rates are not connected to 
application-level metrics such as model accuracy. Recent ML 
reliability work (including [10] and [19], already cited by the 
author) typically evaluates whether hardware-level corruption 
translates into meaningful accuracy degradation, since DNN inference 
exhibits inherent robustness to certain perturbations. An indication 
of how the reported corruption rates relate to inference-level 
outcomes would significantly strengthen the practical relevance of 
the findings.

9. Interpretation of the thermal results

In Figure 3, the peak temperature across all cases is nearly constant 
(between 83.48 and 84.90 degrees Celsius), while the thermal spread 
varies by approximately a factor of four. The statement in Section 6.2 
that workload-driven thermal concentration creates a "hotter fabric" 
may benefit from being rephrased, since the data indicate a more 
uneven rather than a hotter fabric.

10. Specification of experimental conditions

To improve reproducibility, the matrix dimensions of the GEMM 
workloads, the sparsity level and pattern of the sparse workload, and 
the numerical range of the low-dynamic-range workload should be 
specified in Section 4.

11. Clarification of Figure 4

Figure 4 visualizes modeling assumptions about path-role priority 
rather than measured outcomes. This should be stated more explicitly 
in both the caption and the body.

Minor suggestions

The use of "We first examine..." in Section 5.1 may be revised for 
consistency with the single-author nature of the manuscript. The 
arXiv identifier of Reference [20] should be verified.

Closing remarks

Overall, this work explores a topic of growing relevance to low-power 
AI accelerator design, and the cross-layer framing has clear value as 
an early-stage exploration tool. The suggestions above are intended 
to help the author strengthen the rigor, reproducibility, and 
practical relevance of the framework so that it can serve as a more 
solid foundation for the deeper validation steps outlined in the 
future work section.

Importantly, I would like to emphasize that the revised manuscript 
should be supported by additional experimental work rather than by 
textual clarification alone. In particular, the concerns regarding 
the lack of validation against established tools, the statistical 
reliability of the reported corruption rates, the scalability of the 
framework beyond the 16x16 case, and the application-level relevance 
of the silent corruption outcomes cannot be adequately addressed 
through rewording or expanded discussion. A substantive revision 
should include, at minimum, a preliminary cross-comparison with an 
established thermal or fault-injection reference, an increased number 
of trials with reported confidence intervals, and at least one 
additional array configuration to demonstrate the generality of the 
findings.

I look forward to seeing a revised version of the manuscript that 
reflects this experimental strengthening.

Best regards,

Author Response

Please see the attachment

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

The paper proposes a thermal-nonuniformity-aware reliability screening methodology for systolic AI accelerators. The topic is relevant and interesting, but there are several concerns that need to be resolved. More specifically, my comments for the authors are: 

  • The paper says accumulator paths are most vulnerable and forwarding paths moderately vulnerable, but no STA, circuit timing, or empirical calibration is provided.
  • Authors should expand related work and provide some discussion regarding comparisons with other approaches such as HotSpot (HotSpot: a dynamic compact thermal model at the processor-architecture level), RTL power analysis (Roelke, Alec, et al. "Pre-RTL voltage and power optimization for low-cost, thermally challenged multicore chips." 2017 IEEE International Conference on Computer Design (ICCD). IEEE, 2017.), or implementation-grade thermal tools (Floros, George, Nestor Evmorfopoulos, and George Stamoulis. "Efficient IC hotspot thermal analysis via low-rank model order reduction." Integration 66 (2019): 1-8 and Ladenheim, Scott, et al. "The mta: An advanced and versatile thermal simulator for integrated systems." IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 37.12 (2018): 3123-3136). 
  • The main results focus on a 16 × 16 array and only three synthetic GEMM workload classes. This is too narrow for broad claims about AI accelerators. 
  • Package assumptions, ambient temperature, cooling model, floorplan, power density, tile size, and boundary conditions are not sufficiently specified. 
  • The work would be stronger with CNN, Transformer, or recommendation-model workloads rather than only dense, low-DR, and sparse GEMM abstractions. 
  • The exact probability functions, thresholds, perturbation magnitudes, tolerances, and randomization methodology are not clearly defined and authors should elaborate more on these aspects.
  • An ablation study would be interesting. In the current form it is unclear how much each component, activity model, residency cost, relay burden, diffusion model, path weighting, contributes to the final outcomes. 
  • The control/update path is listed in the taxonomy but excluded from the active model
  • Sparse workloads show much larger thermal spread but similar peak temperatures, which is a strange result that needs stronger explanation. 

Author Response

Please see the attachement

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

The authors' revised manuscript has addressed my comments.

Reviewer 3 Report

Comments and Suggestions for Authors

All comments have beed addressed. I suggest acceptance.

Back to TopTop