Thermal Nonuniformity-Aware Reliability Screening for Systolic AI Accelerators
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors
This article focuses on a practical application issue: silent data corruption caused by localized thermal hot spots due to workload in systolic AI accelerators. The authors also propose a cross-layer screening methodology that connects workload, dataflow, power concentration, thermal imbalance, path vulnerability, and final computational errors to identify high-risk scenarios in early design stages. The viewpoints presented in this article are constructive and worth referencing in the development process of systolic AI accelerators. However, I have the following comments:
- The authors proposed a five-stage processing flow for the cross-layer screening methodology in Section 3, conducted experiments on a 16 × 16 systolic array in Section 4, and presented the analysis results in Section 5. However, the authors’ discussion of the methodology is mainly conceptual and lacks technical explanations (e.g., whether computational methods or manual processing are used, whether any EDA tools are employed, etc.). I suggest the authors make arrangements in Sections 3/4/5 to provide such explanations.
- Following up on the previous comment, especially regarding the path-aware vulnerability modeling approach, please clarify how the authors' solution distinguishes between different hardware blocks and PC1~PC4 (algorithmic or user-defined?).
- Although this article mentions the insufficiencies of past literature as well as the practical perspectives and contributions proposed by this work, the content is mainly descriptive. I suggest that the authors provide a comparison table between their proposed solution and previous literature in the later sections of the article. In addition to qualitative explanations, it would be even better if the author could include quantitative data from the experiments and analyses in Sections 4 and 5 and compare them with the solutions in the previous literature.
Author Response
- Clarification of Methodology and Tools
Reviewer Comment: The methodology is mainly conceptual and lacks technical explanations (e.g., whether computational methods or manual processing are used, whether EDA tools are employed).
Response: We thank the reviewer for pointing this out. We have entirely revised Section 3 to clarify that the screening flow operates as a computational pipeline. To make the methodology technically explicit, we have replaced the prose descriptions with exact mathematical formulas (Equations 1–8), detailing the activity extraction, power proxy, iterative thermal diffusion, and stress-to-probability mappings. We also clarified that this is a lightweight screening proxy, explicitly not requiring full EDA tools at this early stage. - Path-Aware Vulnerability Modeling (PC1–PC4)
Reviewer Comment: Clarify how the solution distinguishes between different hardware blocks and PC1~PC4.
Response: We have expanded Section 3.6 to explicitly define the path-class taxonomy. We clarify that the classes (MAC datapath, accumulator update, forwarding path, and control logic) are derived from the functional structure of a synthesizable systolic processing element. We also added Table 2 to explicitly map these hardware logic roles to their corruption interpretations and vulnerability settings. - Comparison Table and Quantitative Data
Reviewer Comment: Provide a comparison table between the proposed solution and previous literature, including quantitative data.
Response: We have added a comprehensive comparison table (Table 1) in Section 2.5, which contrasts our approach against compact thermal modeling (e.g., HotSpot), RTL optimization, and DNN fault-injection studies. Furthermore, to provide quantitative alignment data, we added Section 5.6, which contains a preliminary alignment study (Table 5) comparing our thermal proxy's spatial hotspot correlation directly against a compact reference-style thermal model.
Reviewer 2 Report
Comments and Suggestions for Authors
Dear Author,
This manuscript addresses the connection between workload-dependent thermal nonuniformity and reliability behavior in systolic AI accelerators operating under tight design margins. The cross-layer perspective that links activity, thermal concentration, and path-aware corruption is conceptually interesting and well aligned with current concerns in low-power AI hardware design. The paper is well structured and clearly written, and the positioning of the work as an early-stage screening methodology is appropriate for the level of abstraction adopted.
However, several aspects of the manuscript would benefit from
substantial strengthening before the work can be accepted for
publication. My main suggestions are summarized below.
1. Findings appear to follow directly from modeling assumptions
A major concern is that the main reported findings may follow as
direct consequences of the modeling choices rather than as empirical
discoveries. For example, sparse workloads are constructed to suppress portions of the activity map (Section 3.1), output-stationary execution is defined to create stronger residency concentration (Section 3.1), and accumulator paths are explicitly assigned the highest vulnerability weight (Section 3.5). The subsequent findings that sparse and output-stationary cases exhibit the strongest thermal nonuniformity, and that accumulator behavior dominates corruption propagation, therefore risk being read as restatements of these assumptions. The author is encouraged to clarify which observations represent genuine emergent behavior of the framework, as opposed to direct consequences of its construction.
2. Mathematical formalization of the proxy models
While the proxy nature of the framework is appropriate for early-stage screening, Section 3 currently relies entirely on prose. The inclusion of explicit equations for the power proxy (Section 3.2), the
diffusion-based thermal proxy (Section 3.3), and the stress-to-probability mapping in the vulnerability model (Section 3.5) would substantially improve clarity and reproducibility without altering the screening-level character of the work.
3. Missing physical link between thermal stress and corruption
The framework treats thermal stress as a direct cause of corruption, but heat does not flip bits on its own. In practice, temperature affects circuits indirectly: it slows transistors, increases propagation delay, and reduces the available timing margin, which is what ultimately produces corruption.
The current model skips this chain and jumps from "stress above a threshold" to "corruption probability." Because the timing or voltage margin step is missing, the input could be labeled as any spatial stress quantity, which weakens the "thermal-aware" claim of the framework. References [6] and [7] describe exactly this thermal-to-timing pathway but are not integrated into the model.
A brief, even simplified, treatment of how temperature is assumed to influence timing margins would strengthen the methodological foundation.
4. Validation or calibration against established tools
The work would be considerably strengthened by even a limited
cross-comparison between the proposed thermal proxy and an established
compact thermal model such as HotSpot, or between the path-aware
vulnerability assumptions and RTL-level fault injection on a
representative tile. The author acknowledges this gap in Section 8,
but at least a preliminary alignment study would increase confidence
in the reported trends.
5. Statistical treatment of the experimental results
The reported corruption rates in Figure 5 all fall on multiples of
12.5 percent, suggesting a small number of trials per condition. The
author is encouraged to report the number of trials, the handling of
random seeds, and confidence intervals. This is particularly relevant
for the sparse workload under weight-stationary execution, where the
observed 75 percent silent-corruption rate appears to deviate from
the broader trend that output-stationary execution produces stronger
thermal concentration.
6. Transparency of the calibrated stress conditions
The notion of "calibrated stress conditions" introduced in Section 4
is open to the interpretation that parameters were tuned to produce
specific outcomes. The author is encouraged to describe the
calibration procedure explicitly and to include a brief sensitivity
analysis showing how results change under variation of the calibration
parameters.
7. Array size and dataflow coverage
The experiments are conducted on a 16×16 array, which is considerably
smaller than production-scale systolic arrays (for example, the
256×256 array of TPU v1). Since thermal coupling and edge effects
scale nonlinearly with array size, the generality of the reported
trends would benefit from at least one larger-scale comparison.
Similarly, the evaluation is restricted to weight-stationary and
output-stationary dataflows, while the cited Eyeriss work [3] uses
row-stationary execution. A short discussion of how the framework is
expected to extend to other dataflows would broaden the contribution.
8. Application-level impact of silent corruption
The reported silent corruption rates are not connected to
application-level metrics such as model accuracy. Recent ML
reliability work (including [10] and [19], already cited by the
author) typically evaluates whether hardware-level corruption
translates into meaningful accuracy degradation, since DNN inference
exhibits inherent robustness to certain perturbations. An indication
of how the reported corruption rates relate to inference-level
outcomes would significantly strengthen the practical relevance of
the findings.
9. Interpretation of the thermal results
In Figure 3, the peak temperature across all cases is nearly constant
(between 83.48 and 84.90 degrees Celsius), while the thermal spread
varies by approximately a factor of four. The statement in Section 6.2
that workload-driven thermal concentration creates a "hotter fabric"
may benefit from being rephrased, since the data indicate a more
uneven rather than a hotter fabric.
10. Specification of experimental conditions
To improve reproducibility, the matrix dimensions of the GEMM
workloads, the sparsity level and pattern of the sparse workload, and
the numerical range of the low-dynamic-range workload should be
specified in Section 4.
11. Clarification of Figure 4
Figure 4 visualizes modeling assumptions about path-role priority
rather than measured outcomes. This should be stated more explicitly
in both the caption and the body.
Minor suggestions
The use of "We first examine..." in Section 5.1 may be revised for
consistency with the single-author nature of the manuscript. The
arXiv identifier of Reference [20] should be verified.
Closing remarks
Overall, this work explores a topic of growing relevance to low-power
AI accelerator design, and the cross-layer framing has clear value as
an early-stage exploration tool. The suggestions above are intended
to help the author strengthen the rigor, reproducibility, and
practical relevance of the framework so that it can serve as a more
solid foundation for the deeper validation steps outlined in the
future work section.
Importantly, I would like to emphasize that the revised manuscript
should be supported by additional experimental work rather than by
textual clarification alone. In particular, the concerns regarding
the lack of validation against established tools, the statistical
reliability of the reported corruption rates, the scalability of the
framework beyond the 16x16 case, and the application-level relevance
of the silent corruption outcomes cannot be adequately addressed
through rewording or expanded discussion. A substantive revision
should include, at minimum, a preliminary cross-comparison with an
established thermal or fault-injection reference, an increased number
of trials with reported confidence intervals, and at least one
additional array configuration to demonstrate the generality of the
findings.
I look forward to seeing a revised version of the manuscript that
reflects this experimental strengthening.
Best regards,
Author Response
Please see the attachment
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for Authors
The paper proposes a thermal-nonuniformity-aware reliability screening methodology for systolic AI accelerators. The topic is relevant and interesting, but there are several concerns that need to be resolved. More specifically, my comments for the authors are:
- The paper says accumulator paths are most vulnerable and forwarding paths moderately vulnerable, but no STA, circuit timing, or empirical calibration is provided.
- Authors should expand related work and provide some discussion regarding comparisons with other approaches such as HotSpot (HotSpot: a dynamic compact thermal model at the processor-architecture level), RTL power analysis (Roelke, Alec, et al. "Pre-RTL voltage and power optimization for low-cost, thermally challenged multicore chips." 2017 IEEE International Conference on Computer Design (ICCD). IEEE, 2017.), or implementation-grade thermal tools (Floros, George, Nestor Evmorfopoulos, and George Stamoulis. "Efficient IC hotspot thermal analysis via low-rank model order reduction." Integration 66 (2019): 1-8 and Ladenheim, Scott, et al. "The mta: An advanced and versatile thermal simulator for integrated systems." IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 37.12 (2018): 3123-3136).
- The main results focus on a 16 × 16 array and only three synthetic GEMM workload classes. This is too narrow for broad claims about AI accelerators.
- Package assumptions, ambient temperature, cooling model, floorplan, power density, tile size, and boundary conditions are not sufficiently specified.
- The work would be stronger with CNN, Transformer, or recommendation-model workloads rather than only dense, low-DR, and sparse GEMM abstractions.
- The exact probability functions, thresholds, perturbation magnitudes, tolerances, and randomization methodology are not clearly defined and authors should elaborate more on these aspects.
- An ablation study would be interesting. In the current form it is unclear how much each component, activity model, residency cost, relay burden, diffusion model, path weighting, contributes to the final outcomes.
- The control/update path is listed in the taxonomy but excluded from the active model
- Sparse workloads show much larger thermal spread but similar peak temperatures, which is a strange result that needs stronger explanation.
Author Response
Please see the attachement
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for Authors
The authors' revised manuscript has addressed my comments.
Reviewer 3 Report
Comments and Suggestions for Authors
All comments have beed addressed. I suggest acceptance.
