Human Capital Sustainability Leadership and Employee Job Performance in Public Asset Management: The Mediating Role of Innovative Work Behavior Among Indonesian Civil Servants
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript investigates examines the influence of Human Capital Sustainability Leadership (HCSL) on Employee Job Performance (EJP) and the mediating role of Innovative Work Behavior (IWB) among Indonesian civil servants managing regional government assets (Barang Milik Daerah). Based on Self-Determination Theory, the study uses survey data from 541 civil servants and applies PLS-SEM to test the proposed mediation model. Findings underscore sustainable leadership as a strategic lever for fostering both innovation and performance in public asset stewardship, with implications for sustainable governance and human capital development.
Overall, I am supportive of publication, provided that the following issues are carefully addressed:
- Section 2.2 Human Capital Sustainability Leadership, Lines 97–110, where HCSL is defined as a higher-order leadership construct integrating four dimensions; and Section 3.3 Instruments, Lines 211–223, where the measurement items for HCSL, IWB, and EJP are described. The relevant constructs are described, but the manuscript does not clearly explain whether HCSL was treated as a first-order reflective construct or modeled as a higher-order construct in the PLS-SEM analysis. The authors are advised to further clarify the measurement model specification and the operationalization of HCSL as a higher-order construct.
- Section 3.2, Lines 200–210: This study adopted purposive sampling and obtained 541 valid responses from civil servants involved in regional government asset management. Please provide more details on the sampling procedure, such as how respondents were identified, the total number of questionnaires distributed, and the response rate.
- Section 4.2.1, Lines 287–296: The Cronbach’s α and composite reliability values of all constructs are extremely high, ranging from 0.989 to 0.991. Such high reliability values may indicate item redundancy or insufficient construct breadth. The authors are advised to explain whether item reduction, item parceling, or dimension-level analysis was considered to address this issue.
- Please carefully check the formatting of the references. Some reference entries appear to be merged with adjacent references, for example, Reference 13.
Author Response
Manuscript ID: sustainability-4339306
Human Capital Sustainability Leadership and Employee Job Performance in Public Asset Management: The Mediating Role of Innovative Work Behavior Among Indonesian Civil Servants
We sincerely thank Reviewer 1 for the supportive assessment of our manuscript and for the careful, constructive comments. We have addressed each point in the revised manuscript; the corresponding changes are marked using tracked changes. Our point-by-point responses are provided below.
Comment 1:
Section 2.2 Human Capital Sustainability Leadership, Lines 97–110, where HCSL is defined as a higher-order leadership construct integrating four dimensions; and Section 3.3 Instruments, Lines 211–223, where the measurement items for HCSL, IWB, and EJP are described. The relevant constructs are described, but the manuscript does not clearly explain whether HCSL was treated as a first-order reflective construct or modeled as a higher-order construct in the PLS-SEM analysis. The authors are advised to further clarify the measurement model specification and the operationalization of HCSL as a higher-order construct.
Authors' Response:
We thank the reviewer for this helpful observation. HCSL was modeled as a reflective higher-order construct and estimated using the two-stage approach (Sarstedt et al., 2019): the four first-order dimensions—ethical, mindful, sustainable, and servant leadership—were measured reflectively and then used to form the higher-order HCSL construct. To make this explicit, we have (i) added a dedicated “Instrument Development and Validation” passage to Section 3.3 that states the higher-order specification and the two-stage estimation procedure; (ii) clarified the higher-order modeling in the measurement-model description (Section 4.2); and (iii) revised Figures 1 and 2 so that the four first-order dimensions and their relationship to the higher-order HCSL construct are shown visually.
These changes appear in Section 3.3, Section 4.2, and Figures 1–2 of the revised manuscript.
Comment 2:
Section 3.2, Lines 200–210: This study adopted purposive sampling and obtained 541 valid responses from civil servants involved in regional government asset management. Please provide more details on the sampling procedure, such as how respondents were identified, the total number of questionnaires distributed, and the response rate.
Authors' Response:
We appreciate this request for additional procedural detail, and we have expanded Section 3.2 accordingly. Eligible respondents were identified from the official personnel rosters of units responsible for BMD management across the Provincial Government’s SKPD and UKPD and were required to have at least one year of direct BMD involvement. The questionnaire was disseminated through official BMD coordinator channels, and a total of 614 responses were received; after screening against the eligibility criteria and for completeness, 73 responses were excluded, yielding 541 valid responses (a valid-response rate of 88.1%).
Because distribution was carried out through coordinator channels rather than a fixed, enumerated mailing list, we now state transparently that 614 denotes the number of submissions received rather than a denominator of invitations sent. In addition, we have added the full demographic profile of the sample (gender, age, education, tenure, position, work unit, and BMD experience) in Table A1.
These revisions appear in Section 3.2 and Table A1 of the revised manuscript.
Comment 3:
Section 4.2.1, Lines 287–296: The Cronbach’s α and composite reliability values of all constructs are extremely high, ranging from 0.989 to 0.991. Such high reliability values may indicate item redundancy or insufficient construct breadth. The authors are advised to explain whether item reduction, item parceling, or dimension-level analysis was considered to address this issue.
Authors' Response:
We thank the reviewer for raising this important point, with which we agree. We have substantially revised our treatment of the reliability values in Section 4.2.1. We no longer present them as a strength; instead, we explicitly acknowledge that values exceeding 0.95 constitute a recognized warning signal of indicator redundancy (Hair et al., 2022) and discuss the plausible contributors in this dataset—semantic proximity among the indicators generated to elaborate each dimension, a pronounced ceiling effect (construct means of 4.37–4.45 with a mean item standard deviation below 0.70), sample homogeneity, and potential halo and acquiescence effects associated with single-source self-report.
Regarding the specific remedies the reviewer mentions, we did consider item reduction and dimension-level analysis. A supplementary check indicated that reducing each scale to its originally validated length still yields Cronbach’s α above 0.98 for all three focal constructs, which suggests that the elevated values stem primarily from sample homogeneity and response tendencies rather than from item count alone. Ad hoc item deletion would therefore not resolve the issue and would discard content that was validated through the expert content-review procedure now described in Section 3.3. We accordingly addressed the concern through transparency and critical interpretation rather than post hoc trimming: full indicator-level statistics—loadings, item means, standard deviations, and corrected item–total correlations—for all 109 items are now reported in Appendix A so that readers can assess potential redundancy directly.
These changes appear in Section 4.2.1 and Appendix A of the revised manuscript.
Comment 4:
Please carefully check the formatting of the references. Some reference entries appear to be merged with adjacent references, for example, Reference 13.
Authors' Response:
We thank the reviewer for catching these formatting issues, which we have corrected. We identified and consolidated a duplicate entry: the Government Regulation No. 27/2014 was listed twice (as References 4 and 44). The duplicate has been removed, the corresponding in-text citations have been updated, and the reference list has been renumbered accordingly. We also removed inconsistent list formatting affecting several entries—including those adjacent to Reference 13, which produced the merged appearance the reviewer noted—so that all entries follow a uniform MDPI reference style. The reference list has been checked in full for consistency.
These corrections appear throughout the References section of the revised manuscript.
We are grateful for the reviewer’s time and constructive guidance, which have improved the clarity and rigor of the manuscript.
Reviewer 2 Report
Comments and Suggestions for Authors<Overall Evaluation>
This paper analyzes the relationships between Human Capital Sustainability Leadership (HCSL), Innovative Work Behavior (IWB), and Employee Job Performance (EJP) within the context of Indonesian public asset management. The subject of this study is timely as it attempts to address sustainability-oriented leadership within the public sector context, and the topic also generally aligns well with the scope of the Sustainability journal.
However, the current manuscript exhibits significant limitations in terms of theoretical originality, conceptual depth, and methodological rigor. First, the research model itself is based on a very typical leadership → innovation → performance structure and does not go far beyond merely reaffirming relationships that have been repeatedly verified in existing organizational behavior and HRM research. Second, while the paper repeatedly emphasizes sustainability, public governance, and public asset stewardship, it is difficult to view these elements as being sufficiently theoretically integrated within the actual analysis model. In particular, this study fails to sufficiently explain how the context of public asset management extends or alters the existing leadership-performance relationship; consequently, the theoretical contribution claimed by the paper appears somewhat exaggerated compared to the actual scope of the analysis. Finally, despite being based on a self-report questionnaire regarding a single point in time and respondent, very high correlation coefficients, reliability, and explanatory power are reported, raising significant concerns regarding the possibility of common method inflation and construct redundancy. Overall, while this study holds certain potential in terms of application (public-sector applied study), its academic contribution and methodological persuasiveness are insufficient in its current state. In particular, a more rigorous re-examination of conceptual contribution, measurement validity, and causal interpretation is required. Accordingly, the reviewer's opinion is a ‘Major Revision.’
<Major Concerns>
(1) The core model of this study (HCSL → IWB → EJP) is a very typical mediation structure in the fields of organizational behavior and HRM, and there is already extensive research supporting the logic that leadership promotes innovative behavior and that innovative behavior improves performance. While the authors emphasize the context of public asset management and sustainable governance, the actual analysis does not sufficiently explain what unique mechanisms this context provides to the relationships between variables. The current study does not demonstrate significant differentiation from general public organization survey research, and it is difficult to identify theoretical originality or academic contribution beyond contextual relocation.
(2) The paper claims a significant level of theoretical advancement by emphasizing sustainability leadership, SDT mediation logic, and the innovation–performance paradox. However, the actual study remains at the level of a simple cross-sectional mediation model, and there is a significant lack of examination regarding boundary conditions, competing mechanisms, and contextual contingencies. In particular, discussing the “innovation–performance paradox” requires a more sophisticated analysis regarding procedural rigidity or bureaucratic constraints in the public sector, but the current paper merely presents positive path coefficients. Therefore, the paper's contribution claim needs to be adjusted to a more cautious and realistic level.
(3) All variables (HCSL, IWB, EJP) were measured using the same respondents, at the same time point, and with the same survey method. Nevertheless, very high figures are being reported, for example: (a) HCSL → IWB = .834, (b) R²(IWB) = .695, (c) construct correlations ≈ .83, (d) Cronbach’s α ≈ .99. The authors argue that the CMB problem is not serious by utilizing Kock’s full collinearity VIF, but this approach alone is not considered sufficiently persuasive. The current results strongly suggest not only a substantive effect but also the possibility of method inflation, raising significant concerns regarding common method bias. In particular, as recent PLS-SEM and survey methodology literature require stricter verification of single-source perceptual surveys, the authors need to add a more critical discussion on this matter.
(4) The paper acknowledges a borderline violation of the Fornell-Larcker criterion between HCSL and IWB, but interprets it relatively lightly. In particular, the results showing an HCSL–IWB correlation of .834 and a √AVE(HCSL) of .830 suggest the possibility that the two constructs are not empirically distinct. Furthermore, the fact that results at an α ≈ .99 level were obtained despite using an excessively large number of items—HCSL = 48 items, IWB = 26 items, and EJP = 35 items—strongly suggests the possibility of construct redundancy and item overlap. While the current paper interprets this simply as "high confidence," excessively high internal consistency can actually be a signal of measurement redundancy. The authors are advised to conduct a more rigorous review of higher-order construct specification and discriminant validity.
(5) The authors explain that they used PLS-SEM because it is prediction-oriented research. However, the actual study is strongly theory-testing in nature, and the sample size (n=541) is also sufficient. Since the current model is merely a relatively simple mediation structure, it does not sufficiently explain why PLS-SEM was necessary rather than covariance-based SEM (CB-SEM), so the justification for using PLS-SEM appears insufficient.
(6) Although this study is a cross-sectional self-report survey, the paper describes HCSL as “enhance,” “drive,” and “improve” performance throughout. However, with the current design, reverse causality, omitted variables, and reciprocal relationships cannot be excluded. In particular, there is also a possibility of perceptual consistency bias when using self-rated performance. Therefore, the causal language needs to be modified more carefully and restrictively.
(7) The paper repeatedly utilizes sustainability terminology, but the actual analysis is closer to a typical organizational behavior/performance study. For example, direct connections with environmental sustainability, sustainable asset lifecycle, ESG/public accountability, and long-term resource stewardship are relatively weak. In its current state, it gives the impression that sustainability remains at the level of a framing keyword rather than a substantive analytical dimension.
(8) Although the paper repeatedly emphasizes “compliance-intensive public bureaucracy,” structural factors unique to the public sector, such as bureaucratic rigidity, political accountability, and administrative hierarchy, are not substantially reflected in the model during the actual analysis. Therefore, it gives the impression that the public sector context remains at the level of background explanation rather than a substantive explanatory mechanism.
<Minor Concerns>
(1) There is a lack of descriptive statistics regarding the basic sample profile, such as respondents' gender, age, job title, years of service, and department composition.
(2) Despite the use of a very large number of items, information on actual item wording and operationalization was not sufficiently presented.
(3) Levels of α > .98 and CR > .99 were reported in the pilot stage, which may suggest the possibility of redundancy rather than simply “excellent reliability.”
(4) Please review the consistency of citation ordering and formatting throughout the References section more closely.
(5) In particular, the current figures, including Figures 1 and 2, remain at the level of basic mediation diagrams; more professional and informative visualizations are needed.
(6) Overall, results are being restated in the Discussion section. Sections 5.2–5.5, in particular, have a high proportion of repetition in explaining results and need to be condensed into a more analytical discussion.
Author Response
We are grateful to Reviewer 2 for a thorough and demanding review. The concerns regarding theoretical originality, conceptual depth, and methodological rigor have prompted substantial revisions across the manuscript, particularly to its measurement reporting, theoretical positioning, and calibration of claims. We respond to each major and minor concern below; all changes are marked using tracked changes.
Major Concerns
Major Concern 1:
The core model (HCSL→IWB→EJP) is a very typical mediation structure. While the authors emphasize the public asset management context, the analysis does not sufficiently explain what unique mechanisms this context provides. It is difficult to identify theoretical originality beyond contextual relocation.
Authors' Response:
We thank the reviewer, and we have both recalibrated the contribution and made explicit what the context contributes. The Introduction now frames the contribution as empirical and contextual rather than as a new model. Substantively, Section 2.3 specifies the mechanisms the context provides: the predominance of incremental, compliance-reinforcing process innovation in asset administration, and the BerAKHLAK reform agenda that legitimizes adaptiveness—conditions that together govern whether innovation is converted into rated performance. Section 2.4 articulates how asset stewardship links to sustainability in the resource-governance sense. We thus locate the contribution as context-specific evidence and boundary-condition specification rather than theoretical relocation.
These changes appear in Sections 1, 2.3, 2.4, and 5.7 of the revised manuscript.
Major Concern 2:
The study claims theoretical advancement but remains a simple cross-sectional mediation model, with a significant lack of examination of boundary conditions, competing mechanisms, and contextual contingencies. Discussing the innovation–performance paradox requires a more sophisticated analysis of procedural rigidity or bureaucratic constraints. The contribution claim needs to be adjusted to a more cautious and realistic level.
Authors' Response:
We agree and have addressed this directly. Section 2.3 now develops the boundary conditions of the innovation–performance relationship—drawing on procedural rigidity and bureaucratic constraints (Bysted & Jespersen, 2014; Damanpour & Schneider, 2009) and on public service motivation (Perry & Wise, 1990)—rather than merely asserting positive paths. The Discussion (Sections 5.1–5.5) has been rewritten from restatement toward analysis, and the contribution claims in Section 5.7 and the Conclusions have been adjusted to a more cautious and realistic level consistent with a cross-sectional mediation design.
These changes appear in Sections 2.3, 5.1–5.7, and 6 of the revised manuscript.
Major Concern 3:
All variables were measured from the same respondents, at the same time, with the same method, yet very high figures are reported (HCSL→IWB = .834, R²(IWB) = .695, correlations ≈ .83, α ≈ .99). Reliance on Kock’s full collinearity VIF is not sufficiently persuasive. The results suggest the possibility of method inflation; a more critical discussion of common method bias is needed.
Authors' Response:
We thank the reviewer for this central concern, with which we agree, and we have substantially expanded our critical treatment. Section 3.6 now reports Harman’s single-factor test (first factor 62.6%, above the 50% benchmark), reported transparently; we note that Harman’s test is itself insensitive and that, together with the limited reassurance from Kock’s full-collinearity VIF (3.276, close to the 3.3 threshold), the available diagnostics cannot exclude common-method and halo influence. We have removed the claim that CMB is not a serious concern, and we now state that the high path coefficients, correlations, and reliabilities are consistent with some method-induced inflation superimposed on substantive covariation, amplified by a pronounced ceiling effect.
Section 5.9 treats the single-source, single-occasion, self-report design as a material limitation and, in line with current survey-methodology guidance, recommends multi-source data (supervisor or archival performance ratings), temporal separation, and an a priori marker variable for future work. These changes appear in Sections 3.6, 4.2.1, and 5.9 of the revised manuscript.
Major Concern 4:
The borderline Fornell-Larcker violation between HCSL and IWB is interpreted too lightly (correlation .834 vs. √AVE .830), suggesting the constructs may not be empirically distinct. Moreover, α ≈ .99 with an excessive number of items (HCSL = 48, IWB = 26, EJP = 35) strongly suggests construct redundancy and item overlap. A more rigorous review of higher-order construct specification and discriminant validity is advised.
Authors' Response:
We have engaged with each element of this concern. On reliability, Section 4.2.1 now treats values above 0.95 as a warning signal of redundancy (Hair et al., 2022) rather than a strength, reports indicator-level statistics for all 109 items in Appendix A, and notes that reduction to validated scale lengths still yields α above 0.98—indicating the values reflect sample homogeneity rather than item count. On discriminant validity, Section 4.2.3 reports the Fornell-Larcker result transparently as a marginal violation, presents a clean item-level cross-loading analysis (Appendix B), and—together with HTMT below 0.85—concludes that HCSL and IWB are empirically distinct but proximate, with the implication for H1 discussed in Section 5.2. On higher-order specification, Section 3.3 documents the reflective higher-order modeling via the two-stage approach (Sarstedt et al., 2019), now also depicted in Figures 1–2.
These changes appear in Sections 3.3, 4.2.1, 4.2.3, 5.2, Appendices A–B, and Figures 1–2.
Major Concern 5:
The authors justify PLS-SEM as prediction-oriented, but the study is strongly theory-testing, the sample (n=541) is sufficient, and the model is a relatively simple mediation. The justification for using PLS-SEM rather than CB-SEM appears insufficient.
Authors' Response:
We have strengthened the justification in Section 3.7. We clarify that PLS-SEM was chosen for the characteristics of the model and data rather than for sample size (which at 541 is ample for either approach): the reflective higher-order construct estimated via the two-stage approach, the markedly non-normal Likert distributions arising from a ceiling effect, and the study’s predictive aim (assessed via PLSpredict). We also acknowledge explicitly that, as a theory-testing model of moderate complexity, the model could be estimated with CB-SEM, and we do not claim that the conclusions depend on the estimator.
These changes appear in Section 3.7 of the revised manuscript.
Major Concern 6:
Although the study is a cross-sectional self-report survey, the paper describes HCSL as “enhance,” “drive,” and “improve” performance. Reverse causality, omitted variables, and reciprocal relationships cannot be excluded, and there is a possibility of perceptual consistency bias with self-rated performance. The causal language needs to be modified more carefully and restrictively.
Authors' Response:
We agree and have revised the causal language throughout, replacing “enhance,” “drive,” and “improve” with associational terms (“is associated with,” “predicts,” “relates to”). Section 5.9 now states that the cross-sectional design precludes causal inference and that reverse causality, omitted variables, reciprocal relationships, and perceptual-consistency bias in self-rated performance cannot be excluded.
These changes appear in the Abstract, Sections 5.1–5.9, and Section 6 of the revised manuscript.
Major Concern 7:
Sustainability terminology is used repeatedly, but the analysis is closer to a typical organizational behavior/performance study. Direct connections with environmental sustainability, sustainable asset lifecycle, ESG/public accountability, and long-term resource stewardship are weak; sustainability remains a framing keyword rather than a substantive analytical dimension.
Authors' Response:
We thank the reviewer, and we have clarified rather than inflated the sustainability framing. Section 2.4 now makes explicit that “sustainability” in this study operates in two specific senses—the sustainable development of human capital (the leadership construct) and the sustainable stewardship of public resources (the asset-management performance domain, concerning the efficient, accountable, long-horizon management of publicly owned resources)—rather than in the environmental sense. We do not claim to address environmental sustainability, ESG, or the physical asset lifecycle directly, and we have removed phrasing that implied a broader sustainability scope than the analysis supports. This positions the contribution within the resource-governance and human-capital dimensions of sustainability, which we now develop substantively.
These changes appear in Sections 1, 2.4, and 5.6 of the revised manuscript.
Major Concern 8:
Although the paper emphasizes “compliance-intensive public bureaucracy,” structural factors such as bureaucratic rigidity, political accountability, and administrative hierarchy are not substantially reflected in the model. The public sector context remains background explanation rather than a substantive explanatory mechanism.
Authors' Response:
We agree that the public-sector context should function as more than background, and we have revised accordingly. Section 2.3 now integrates the public-sector context—compliance intensity, the BerAKHLAK reform agenda, and the procedural constraints emphasized in the public-sector innovation literature—into the hypothesis development as boundary conditions that shape when innovation translates into rated performance, rather than as descriptive scene-setting. We acknowledge, as a limitation (Section 5.9), that structural factors such as political accountability and administrative hierarchy are not modeled as measured variables; we therefore frame them as theoretical mechanisms and as targets for future multi-level research rather than claiming to have tested them.
These changes appear in Sections 2.3 and 5.9 of the revised manuscript.
Minor Concerns
Minor Concern 1:
There is a lack of descriptive statistics regarding the basic sample profile, such as gender, age, job title, years of service, and department composition.
Authors' Response:
We have added the full demographic profile—gender, age, education, tenure in government, position, type of work unit, and experience in BMD management—in Table A1, summarized in Section 4.1. These changes appear in Section 4.1 and Table A1.
Minor Concern 2:
Despite the very large number of items, information on actual item wording and operationalization was not sufficiently presented.
Authors' Response:
The complete wording of all 109 indicators, in both Bahasa Indonesia (as administered) and English, is now provided in Appendix D, with indicator-level statistics in Appendix A. These additions appear in Appendices A and D.
Minor Concern 3:
Levels of α > .98 and CR > .99 were reported at the pilot stage, which may suggest redundancy rather than simply “excellent reliability.”
Authors' Response:
We agree, and we have revised the pilot-stage framing in Section 3.5 so that the very high pilot reliabilities are no longer described simply as “excellent”; consistent with Section 4.2.1, we note that reliabilities of this magnitude also warrant scrutiny for potential redundancy. These changes appear in Section 3.5.
Minor Concern 4:
Please review the consistency of citation ordering and formatting throughout the References section more closely.
Authors' Response:
We have corrected the references: a duplicate entry was consolidated, inconsistent list formatting was removed, and the citation ordering and formatting were checked in full for consistency with the MDPI style. These corrections appear throughout the References section.
Minor Concern 5:
The current figures, including Figures 1 and 2, remain at the level of basic mediation diagrams; more professional and informative visualizations are needed.
Authors' Response:
Figures 1 and 2 have been redrawn as more professional and informative visualizations that display the reflective higher-order structure of HCSL, the hypothesized paths, and—for Figure 2—the standardized path coefficients, explained variance, and the mediated indirect effect. These appear as Figures 1 and 2.
Minor Concern 6:
Results are being restated in the Discussion. Sections 5.2–5.5, in particular, have a high proportion of repetition and need to be condensed into a more analytical discussion.
Authors' Response:
We agree, and Sections 5.2–5.5 have been condensed and rewritten from result-restatement toward analytical discussion, interrogating the magnitude, mechanism, and boundary conditions of the findings rather than re-reporting coefficients. These changes appear in Sections 5.1–5.5.
We thank the reviewer for the rigor of this assessment. The revisions prompted by these comments have, in our view, produced a substantially more rigorous, transparent, and appropriately calibrated manuscript.
Reviewer 3 Report
Comments and Suggestions for AuthorsDear Authors,
I found your article timely, with a clear theoretical anchoring in Self-Determination Theory, four explicitly formulated hypotheses, an adequate sample (N = 541), and a systematic two-stage PLS-SEM procedure. The focus on civil servants engaged in public asset stewardship is a genuinely under-explored domain. However, several methodological concerns require substantive revision before the article can be considered for publication. My concerns are not marginal; they bear directly on the validity of the empirical claims.
- Unexplained inflation of scale length
The article relies on validated instruments but applies them in expanded forms whose construction is not explained. HCSL is operationalised through 48 indicators against the 16 items of the original Di Fabio and Peiro (2018) scale; IWB through 26 items against the 10 of De Jong and Den Hartog (2010); and EJP through 35 indicators against the 18 of the IWPQ (Koopmans et al., 2014). Respondents thus completed approximately 109 Likert items derived from instruments whose validated length is far shorter. The manuscript does not explain how the additional items were generated, on what conceptual basis they were added, nor whether the factor structure of the extended scales was confirmed through EFA or CFA before structural modelling. This is a foundational concern: without evidence that the expanded instruments still measure the intended constructs, every subsequent inference rests on an uncertain measurement basis. I recommend a transparent account of the scale development procedure in Section 3.3, supported by a confirmatory factor analysis of the higher-order measurement model.
- Anomalously high internal consistency values
Cronbach's alpha values for the three focal constructs range from 0.989 to 0.990, and composite reliability values from 0.990 to 0.991. These values exceed the 0.95 threshold that the methodological literature (Hair et al., 2022; Sarstedt et al., 2019) flags as indicative of indicator redundancy. The manuscript acknowledges this but dismisses the concern by invoking higher-order constructs and sample homogeneity. This explanation is insufficient. Reliability values approaching 0.99 are exceedingly rare even in large validated instruments and most often reflect redundancy, halo effects, or response acquiescence - all plausible given the length and homogeneity of the questionnaire. I recommend reporting indicator-level statistics in an appendix so readers can assess potential item duplication, and engaging critically with these values rather than presenting them as uniformly positive. Excessively high reliability should be treated as a warning signal, not as a measurement strength.
- Borderline discriminant validity between HCSL and IWB
Three pieces of evidence converge on a single concern. The Fornell-Larcker criterion fails marginally for the HCSL-IWB pair: the inter-construct correlation (0.834) exceeds the square root of HCSL's AVE (0.830). The HTMT value (0.839) sits just below the conservative 0.85 threshold. The structural path coefficient HCSL to IWB (beta = 0.834, with R-squared for IWB = 0.695) is extraordinarily high. Taken together, these indicators raise a serious question about whether HCSL and IWB are empirically distinguishable constructs in this sample. The current explanation - conceptual proximity in public sector settings - is theoretically plausible but does not resolve the measurement problem. I suggest examining which specific items contribute most heavily to cross-loadings, assessing whether some items in the expanded HCSL scale may capture behavioural outcomes that conceptually overlap with IWB, and explicitly discussing the implications of this boundary case for the interpretation of H1. The magnitude of the coefficient warrants a more nuanced reading than the current straightforward confirmation.
- Insufficient treatment of common method bias
All three constructs were measured through self-report from a single respondent in a single sitting - a textbook configuration for common method bias. The procedural remedies described in Section 3.6 (anonymity assurance, separation of sections, selective reverse coding) are minimal, and the statistical diagnostic relied upon - Kock's full-collinearity VIF test - is widely regarded as a weak instrument for detecting CMB. The highest reported VIF (3.276) is so close to the 3.3 threshold that the test offers little reassurance, and the high inter-construct correlations in Tables 2 and 3 are consistent with - though not proof of - method-induced inflation. Given the magnitude of the reported path coefficients, the burden of demonstrating that CMB cannot account for a substantial portion of the observed effects is non-trivial. I recommend adding a full Harman single-factor test and ideally a marker-variable analysis using a theoretically unrelated construct. At minimum, the limitations section should acknowledge the single-source design as a material constraint and propose multi-source data collection (e.g., supervisor ratings of EJP) for future work.
- Literature review: descriptive rather than synthetic
Section 2 is well organised and references are reasonably current, yet the review functions more as successive characterisation of the focal constructs than as critical synthesis. There is little engagement with competing or alternative conceptualisations: HCSL is presented as a consensus construct without serious comparison with responsible leadership, ethical leadership in its standalone form, or sustainability-augmented transformational leadership. The 'innovation-performance paradox' is invoked as a contrast but its boundary conditions are not theoretically developed. The well-documented tensions in the public-sector innovation literature - Bysted and Jespersen, Damanpour, and the broader public service motivation tradition - are cited but not unpacked. Finally, the public-sector dimension itself is treated as background rather than as an object of analysis: hierarchical orientation, accountability pressures, and the specifics of the BerAKHLAK reform agenda are mentioned but not theoretically integrated into the hypothesis development. A more synthetic review would identify the principal positions in the leadership-innovation-performance literature, articulate the tensions among them, and locate this study's contribution as a response to those tensions, rather than as a justification for the chosen variables.
My final remarks
The article addresses a relevant question, uses appropriate methodology, and demonstrates analytical competence. The five concerns outlined above - the unexplained scale expansion, the anomalous reliability values, the borderline discriminant validity between HCSL and IWB, the limited treatment of common method bias, and the descriptive character of the literature review - do not invalidate the research project, but they cumulatively raise substantial questions about the validity of the empirical claims as currently presented. Each concern calls for a more rigorous defence of the measurement model or a deeper conceptual scaffolding for the analysis. The empirical material is rich; what it requires now is a measurement layer and an interpretative layer that match its scale and ambition. I encourage the authors to engage seriously with these issues - the revisions would substantially strengthen the contribution and make the findings more interpretable and citable.
Good luck!
Author Response
We are grateful to Reviewer 3 for a rigorous and constructive review. The five concerns raised bear directly on the validity of our empirical claims, and we have engaged with each of them seriously through substantive revision rather than rebuttal. Our point-by-point responses follow; all changes are marked using tracked changes.
Concern 1:
Unexplained inflation of scale length. The article relies on validated instruments but applies them in expanded forms whose construction is not explained. HCSL is operationalised through 48 indicators against the 16 items of the original Di Fabio and Peiro (2018) scale; IWB through 26 items against the 10 of De Jong and Den Hartog (2010); and EJP through 35 indicators against the 18 of the IWPQ (Koopmans et al., 2014). The manuscript does not explain how the additional items were generated, on what conceptual basis, nor whether the factor structure of the extended scales was confirmed through EFA or CFA before structural modelling. I recommend a transparent account of the scale development procedure in Section 3.3, supported by a confirmatory factor analysis of the higher-order measurement model.
Authors' Response:
We thank the reviewer for this foundational concern, which we have addressed directly. We have added an “Instrument Development and Validation” subsection to Section 3.3 documenting the scale-development procedure that was omitted from the original submission. Item generation proceeded deductively from the established dimensional structure of each parent instrument; no new dimensions were introduced. Instead, each theoretically specified dimension was elaborated into several context-specific behavioral indicators expressing the construct within the regional-asset-management (BMD) setting, following established guidance on measure development (Hinkin, 1998; MacKenzie et al., 2011; Carpenter, 2018) and on cross-cultural adaptation (Beaton et al., 2000). The content validity of the resulting items was established through expert review prior to data collection, with the minor refinements recommended by the expert incorporated before pilot testing.
On confirmation of the factor structure: HCSL was estimated as a reflective higher-order construct using the two-stage approach (Sarstedt et al., 2019), and the full measurement model was assessed against established thresholds (outer loadings ≥ 0.708, AVE ≥ 0.50, and reliability), which serves as the variance-based counterpart to covariance-based confirmatory factor analysis. To allow independent scrutiny of the expanded instruments, indicator-level statistics for all 109 items (Appendix A) and the complete bilingual item wording (Appendix D) are now reported.
These changes appear in Section 3.3, Section 4.2, and Appendices A and D of the revised manuscript.
Concern 2:
Anomalously high internal consistency values. Cronbach’s alpha values range from 0.989 to 0.990 and composite reliability from 0.990 to 0.991, exceeding the 0.95 threshold flagged as indicative of redundancy. The manuscript acknowledges this but dismisses the concern by invoking higher-order constructs and sample homogeneity. I recommend reporting indicator-level statistics in an appendix and engaging critically with these values. Excessively high reliability should be treated as a warning signal, not as a measurement strength.
Authors' Response:
We fully agree, and we have reframed our treatment accordingly. Section 4.2.1 now treats the high reliabilities as a warning signal rather than a strength, acknowledging per Hair et al. (2022) that values above 0.95 typically indicate indicator redundancy. We discuss the plausible contributors in this dataset—semantic proximity among the indicators elaborating each dimension, a pronounced ceiling effect (construct means of 4.37–4.45 with a mean item standard deviation below 0.70), respondent homogeneity, and potential halo and acquiescence effects from single-source self-report.
As the reviewer recommends, we now report indicator-level statistics—loadings, item means, standard deviations, and corrected item–total correlations—for all 109 items in Appendix A, enabling direct assessment of potential item duplication. We further note that a supplementary check showed that reducing each scale to its originally validated length still yields Cronbach’s α above 0.98 for all three focal constructs, indicating that the elevated values stem chiefly from sample homogeneity and response tendencies rather than from item count alone.
These changes appear in Section 4.2.1 and Appendix A of the revised manuscript.
Concern 3:
Borderline discriminant validity between HCSL and IWB. The Fornell-Larcker criterion fails marginally (inter-construct correlation 0.834 exceeds √AVE(HCSL) 0.830); HTMT (0.839) sits just below the 0.85 threshold; and the path HCSL→IWB (beta = 0.834, R² = 0.695) is extraordinarily high. I suggest examining which specific items contribute most heavily to cross-loadings, assessing whether some HCSL items capture behavioural outcomes that overlap with IWB, and explicitly discussing the implications for the interpretation of H1.
Authors' Response:
We thank the reviewer for this careful and well-justified concern, which we have engaged with directly, including the specific item-level analysis recommended. Section 4.2.3 now reports the Fornell-Larcker result transparently as a genuine, if marginal, violation (0.834 vs. 0.830) rather than as merely “borderline.” Following the reviewer’s suggestion, we examined the indicator cross-loadings to determine whether specific HCSL items capture behavior overlapping with IWB. This examination found that every HCSL indicator loaded more strongly on HCSL than on IWB—that is, the cross-loadings criterion for discriminant validity is satisfied throughout—and the complete cross-loading matrix is now provided in Appendix B.
The evidence is therefore mixed: of the three diagnostics, two are met (HTMT = 0.839 < 0.85; clean cross-loadings) and one is marginally not met (Fornell-Larcker). We interpret this convergence as indicating that HCSL and IWB are empirically distinct but conceptually proximate constructs. We have added a passage to Section 2.2 positioning HCSL relative to adjacent leadership constructs and explaining why such proximity is theoretically expected, and we discuss the implication for H1 explicitly in Section 5.2: the very strong HCSL→IWB path is read as reflecting both a genuine need-supportive mechanism and this conceptual adjacency—together with possible method-related inflation—rather than as an exceptionally large unique effect.
These changes appear in Sections 2.2, 4.2.3, 5.2, and Appendix B of the revised manuscript.
Concern 4:
Insufficient treatment of common method bias. All constructs were measured through self-report from a single respondent in a single sitting. The procedural remedies are minimal and the statistical diagnostic—Kock’s full-collinearity VIF—is a weak instrument; the highest VIF (3.276) is so close to the 3.3 threshold that it offers little reassurance. I recommend adding a full Harman single-factor test and ideally a marker-variable analysis; at minimum, the limitations section should acknowledge the single-source design and propose multi-source data collection for future work.
Authors' Response:
We thank the reviewer, and we have strengthened both the analysis and its interpretation. As recommended, we added Harman’s single-factor test; the first unrotated factor accounted for 62.6% of the variance, exceeding the 50% benchmark. We report this transparently and do not interpret it as exonerating. Rather, we note that Harman’s test is widely regarded as insensitive, that in a 109-indicator, highly reliable, homogeneous instrument a large first factor is partly an artifact of the test, and that—together with the limitations of Kock’s full-collinearity approach (highest VIF 3.276, close to the 3.3 threshold)—the available diagnostics cannot rule out common-method and halo influence.
Accordingly, we have removed the original claim that CMB “is not a serious concern,” acknowledged that the high correlations and reliabilities are consistent with some method-induced inflation, and stated that we interpret the effect magnitudes conservatively (Section 3.6). With respect to a marker variable, no theoretically unrelated marker was included a priori in the present design; we acknowledge this limitation and, as the reviewer recommends, Section 5.9 now treats the single-source, single-occasion design as a material constraint and proposes multi-source data—including supervisor or archival ratings of performance—an a priori marker variable, and temporally separated measurement for future research.
These changes appear in Sections 3.6 and 5.9 of the revised manuscript.
Concern 5:
Literature review: descriptive rather than synthetic. Section 2 functions more as successive characterisation of the focal constructs than as critical synthesis. There is little engagement with competing conceptualisations (responsible leadership, ethical leadership in its standalone form, sustainability-augmented transformational leadership); the innovation-performance paradox is invoked but its boundary conditions are not developed; the public-sector tensions (Bysted and Jespersen, Damanpour, public service motivation) are cited but not unpacked; and the public-sector dimension is treated as background rather than as an object of analysis.
Authors' Response:
We thank the reviewer for this constructive critique, which prompted a substantial revision of the theoretical sections toward synthesis. Section 2.2 now positions HCSL explicitly against adjacent leadership conceptualizations—responsible leadership (Pless & Maak, 2011), ethical leadership as a standalone construct (Brown et al., 2005), and sustainability-/environmentally-specific transformational leadership (Robertson & Barling, 2013)—articulating both HCSL’s distinctiveness and the conceptual proximity that bears on our measurement results.
Section 2.3 now develops the boundary conditions of the innovation–performance paradox rather than merely invoking it, drawing on the public-sector innovation literature (Bysted & Jespersen, 2014; Damanpour & Schneider, 2009) and on public service motivation (Perry & Wise, 1990) to specify the conditions under which innovation is likely to translate into rated performance in this setting. The public-sector dimension—including the BerAKHLAK reform agenda—is now treated as an explanatory mechanism integrated into the hypothesis development rather than as background, and Section 2.4 makes the sustainability framing substantive by linking asset stewardship to the resource-governance dimension of sustainability. Finally, the Discussion (Sections 5.1–5.5) has been restructured from restatement toward analysis, locating the study’s contribution as a calibrated, context-specific response to these tensions.
These changes appear in Sections 1, 2.2, 2.3, 2.4, and 5.1–5.5 of the revised manuscript.
We thank the reviewer again for the depth of this review. The revisions prompted by these five concerns have, in our view, materially strengthened the measurement transparency and interpretive rigor of the manuscript.
Reviewer 4 Report
Comments and Suggestions for AuthorsHello Dear Auhor(s),
1. The paper makes an incremental empirical contribution and does not advance the field theoretically or methodologically. The study tests the mediation effect of HCSL on the relationship between Innovative Work Behavior and Employee Job Performance among Indonesian civil servants. The paper topic is relevant, but the theoretical significance is overstated since the mediation model proposed seems to confirm widely recognized hypotheses of leadership–innovation–performance rather than introduce a new model.
2. The primary conceptual weakness is that even though HCSL, IWB, and EJP are presented as distinct constructs, empirically, they share substantial conceptual similarities. The fact that reliability indices were very high, coupled with the problematic Fornell-Larcker criterion for the HCSL-IWB correlation, suggests conceptual redundancy or even inflation of one construct into two others. The research has touched upon this issue briefly, while in my opinion, it should be discussed in detail since the three constructs are not independent.
3. The theoretical basis is based on Self-Determination Theory, but SDT itself is not operationalized in any form. While it is argued that HCSL promotes the psychological needs for autonomy, competence, and relatedness, thus stimulating innovation and performance, psychological needs satisfaction is not included in the model. Thus, there is an implicit theoretical hole in the manuscript; the proposed mechanism underlying the theory is simply assumed without being examined empirically.
4. The methodology is exposed to the problem of common method variance. Both independent and dependent variables are self-reported, collected from the same sample, at the same point in time, using the same type of Likert scales. The application of full collinearity VIF test is not enough to overcome this limitation. 4. Since the path between HCSL and IWB is very strong, it is possible that the results may be biased by response consistency and social desirability.
5. The limitations associated with the use of the sampling strategy limit the potential for generalizability of the findings. The research utilizes purposeful sampling in one context, which is the provincial government of DKI Jakarta, and one particular area of administration, which is the management of public assets. Thus, the manuscript cannot make assertions regarding sustainability leadership in the public sector in general terms.
6. The operationalization of the constructs used in this study needs to be supported by a stronger rationale. Specifically, HCSL is measured using 48 indicators, while IWB and EJP have 26 and 35 indicators, respectively, making a very large questionnaire. It needs to be shown why such extensive operationalizations were required, and how item redundancy might have been considered or addressed.
7. The statistics provided in the manuscript provide very strong evidence, although not entirely convincing. High Cronbach’s alpha coefficients beyond 0.98 can point towards redundancy of the constructs rather than high quality of their measurement. However, the research seems to present these coefficients as positive.
8. The discussion overextends the findings. Claims that the study reframes the innovation–performance paradox are not sufficiently supported. A single cross-sectional study in one Indonesian public asset management context cannot resolve or reframe a broader theoretical debate. At most, it provides context-specific evidence that IWB is positively associated with self-reported performance.
9. The manuscript’s practical implications are plausible but insufficiently grounded in the empirical design. Recommendations about leadership development, sustainability investment, and performance management go beyond what the data can directly establish. The study does not test interventions, leadership training, institutional reforms, or objective asset-management outcomes.
10. An improvement to the research would be the integration of more recent and conceptually advanced literature that strengthens its theoretical grounding in policies e.g., https://doi.org/10.3390/world6040150, https://doi.org/10.3390/su17052237.
11. Another significant limitation pertains to the lack of objective performance metrics. Employee Job Performance was self-rated, while the context of public asset management might have provided more tangible performance indicators like timeliness, reporting effectiveness, compliance results, maintenance effectiveness, or audit indicators. In this regard, it is difficult to establish a direct link between leadership effectiveness and asset management practices.
12. The paper's structure and organization are adequate, but its rhetoric seems too confident. The writing often equates statistical association with a managerial cause-and-effect relationship. The conclusions are exaggerated and need to be significantly revised, along with a better differentiation of empirical findings from theory.
In general, good effort. Please, fix the prior issues.
Best Regards,
Author Response
We thank Reviewer 4 for the careful reading and the candid, constructive feedback. We have addressed each of the twelve points through revision, with particular attention to recalibrating the manuscript’s claims and language. Our point-by-point responses follow; all changes are marked using tracked changes.
Point 1:
The paper makes an incremental empirical contribution and does not advance the field theoretically or methodologically. The paper topic is relevant, but the theoretical significance is overstated since the mediation model proposed seems to confirm widely recognized hypotheses of leadership–innovation–performance rather than introduce a new model.
Authors' Response:
We thank the reviewer, and we have recalibrated the contribution throughout. The Introduction now states explicitly that the contribution is empirical and contextual rather than a claim to novelty in the leadership–innovation–performance structure, which is well established. The revised contribution emphasizes context-specific evidence from an under-examined public-asset-management setting, the specification of boundary conditions for the innovation–performance relationship, and the clarification of the relative magnitudes of the direct and mediated pathways—while explicitly bounding these contributions by the study’s design. The Theoretical Contributions (Section 5.7) and Conclusions have been revised in the same spirit.
These changes appear in Section 1, Section 5.7, and Section 6 of the revised manuscript.
Point 2:
Even though HCSL, IWB, and EJP are presented as distinct constructs, empirically they share substantial conceptual similarities. The very high reliability indices, coupled with the problematic Fornell-Larcker criterion for the HCSL-IWB correlation, suggest conceptual redundancy or inflation. This should be discussed in detail since the three constructs are not independent.
Authors' Response:
We agree this warranted fuller treatment, and we have provided it. Section 4.2.3 now reports the Fornell-Larcker result transparently as a marginal violation and presents an item-level cross-loading examination (Appendix B) showing that every indicator loads most strongly on its own construct. Section 2.2 positions HCSL against adjacent leadership constructs and explains why the HCSL–IWB proximity is theoretically expected, and Section 5.2 interprets the strong HCSL→IWB path as reflecting both a genuine mechanism and conceptual adjacency—together with possible method-related inflation—rather than an exceptionally large unique effect. We conclude that the constructs are empirically distinct but proximate, and we discuss this boundary case explicitly.
These changes appear in Sections 2.2, 4.2.3, and 5.2 of the revised manuscript.
Point 3:
The theoretical basis is Self-Determination Theory, but SDT itself is not operationalized in any form. Psychological needs satisfaction is not included in the model. There is an implicit theoretical hole; the proposed mechanism is simply assumed without being examined empirically.
Authors' Response:
We thank the reviewer for identifying this gap, which we now address candidly. Section 2.1 clarifies that HCSL can be read as an operationalization of need-supportive leadership—its mindful dimension as autonomy support, its sustainable dimension as competence development, and its ethical and servant dimensions as relatedness—and that SDT functions in this study as an explanatory lens rather than a directly tested mechanism. We state explicitly that the satisfaction of the three psychological needs was not measured, and Section 5.9 identifies the direct measurement of need satisfaction as a priority for future research.
These changes appear in Sections 2.1 and 5.9 of the revised manuscript.
Point 4:
The methodology is exposed to common method variance. Both independent and dependent variables are self-reported, from the same sample, at the same point in time, using the same type of Likert scales. The full collinearity VIF test is not enough. Since the HCSL–IWB path is very strong, the results may be biased by response consistency and social desirability.
Authors' Response:
We agree and have strengthened our treatment. Section 3.6 now adds Harman’s single-factor test (first factor 62.6%, reported transparently), notes the limitations of both Harman’s test and Kock’s full-collinearity VIF, and removes the original claim that CMB is not a serious concern. We acknowledge that the strong correlations are consistent with response-consistency and social-desirability influences, and we interpret the effect magnitudes conservatively. Section 5.9 treats the single-source, single-occasion design as a material limitation and proposes multi-source and temporally separated data for future work.
These changes appear in Sections 3.6 and 5.9 of the revised manuscript.
Point 5:
The sampling strategy limits generalizability. The research uses purposive sampling in one context (the provincial government of DKI Jakarta) and one particular administrative area (BMD management). Thus, the manuscript cannot make assertions regarding sustainability leadership in the public sector in general terms.
Authors' Response:
We accept this limitation and have made it explicit. Section 5.9 now states that the single-jurisdiction (DKI Jakarta), single-domain (BMD), purposively sampled design limits generalizability to other regional governments and public-sector settings, and we propose replication across Indonesian regions and public-sector domains. We have correspondingly tempered general claims about public-sector sustainability leadership throughout the manuscript.
These changes appear in Section 5.9 of the revised manuscript.
Point 6:
The operationalization of the constructs needs a stronger rationale. HCSL is measured using 48 indicators, while IWB and EJP have 26 and 35 indicators, making a very large questionnaire. It needs to be shown why such extensive operationalizations were required, and how item redundancy might have been considered or addressed.
Authors' Response:
This concern is closely related to the scale-development question, which we now address in full. Section 3.3 documents the deductive, dimension-based generation of the expanded items, the expert content-validation step, and the two-stage higher-order estimation; Appendix A reports indicator-level statistics for all 109 items so that redundancy can be assessed directly. We also report that reducing each scale to its originally validated length still yields Cronbach’s α above 0.98, indicating that the high values reflect sample characteristics rather than item count; we therefore addressed redundancy through transparency rather than ad hoc item deletion, which would discard expert-validated content.
These changes appear in Section 3.3, Section 4.2.1, and Appendix A of the revised manuscript.
Point 7:
The statistics provide very strong, although not entirely convincing, evidence. High Cronbach’s alpha coefficients beyond 0.98 can point towards redundancy of the constructs rather than high quality of measurement. The research seems to present these coefficients as positive.
Authors' Response:
We agree. Section 4.2.1 now treats the high reliabilities as a warning signal of possible redundancy rather than as a positive feature, discusses the likely contributors (semantic proximity, ceiling effects, homogeneity, and single-source halo/acquiescence), and reports indicator-level statistics in Appendix A for direct scrutiny.
These changes appear in Section 4.2.1 and Appendix A of the revised manuscript.
Point 8:
The discussion overextends the findings. Claims that the study reframes the innovation–performance paradox are not sufficiently supported. A single cross-sectional study in one Indonesian public asset management context cannot resolve or reframe a broader theoretical debate. At most, it provides context-specific evidence that IWB is positively associated with self-reported performance.
Authors' Response:
We thank the reviewer, and we have removed the overreaching claims. The manuscript no longer states that the study “reframes” the innovation–performance paradox. Sections 5.3 and 5.7 and the Conclusions now present the finding as context-specific evidence that, under the conditions characterizing regional asset management, innovation and performance are positively linked—explicitly framing this as informing the debate about when the paradox holds, not as resolving it.
These changes appear in Sections 5.3, 5.7, and 6 of the revised manuscript.
Point 9:
The practical implications are plausible but insufficiently grounded in the empirical design. Recommendations about leadership development, sustainability investment, and performance management go beyond what the data can directly establish. The study does not test interventions, leadership training, institutional reforms, or objective asset-management outcomes.
Authors' Response:
We agree and have reframed Section 5.8. The practical implications are now presented as tentative, with an explicit caveat that the study did not test interventions, training, institutional reforms, or objective asset-management outcomes and therefore cannot establish their effects directly; the recommendations are framed as hypotheses for practice rather than as prescriptions.
These changes appear in Section 5.8 of the revised manuscript.
Point 10:
An improvement would be the integration of more recent and conceptually advanced literature that strengthens the theoretical grounding in policies (e.g., https://doi.org/10.3390/world6040150, https://doi.org/10.3390/su17052237).
Authors' Response:
We thank the reviewer for these suggestions. We have considered both recommended works and cited them where they strengthen the theoretical and policy grounding of the study. These additions appear in the revised manuscript and References.
Point 11:
Another significant limitation pertains to the lack of objective performance metrics. Employee Job Performance was self-rated, while the public asset management context might have provided more tangible performance indicators like timeliness, reporting effectiveness, compliance results, maintenance effectiveness, or audit indicators.
Authors' Response:
We agree this is an important limitation. Section 5.9 now states explicitly that Employee Job Performance was self-rated and that the public-asset-management context offers more objective indicators—reporting timeliness, audit and reconciliation outcomes, maintenance effectiveness, and unqualified-opinion (WTP) attainment—which would provide a sterner test of the leadership–performance link and which we recommend for future multi-source designs.
These changes appear in Section 5.9 of the revised manuscript.
Point 12:
The paper’s structure is adequate, but its rhetoric seems too confident. The writing often equates statistical association with a managerial cause-and-effect relationship. The conclusions are exaggerated and need to be significantly revised, along with a better differentiation of empirical findings from theory.
Authors' Response:
We thank the reviewer, and we have revised the manuscript’s language throughout to avoid equating statistical association with causation. Causal verbs in the Abstract, the discussion, and the Conclusions have been replaced with associational language (“is associated with,” “predicts,” “relates to”); the cross-sectional design’s limits on causal inference are stated in Section 5.9; and we now consistently differentiate empirical findings from theoretical interpretation—for example, describing the design as consistent with the SDT mechanism rather than as a test of it. The Conclusions have been recalibrated accordingly.
These changes appear in the Abstract, Sections 5.1–5.9, and Section 6 of the revised manuscript.
We thank the reviewer again for the constructive guidance. The revisions have, in our view, produced a more measured and defensible manuscript whose claims are appropriately matched to its evidence.
Round 2
Reviewer 2 Report
Comments and Suggestions for Authors<Overall Evaluation>
This manuscript has been significantly improved compared to the previous version. The authors have made considerable efforts to address the reviewers' concerns, such as readjusting the claims of research contribution, strengthening the theoretical framework, expanding the discussion on general methodological biases and measurement limitations, clarifying the sustainability perspective, and making substantial revisions to the discussion and limitations sections. Nevertheless, several significant concerns remain. While these issues do not invalidate the study itself, they limit the validity of the conclusions that can be drawn from the results; therefore, the remaining limitations must be explained and resolved more clearly before final acceptance for publication. Accordingly, my decision is '(strong) minor revisions'.
<Major Concerns>
(1) The current version of the manuscript appropriately acknowledges the abnormally high confidence coefficient, the strong HCSL-IWB relationship, and the possibility of common method inflation. However, the implications of these results need to be discussed more clearly. Very high confidence coefficients (α ≈ .99; CR ≈ .99), very high inter-construct correlations, strong HCSL → IWB path coefficients (β = .834), and high explanatory variance (R² = .695) suggest that conceptual overlap, response consistency tendencies, and common source effects may have contributed to the observed relationship. While the authors acknowledge these issues, the discussion should more clearly emphasize that the magnitude of the reported effect may be an upper limit estimate rather than an accurate representation of the underlying relationship.
(2) This paper clearly indicates that the violation of the Fornell-Larcker assumption between HCSL and IWB is minimal and presents evidence supporting this through HTMT and cross-loading analysis. Nevertheless, the empirical distinction between these two constructs remains relatively weak. Therefore, rather than unconditionally interpreting the reported relationship as evidence of a strong and unique effect of leadership on innovation, the authors should further clarify that the relationship between HCSL and IWB is observed under significant conceptual and empirical proximity between the two constructs.
(3) This paper explains that it did not fundamentally redesign the existing scale but rather refined it to fit the context. However, the resulting measurement model includes a total of 109 indicators across three constructs, 48 of which are indicators for the HCSL. The authors need to clarify the conceptual logic underlying this extensive expansion and explain more specifically how the modified indicators maintain the theoretical boundaries of the existing tool. Readers may be confused as to whether this study is a modified version of the existing scale or is substantially developing a new measurement tool.
(4) The revised manuscript has significantly improved the discussion of context. However, key contextual elements repeatedly emphasized throughout the manuscript—such as bureaucratic rigidity, pressure of responsibility, administrative hierarchy, and public asset management—are not directly reflected in the empirical model. Consequently, while this study demonstrates that the proposed relationship exists within the public asset management environment, it does not directly verify whether this context shapes that relationship. The authors should continue to emphasize these differences and avoid implying that the contextual mechanism itself has been empirically verified.
(5) Employee job performance still relies entirely on self-reporting. Given that leadership perceptions, innovative behaviors, and performance evaluations were collected simultaneously from the same respondents, perception consistency bias remains a valid explanation for some of the observed relationships. The manuscript should more clearly mention the possibility that the use of self-assessed performance influenced the results as one of the limitations.
<Minor Concerns>
(1) Although claims of causality have been significantly reduced in the manuscript, some passages in the discussion and practical implications sections still imply that HCSL necessarily improves or enhances performance outcomes. All interpretations must be strictly based on association, and expressions implying causality should be completely removed.
(2) Some recommendations were presented in a manner that could be interpreted as evidence-based prescriptions. Given the observational study design, practical implications should be presented as valid implications rather than proven interventions.
(3) The revised manuscript has significantly improved the sustainability framework. However, the authors need to make it clearer that this study addresses sustainability primarily in terms of human capital sustainability and public resource management, rather than environmental sustainability, ESG performance, or sustainable asset lifecycle management.
(4) The public asset management environment provides opportunities to utilize objective indicators such as audit results, reporting accuracy, asset utilization, maintenance efficiency, and compliance performance. Discussions for future research need to address objective performance indicators, including these metrics, more clearly.
(5) The manuscript contains remnants of the change tracking feature, duplicate phrases, formatting inconsistencies, and traces of editing. It must undergo a careful editorial review before publication. Authors are encouraged to improve readability for reviewers by providing both versions of the word processing program with and without the tracking feature enabled.
(6) Authors must verify that duplicate graphic elements, formatting inconsistencies, and remnants of previous versions in Figures 1 and 2 have been completely removed from the final manuscript.
(7) Despite the revisions, it appears that duplicate entries, numbering inconsistencies, and formatting errors remain in the bibliography. The bibliography must be carefully reviewed in accordance with journal guidelines.
(8) I appreciate the inclusion of detailed appendices. However, please improve the connectivity between the main text and the appendices by adding reference markers that guide readers to clearly refer to relevant appendices when discussing measurement quality, cross-loading, indicator-level statistics, etc.
Author Response
Please see the attachment
Author Response File:
Author Response.docx
Reviewer 3 Report
Comments and Suggestions for AuthorsDear Authors,
Thank you for your work.
This version of your article is very good and meets all of the journal's requirements.
Best regards!
Author Response
Reviewer's comment: Dear Authors, thank you for your work. This version of your article is very good and meets all of the journal's requirements.
Response: We sincerely thank the reviewer for the positive evaluation of the revised manuscript and for the constructive feedback provided during the first review round, which contributed substantially to the improvement of this work.
For completeness, we note that in response to Reviewer 2’s remaining comments, the final version incorporates further clarifications of the interpretive boundaries of our findings (upper-bound reading of effect magnitudes, the conceptual proximity of HCSL and IWB, the untested contextual mechanism, and the self-report limitation), together with a full editorial and bibliographic cleanup. None of these changes alters the substance, methods, or conclusions of the version the reviewer evaluated.
We are grateful for the reviewer’s time and consideration.
Author Response File:
Author Response.docx
