Review Reports
- Kuo-Yi Jade Chang 1,*,
- Jennifer Smith-Merry 1 and
- Ying Li 2
- et al.
Reviewer 1: Helen Lockett Reviewer 2: Anonymous Reviewer 3: António Duarte Santos
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThank you for asking me to review this manuscript. The review covers a topic of much interest and relevance internationally, to understand different approaches to understanding the costs and benefits of supported employment services. I spent some time.
The review is well conducted and the manuscript very well written.
I commend the authors for such a thorough review.
I have no major or minor changes
Author Response
Please see attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsGeneral Assessment
This manuscript addresses a genuine and underexplored gap in the literature on supported employment for people with severe mental illness. A methodological review focused specifically on costing and outcome measurement practices across RCTs is timely and potentially useful for both researchers and policymakers. The practical guidance synthesized in Table 4 is the manuscript's most valuable contribution and is clearly structured. The supplementary material (Supplementary Material 1) provides a detailed and systematic extraction of study characteristics, outcome measures, data collection methods, and quality assessments across all 32 included trials, and represents a substantive addition to the manuscript. However, several concerns — some of which are directly evidenced by the supplementary data itself — must be addressed before the manuscript can be considered for publication.
Major Concerns
1. The relationship between data and recommendations is not sufficiently demonstrated
The practical guidance presented in Section 4.2 and Table 4 constitutes the manuscript's primary contribution. However, as currently written, it is not clear how this guidance is derived from the 32 included trials rather than from the existing methodological literature — particularly Drummond et al. (2015) and the CHEERS 2022 guidelines, both of which are already cited. The authors should explicitly demonstrate, for each recommendation, how the evidence extracted from the included studies supports or necessitates that specific guidance. Without this, the guidance reads as a restatement of established health economic standards applied to a new context, rather than as an original contribution emerging from this review's findings.
2. Internal inconsistencies between the quality assessment scores and the extracted data in the supplementary material
This is the most substantive concern identified upon reading Supplementary Material 1 in conjunction with the main manuscript. Two cases are particularly problematic. First, for Bond et al. (2016), the supplementary material explicitly notes under data collection methods: "Detailed descriptions of data collection methods for each outcome measure are not provided in this paper and are instead referenced to earlier studies." Despite this, the study receives affirmative quality ratings on both outcome identification and outcome measurement. This is internally inconsistent: if data collection methods cannot be assessed from the published report, a rating of "Yes" on measurement quality cannot be justified without retrieving and reviewing those earlier publications — a step the authors explicitly state they did not take (Section 4.5, third limitation). Second, for DeTore et al. (2023), the supplementary records that data collection methods are "Not explicitly reported" for both outcome domains, yet the study again receives affirmative quality ratings. The same logical inconsistency applies.
These cases suggest that the quality assessment criteria were not applied uniformly across studies, or that the threshold for an affirmative rating was insufficiently defined. The authors should either revise the quality ratings for these studies with appropriate justification, or explicitly clarify the decision rule applied when methodological details were absent from the primary publication.
3. The modified CHEC criteria are insufficiently described in the main text
Reading Supplementary Material 1 makes it possible to infer that the quality assessment was based on three criteria: (1) whether all important and relevant costs were fully identified in relation to the stated perspective and research question; (2) whether costs were measured using appropriate physical units; and (3) whether sources of valuation and reference years were clearly reported. However, none of this is stated explicitly in the main manuscript. Section 2.3 describes the assessment as using "a modified subset of items from the CHEC list" without specifying which items were retained, which were excluded, and why. This omission significantly limits the reproducibility of the quality assessment. The authors should describe the criteria explicitly in the main text, explain the rationale for selecting them, and define the decision rules applied, including how partial or unclear compliance was handled.
4. The data extraction procedure requires clarification
The manuscript describes an unusual extraction procedure in which two separate reviewer pairs extracted data on different domains — one pair on outcomes, one on costing — with a fifth author validating all data. The rationale for this division is not explained, nor are its potential implications for consistency discussed. Standard practice requires that each study be extracted by at least two independent reviewers across all domains, with explicit inter-rater reliability reporting. The authors should clarify this procedure and address whether consistency across domains was formally assessed.
5. The economic analysis subsample is too small to support the breadth of conclusions drawn
Only seven of the 32 included studies incorporated an economic analysis. While the authors acknowledge this implicitly, they do not sufficiently reflect on how it limits the generalizability of their methodological observations regarding costing practices. Several claims in Section 3.3 are presented with a degree of confidence that the evidence base does not fully support. The authors should more consistently qualify their observations about economic analyses as preliminary or illustrative, given the small and heterogeneous subsample.
6. Repetition of the central argument without analytical development
The manuscript's core interpretive claim — that methodological heterogeneity across studies reflects differing research questions and policy contexts rather than inconsistent intervention effectiveness — is introduced in the abstract, restated in the Introduction:
- This heterogeneity should not be interpreted as a flaw in the evidence base or as an indication that supported employment is inconsistently effective. Rather, it reflects the fact that evaluations have been conducted to address different research questions, policy objectives, and knowledge gaps...
and repeated again in Section 4.1:
- This heterogeneity is neither unexpected nor inherently problematic. Rather, it reflects the fact that evaluations have been designed to address different research questions, inform different policy decisions, and operate within distinct institutional, labor market, and fiscal contexts.
This is a legitimate and important argument, but it is asserted rather than demonstrated. The manuscript would be considerably strengthened if the authors used their extracted data to substantiate this claim — for example, by systematically mapping methodological choices onto the stated research questions and policy contexts of each study, thereby showing empirically that variation is explained by context rather than by arbitrary or inconsistent design decisions.
Minor Concerns
7. The title is partially misleading
The title "How Do We Know It Works?" implies a broader assessment of intervention effectiveness, whereas the manuscript is exclusively concerned with methodological evaluation practices. A title more accurately reflecting the manuscript's scope would improve clarity and avoid attracting readers with different expectations.
8. The analytical perspective classification requires more precision
Two studies — Heslin et al. (2011) and Howard et al. (2010) — are described as not explicitly reporting their analytical perspective, with the authors inferring it from the cost categories included. This inference is clearly documented in the supplementary material, where both studies are rated as "Unable to judge" on cost identification quality. However, the main text does not consistently flag this distinction between stated and inferred perspectives, and the implications for the quality ratings assigned to these two studies are not discussed. The authors should ensure that this ambiguity is clearly signalled wherever these studies are cited in the economic analysis discussion.
9. The discussion of follow-up duration recommendations lacks grounding in the review's own findings
Section 4.2 recommends follow-up periods as necessary for robust evaluation, citing McGurk et al. (2016) and Christensen et al. (2023). While these references support the recommendation, the authors do not draw on the distribution of follow-up periods across their own 32 included trials to contextualize this guidance — for instance, by reporting how many studies fell below the recommended threshold and what methodological consequences this had for the outcomes they were able to assess. This information is available in the extracted data and would strengthen the empirical grounding of the recommendation considerably.
Conclusion
This manuscript addresses a real and underserved gap in the methodological literature on supported employment evaluation. The authors demonstrate solid command of health economic methods, and Supplementary Material 1 represents a careful and detailed extraction effort. However, the manuscript in its current form does not sufficiently demonstrate that its practical guidance is empirically derived from the review findings, and the quality assessment contains internal inconsistencies that undermine a core component of the analysis.
Author Response
Please see attachment.
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsComments
The distinction between a methodological review vs systematic review or meta-analysis is clear and the focus is consistent in the methodological decisions (e,g., cost, outcomes, data collection). The use of a recognized framework such as CHEC enhances the creditability of the article.
However, the study shows a dependency on the scope review, which may constitute a limitation, as the current work risks inheriting the constraints from the previous studies. Additionally, there appears to be some mixing between different types of evidences with different methodological levels (a secondary analysis doesn’t have the same weight as an RCT - see suggestion).
Regarding the data extraction, even though it was made by a pair of reviewers and then validated by an independent reviewer its missing a structured protocol to extract the data. Key aspects remain unspecified, like what is defined as outcome measurement, and how are the ambiguous studies dealt with. Given that, this is a methodological review, reproducibility is essential, and without specification its not possible to have the same results. The same applies to other parts of the article where it lacks the specification of how decisions were made.
For the economic analysis, there is an excellent distinction between CEA vs CUA vs CBA vs CCA, but could benefit if it explicitly said what are the actual implications of the mismatch between stated SROI and actual implementation. While the issue is well described, the consequences of are not fully explored.
Finally, the section Data Sources and Data Collection Approaches could benefit from a more analytical view instead of being predominantly descriptive.
Suggestions
The article clearly states that its focus is a methodological review. However, there is a potential to mix evidence types with different methodological levels. This concern arises from the selection criteria stablished in section 2.1, where studies included within the same poll may differ substantially in their objectives, methods and rigor levels. While it is not necessary to exclude these from the article, it is important to justify their inclusion and the rationale behind it. Alternatively, the analysis can be strengthen by distinguishing the outcomes from RCT and from secondary analysis.
The article would also benefit with the use of data from the studies in order to prove some of the general affirmations like "European studies showed a greater tendency…".
To enhance clarity for the readers, the economic analysis would benefit if a resume table was added to summarize the descriptive component of the methodological problems. More broadly, some sections are too long and descriptive and adopting a more analytical perspective would further strengthen the article.
Comments on the Quality of English LanguageI am not of English origin, nor do I live in a country whose official language
is English.
Author Response
Please see attachment.
Author Response File:
Author Response.pdf
Round 2
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have responded thoughtfully to the first-round review. The majority of concerns have been addressed in a satisfactory or partially satisfactory manner, and the manuscript is considerably improved. However, a direct comparison between the cover letter and the revised supplementary material reveals several discrepancies — including one case where a revision explicitly declared in the response to reviewers is not present in the submitted file. These issues must be resolved before the manuscript can be accepted.
Comment 1 (Derivation of recommendations). Adequately addressed. Cross-references to the results sections have been added throughout Section 4.2, and citations to external methodological literature are appropriately distinguished from data-grounded recommendations.
Comment 2 (Quality assessment inconsistencies: Bond et al. 2016, DeTore et al. 2023). The authors' conceptual defence — that the criterion concerns alignment between identified and measured outcomes, not the adequacy of data collection procedures — is legitimate and consistent with the CHEC assessment instructions. However, the response states that "we have added a clarifying note to cell M2 of the Supplementary Material" to prevent misinterpretation of this criterion. Upon inspection of the submitted supplementary file, this note is not present. Cell M2 (the column header "Was the outcome measurement directly based on the identified outcomes?") contains only the original question text, with no clarification added. This declared revision must be implemented. Additionally, the logical distinction between data collection adequacy and outcome-measurement alignment — which is central to the defence of these two ratings — should be stated explicitly in Section 2.3 of the main text, not only in the supplementary material, which many readers will not consult.
Comment 3 (CHEC criteria description). Well addressed. Section 2.3 now explicitly names the five selected CHEC items, provides the rationale for their selection, and defines the four-point response scale. This substantially improves reproducibility.
Comment 4 (Data extraction procedure). Adequately addressed. The domain-specific division is now explained and justified in Section 2.2. The absence of formal inter-rater reliability statistics is defended with reference to Cochrane and PRISMA guidance — a defensible position.
Comment 5 (Economic subsample size). Adequately addressed. The clarifying statement at the opening of Section 3.3 — explicitly framing the observations as descriptive and illustrative — is appropriate and sufficient.
Comment 6 (Repetition without analytical development). Partially addressed. The two concrete examples added to Section 4.1 (cognitive remediation trials and Holmås et al.'s welfare-system context) are well chosen and improve the argument. However, the central claim — that methodological variation is explained by contextual differences rather than being arbitrary — remains largely asserted across 32 studies on the basis of two illustrative cases. A brief tabular mapping in the supplementary material, showing research question, key methodological choices, and stated policy/institutional context for each included study, would make this claim empirically defensible. This is recommended.
Comment 7 (Title). Resolved. The revised title accurately reflects the manuscript's scope.
Comment 8 (Analytical perspective classification). Well addressed. The revised approach — recording unstated perspectives as "Not explicitly stated by the authors" and rating cost identification as "Unclear" accordingly — is methodologically sound and consistently applied across the three relevant studies (Heslin et al. 2011, Holmås et al. 2021, Howard et al. 2010). The identification and correction of the previously omitted Holmås et al. case demonstrates careful revision.
Comment 9 (Follow-up duration). Resolved. The cross-reference to Section 3.4.5 has been added and the connection between empirical findings and practical guidance is now explicit.
Minor Issues Identified in the Supplementary Material
Beyond the missing M2 note described under Comment 2, inspection of the supplementary file reveals the following additional issues.
1. Inconsistent classification of Christensen et al. (2019/2020). The supplementary material records the analysis type as "CUA & CEA", while the main manuscript (Table 3 and body text) consistently uses "CEA & CUA". While the substantive difference is minor, the ordering should be standardised across both documents.
2. Logical tension in the Hoffman et al. (2014) quality ratings. The outcome identification criterion (Column 11) is rated "No" — correctly, given the mismatch between the stated SROI perspective and the narrow set of outcomes actually captured. However, the outcome measurement criterion (Column 12) is rated "Yes". Under the CHEC framework, outcome measurement quality is logically conditional on the adequacy of outcome identification: if important outcomes were not identified, their measurement cannot be assessed as fully adequate. The authors should consider whether "Partially" would be a more internally consistent rating for Column 12 in this case, and provide a brief justification for whichever rating is retained.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf