2. Materials and Methods
2.1. Study Design and Relation to the Prior Report
The present study employed a retrospective observational comparative design based on archival clinical records derived from routine care within a school-based mental health service system. The study did not involve prospective recruitment, random assignment, or protocol-driven intervention delivery. All data were generated as part of standard clinical practice before the present analyses were conceived.
This analysis was conducted within the same clinical service context, ethics framework, and general study period as our related elementary-school retrospective observational study (
Kwak & Ahn, 2026), but was restricted to a distinct middle- and high-school sample. The two studies therefore involved developmentally different populations drawn from separate analytic samples, and no participants included in the elementary-school study were included in the present analysis.
Although a pooled analysis across the elementary-school and adolescent cohorts was considered, the two cohorts were analyzed separately because they represented developmentally and clinically distinct populations. The elementary-school cohort and the present middle- and high-school cohort differed in developmental stage, clinical severity, and the likely salience of self-related versus behavioral-regulation domains. Pooling the samples could have obscured these developmentally meaningful clinical differences and produced an overall estimate that was less informative for school-based intervention planning. Therefore, the present study was designed as a separate cohort analysis focused on the adolescent sample, while using the prior elementary-school study as a related developmental and service-context reference point.
Accordingly, the present analyses were intended to characterize patterns of change across two clinically implemented intervention tracks in a naturalistic setting and to examine whether intervention-associated differences were more evident in specific outcome domains rather than at the global level. Given the retrospective and clinically assigned nature of the study, the findings are interpreted as observational associations within a real-world clinical context. No causal inferences regarding treatment effectiveness, comparative superiority, or differential efficacy between interventions are intended.
2.2. Setting and Participants
Clinical records were obtained from a school-based suicide prevention center (KASS), commissioned by the Chungcheongnam-do Office of Education, Republic of Korea. The center provides crisis-oriented assessment and brief intervention for students referred through school-based pathways because of suicidal ideation, NSSI, and related emotional or behavioral concerns.
Chungcheongnam-do is a large provincial education district comprising urban, suburban, and rural school communities. The center functioned as a regional referral and intervention hub, receiving students identified by schools as being at risk for suicidal ideation, NSSI, or related difficulties. During the study period, the service system covered middle- and high-school students referred from schools across the province, thereby providing a naturalistic context for examining intervention patterns under routine school-based care conditions.
For the present study, archival clinical records were reviewed for middle- and high-school students who received services at the center between 2022 and 2024.
Cases were considered eligible for inclusion if records documented the following: (1) enrollment in middle or high school; (2) sufficient literacy to complete Korean self-report measures; (3) documented NSSI within the six months preceding intake; (4) suicidal ideation without acute intent or an imminent suicide plan at intake; and (5) participation in at least six therapeutic sessions with both pre- and post-intervention assessment data available. These criteria were applied to identify cases in which clinically meaningful engagement with the intervention had occurred and outcome data were available for analysis.
Cases were excluded if records indicated: (1) acute suicidal intent requiring emergency or inpatient care; (2) a recent medically serious suicide attempt; (3) severe psychiatric instability requiring immediate intensive treatment; (4) absence of documented NSSI; (5) discontinuation of treatment prior to completion of the minimum session threshold; or (6) transfer to a higher level of care during the treatment episode.
After application of these criteria, the final analytic sample consisted of 112 adolescents, including 52 in the SPT-SAFE group and 60 in the DBT-BI group. The participant flow and analytic inclusion process are presented in
Figure 1.
The final analytic sample represents a treatment-engaged completer sample rather than a representative sample of all 491 school-linked referrals. The reduction from 491 referrals to 112 analytic cases reflects multiple sequential stages: 156 referrals did not proceed to baseline assessment or the target intervention pathway; 143 of the remaining 335 assessed cases were excluded prior to analytic inclusion; 192 adolescents initiated SPT-SAFE or DBT-BI; and 80 of these initiators were excluded from the final analytic sample, including 20 who were transferred or escalated to higher-intensity care during the treatment episode. A full comparison of included versus excluded cases was not feasible because structured demographic, clinical, and outcome data were not consistently available across the excluded referral stages, particularly for referrals that did not proceed to baseline assessment or the target intervention pathway. The complete participant flow is presented in
Figure 1, and safety-related events and higher-intensity care escalation are summarized in
Supplementary Table S3.
2.3. Clinical Assignment and Decision-Making
Clinical assignment was not randomized but was determined as part of routine clinical decision-making within the school-based service system. Assignment occurred before outcome evaluation and independently of the present analyses. All participants underwent a standardized intake process, including structured suicide risk assessment, clinical interviews, and caregiver consultation, through which relevant clinical and contextual information was documented.
Assignment to SPT-SAFE or DBT-BI was based on multiple clinically relevant considerations rather than a single predefined criterion. These considerations included the severity and pattern of suicidal ideation and self-injury, behavioral dysregulation and impulsivity, need for structured behavioral stabilization, capacity for verbal versus symbolic expression, developmental and communication characteristics, prior treatment history, and comorbid emotional or behavioral difficulties. Contextual factors, including caregiver preference, school coordination needs, therapist availability, and service-level considerations, were also taken into account.
Adolescents presenting with acute suicide risk requiring immediate stabilization, such as imminent suicidal intent, medically serious suicide attempts, or need for emergency intervention, were not allocated to either intervention track and were referred to higher-intensity services according to standard clinical protocols. Accordingly, the present sample reflects adolescents considered clinically appropriate for outpatient-level structured intervention following initial risk screening and triage.
Baseline demographic and clinical characteristics relevant to treatment assignment are summarized in
Table 1. These variables were examined to describe baseline comparability between groups, not to establish exchangeability. Because treatment assignment was embedded in routine clinical judgment, comparability on measured variables does not eliminate the possibility of clinically meaningful differences in unmeasured or partially documented factors. In particular, factors such as behavioral dysregulation, impulsivity, expressive style, or family and school context may have influenced both clinical assignment and subsequent treatment response.
Because these same clinical features may also predict subsequent change in impulsivity and behavioral regulation, the assignment process introduces a specific risk of confounding by indication. In particular, adolescents assigned to DBT-BI may have been those whom clinicians perceived as requiring more structured behavioral stabilization because of greater behavioral dysregulation, crisis instability, or readiness for skills-based work. Therefore, baseline comparisons were used only to describe measured group differences at intake and were not interpreted as demonstrating exchangeability between the intervention groups.
Although efforts were made within the service system to maintain a clinically balanced distribution of cases across intervention tracks, the non-randomized nature of assignment introduces the possibility of residual confounding, including confounding by indication. Accordingly, observed between-group differences are interpreted as observational associations within a naturalistic clinical context rather than as evidence of causal effects, comparative treatment superiority, or differential efficacy between interventions.
2.4. Interventions
Both interventions were delivered individually in weekly sessions of approximately 45–50 min. In the final analytic sample, participants received 6–12 sessions, with a mean dose of approximately 9–10 sessions in each intervention pathway.
2.4.1. Shared Safety Procedures
Before intervention-specific treatment, all adolescents received standardized safety-oriented procedures as part of routine school-based suicide risk management. These procedures included structured suicide/NSSI risk assessment, individualized safety planning, session-by-session monitoring of suicidal ideation and NSSI-related risk indicators, caregiver/school coordination when clinically indicated, and referral or escalation to psychiatric or emergency services when acute risk was identified. These safety procedures were implemented across both intervention tracks and were not specific to either SPT-SAFE or DBT-BI.
2.4.2. SPT-SAFE
SPT-SAFE was delivered as a school-based adaptation of nondirective sandplay therapy that integrated ongoing clinical attention to suicidal ideation and NSSI risk. This format reflected the routine school-based crisis-oriented service context, in which adolescents required both a protected symbolic-expressive space and continuous safety monitoring. Sessions generally included: (1) suicide/NSSI risk monitoring and emotional check-in; (2) tactile grounding or session introduction through contact with sand; (3) construction of a sand scene using miniature figures; (4) reflective observation of the sand scene; (5) optional verbal sharing or meaning exploration guided by the adolescent’s readiness; (6) review of emotional state and safety status; and (7) documentation of session content, safety status, and clinically indicated follow-up actions.
Therapists maintained core nondirective sandplay principles, including an empathic and nonjudgmental witnessing stance, respect for symbolic expression, and avoidance of imposed interpretation. Adolescents were not required to verbalize internal thoughts, painful emotions, shame-related experiences, suicidal content, or the personal meaning of symbolic scenes unless they chose to do so. However, therapists maintained explicit safety monitoring throughout the intervention and shifted to safety-oriented clinical procedures, including risk assessment, stabilization when needed, caregiver/school coordination, and referral or escalation to higher-intensity services when acute risk indicators emerged. Accordingly, SPT-SAFE should be understood as a safety-integrated, school-based adaptation of nondirective sandplay therapy, not as sandplay therapy delivered without suicide/NSSI risk management.
2.4.3. DBT-BI
DBT-BI was delivered as a brief, school-based, DBT-informed individual intervention rather than as the full multicomponent DBT-A package. This abbreviated format reflected the routine school-based service context, in which adolescents were referred for crisis-oriented assessment and short-term intervention under constraints of limited time, caregiver–school coordination, and ongoing safety monitoring. Sessions therefore focused on DBT-informed individual-session procedures most directly relevant to immediate risk stabilization: (1) brief mindfulness-based stabilization; (2) review of suicidal ideation and NSSI using diary-card or monitoring information; (3) prioritization of target risk behaviors; (4) DBT-based chain analysis of recent risk episodes; and (5) collaborative practice of selected skills targeting identified risk behaviors. Skills practice was collaboratively selected according to the adolescent’s target risk behaviors and immediate stabilization needs, drawing as appropriate on emotion regulation, distress tolerance, problem solving, and impulse-control strategies. DBT-BI did not include weekly multifamily skills training groups, between-session telephone coaching, a formal DBT therapist consultation team, or a long-term comprehensive DBT-A treatment structure. Accordingly, DBT-BI should be understood as a DBT-informed brief intervention adapted to a school-based crisis-oriented service context, not as a test of comprehensive DBT-A.
2.4.4. Therapist Teams, Supervision, and Fidelity/Adherence Monitoring
To reduce treatment contamination, SPT-SAFE and DBT-BI were delivered by separate therapist teams with no provider overlap. Each intervention was delivered by five licensed clinical or counseling psychologists. Across the final analytic sample, this corresponded to an average cumulative load of approximately 10–12 participants per therapist within each intervention track during the study period; this should not be interpreted as a concurrent caseload estimate. Model-specific supervision was provided through regular case consultation meetings focused on adherence to core intervention principles, safety monitoring, and clinical risk management. Manual-derived fidelity/adherence checklists for SPT-SAFE and DBT-BI are provided in
Supplementary Tables S1 and S2, and approximately 20% of sessions were independently reviewed for adherence. However, because this was a retrospective clinical record study, session-level fidelity ratings were not systematically compiled into a quantitative database. Accordingly, aggregate fidelity scores cannot be reported and therapist-level effects cannot be formally modeled.
2.5. Outcome Measures
The assessment battery and scoring procedures followed the same general framework as our related elementary-school retrospective observational study (
Kwak & Ahn, 2026). For all measures, higher scores indicated greater levels of the corresponding construct unless otherwise stated. For interpretive clarity, the seven outcome measures were grouped into four clinically defined domains prior to analysis: self-harm-related outcomes (SIQ-JR, FASM), emotional distress (CES-DC, STAI-T), behavioral dysregulation (AQ, BIS), and self-related functioning (PHCSCS). This grouping was based on the clinical function represented by each measure and on the theoretical relevance of these domains to school-based intervention planning for adolescents with suicidal ideation and NSSI.
The FASM was used as a dimensional self-injury frequency measure, not as a DSM-based diagnostic instrument for NSSI disorder. In this retrospective clinical record review, NSSI was defined as self-injurious behavior documented as occurring without suicidal intent. Inclusion in the analytic sample was determined retrospectively from intake clinical records, clinical interviews, documented risk assessments, and FASM suicidal intent follow-up item responses, rather than from FASM frequency responses alone. Cases with documented suicidal intent, imminent suicidal intent, or medically serious suicide attempts were referred to higher-intensity services in routine care and were not retained in the analytic sample. Because some behaviors included in the original FASM may require clinical contextualization to distinguish self-injurious acts from socially sanctioned body modification, FASM scores were interpreted as a dimensional self-injury frequency index rather than as a DSM-based diagnosis.
2.5.1. Primary Outcomes
Suicidal ideation. Suicidal ideation was assessed using the Suicidal Ideation Questionnaire–Junior (SIQ-JR;
Reynolds, 1988), a 15-item self-report measure rated on a 7-point scale. Higher scores indicate greater suicidal ideation. A Korean adolescent version has been reported (
Y. S. Lee et al., 2004).
NSSI. NSSI was assessed using the Functional Assessment of Self-Mutilation (FASM;
Lloyd, 1997), a self-report measure of self-injurious behaviors and their associated functions. A Korean version was used in the present study (
Kwon & Kwon, 2017). For the present analyses, an overall NSSI frequency index was computed by summing method-specific frequency ratings, with higher scores indicating greater NSSI frequency.
2.5.2. Secondary Outcomes
Depressive symptoms. Depressive symptoms were assessed using a Korean-language version of the 20-item CES-DC, adapted from the original CES-D for children and adolescents (
Radloff, 1977;
Faulstich et al., 1986). Higher scores indicate greater depressive symptoms.
Trait anxiety. Trait anxiety was assessed using a Korean-language version of the 20-item Trait Anxiety scale of the State-Trait Anxiety Inventory for Children (STAI-T;
Spielberger et al., 1973). Items are rated on a 3-point scale, with higher scores indicating greater trait anxiety.
Aggression. Aggression was assessed using the Aggression Questionnaire (AQ;
Buss & Perry, 1992). The Korean version consists of 27 items rated on a 5-point scale, with higher scores indicating greater aggression. Acceptable internal consistency has been reported in Korean samples (
Seo & Kwon, 2002).
Impulsivity. Impulsivity was assessed using a 23-item shortened Korean version of the Barratt Impulsiveness Scale-11 (BIS-11;
Patton et al., 1995). Items are rated on a 4-point scale, yielding a total score range of 23–92, with higher scores indicating greater impulsivity. The Korean BIS-11 has demonstrated acceptable reliability and validity in prior Korean samples (
S. R. Lee et al., 2012).
Self-concept. Self-concept was assessed using a modified PHCSCS-based self-concept index consisting of 30 items. Item content was retained from the Piers–Harris self-concept item pool (
Piers & Herzberg, 2002), drawn from Korean-translated PHCSCS/PHCSCS-2 materials informed by prior Korean adaptation work (
Choi, 2013). The 30-item short-form structure was informed by
Veiga and Leite (
2016). The response format was changed from the original dichotomous format to a 6-point Likert scale. In the present sample, internal consistency was good (Cronbach’s α = 0.88). Scores were computed by summing item responses; higher scores indicate a more positive self-concept. Because this exact format has not been independently validated, self-concept findings are interpreted only as exploratory secondary outcome data and are not compared with PHCSCS-2 norms or cutoffs.
2.6. Statistical Analysis
All analyses were conducted using IBM SPSS Statistics, version 25. Baseline differences between the two intervention groups were examined descriptively using independent-samples t-tests; Welch’s t-test was applied when the homogeneity of variance assumption was not met. Because intervention assignment was not randomized, these baseline comparisons were not treated as evidence of exchangeability between groups, but rather as descriptive indicators of the magnitude of differences on measured variables at intake.
The primary longitudinal analyses used 2 × 2 mixed-design analyses of variance (ANOVAs) for each outcome, with Group (SPT-SAFE vs. DBT-BI) as the between-subjects factor and Time (pre- vs. post-intervention) as the within-subjects factor. The effect of principal interest was the Group × Time interaction, which was used to examine whether patterns of pre–post change differed between intervention groups across outcome domains.
To complement these analyses, one-way analyses of covariance (ANCOVAs) were conducted for each post-intervention outcome, with group entered as the fixed factor and the corresponding baseline score entered as a covariate. These ANCOVAs were intended as supplementary baseline-adjusted analyses rather than as the primary inferential framework. Accordingly, if the two analytic approaches diverged, interpretation prioritized the Group × Time interaction model.
Because treatment assignment was clinically determined, baseline-adjusted ANCOVAs were not interpreted as removing confounding or establishing causal effects. They were used only as supplementary sensitivity analyses to examine whether the observed pattern was similar after adjustment for the corresponding baseline outcome score. Propensity-score adjustment was considered but not implemented as a primary analytic strategy because several clinically relevant assignment factors—such as crisis instability, clinician judgment regarding need for structured stabilization, family engagement, treatment motivation, and therapist availability—were not systematically available in the retrospective dataset. Under these conditions, a propensity-score model based only on a limited set of measured baseline variables could have provided a misleading impression of causal adjustment. Caregiver preference was recorded only as a contextual assignment consideration rather than as a structured, analyzable variable, and related constructs such as family engagement, therapeutic alliance, and treatment motivation were not systematically measured. Consequently, the magnitude of caregiver-preference-related self-selection bias could not be defensibly modeled; qualitatively, its direction is ambiguous because greater engagement among families preferring one intervention could amplify apparent improvement, whereas preferences driven by greater perceived severity or crisis concern could bias estimates in the opposite direction.
Effect sizes for the ANOVA and ANCOVA models were reported as partial η
2. Statistical significance was set at
p < 0.05, two-tailed. Because multiple psychological outcomes were examined in this exploratory retrospective analysis, findings were interpreted cautiously, with emphasis on multiplicity-adjusted conclusions rather than isolated nominal
p-values. Bonferroni and Benjamini–Hochberg false discovery rate procedures were applied as post hoc sensitivity checks for the seven Group × Time interaction tests and the seven supplementary ANCOVA group effects. Findings that did not remain significant after these multiplicity checks were interpreted as uncorrected exploratory signals rather than confirmatory evidence. To provide precision estimates,
Supplementary Table S5 reports within-group mean changes, within-group paired Cohen’s d, and between-group change differences with 95% confidence intervals for all seven outcomes.
Linear mixed-effects models were considered in response to the concern regarding repeated-measures ANOVA. However, the present dataset included only two assessment time points and a complete-case analytic sample. Under this data structure, the fixed Group × Time contrast estimated in a basic linear mixed-effects model with a subject-level random intercept would be expected to yield conclusions substantively similar to those from the mixed-design ANOVA, whereas the principal limitations of the study arise from non-randomized clinical assignment, multiplicity, and unmeasured confounding rather than from the repeated-measures model structure itself. We therefore retained the mixed-design ANOVA as the primary descriptive longitudinal model and used baseline-adjusted ANCOVAs and expanded covariate-adjusted sensitivity analyses as supplementary checks. Missing data were handled using a complete-case approach for the relevant analyses.
Missing post-intervention data were not assumed to be missing completely at random, as the absence of post-treatment assessment may have reflected early discontinuation, incomplete assessment, disengagement, or escalation to higher-intensity care. Because post-intervention outcomes were unavailable for excluded cases and structured baseline data were inconsistent across exclusion stages, formal MCAR testing and imputation were not considered defensible. Complete-case findings should therefore be interpreted as applying to treatment-engaged completers rather than to all referred or treatment-initiating adolescents.
4. Discussion
4.1. Overall Outcome Patterns
The present study examined outcome patterns associated with SPT-SAFE and DBT-BI in adolescents with suicidal ideation and non-suicidal self-injury in a school-based clinical setting. Across outcome domains, both interventions were associated with significant pre–post improvements, including reductions in suicidal ideation, NSSI frequency, and emotional distress. Between-intervention differences were generally limited.
An uncorrected Group × Time interaction was observed only for impulsiveness, with the DBT-BI group showing a descriptively larger reduction than the SPT-SAFE group. However, this effect was small in magnitude and did not remain statistically significant after Bonferroni or false discovery rate correction. Accordingly, the impulsivity finding should be interpreted as a preliminary, hypothesis-generating signal rather than evidence of differential efficacy.
Taken together, the findings indicate broadly comparable overall changes across the two intervention tracks, with limited evidence of systematic between-intervention differentiation. This pattern is consistent with a broader body of research showing that psychological interventions for adolescent self-harm and suicidality typically produce modest but clinically relevant improvements, often with limited differentiation between active treatment conditions (
Fox et al., 2020;
Ougrin et al., 2015;
Wright-Hughes et al., 2025). In this context, the present results are best understood not as evidence of treatment superiority, differential effects across outcome domains, or equivalence between the two pathways, but as short-term improvements observed in both groups with limited evidence of between-group differences.
4.2. Shared Intervention Elements and Clinical Viability in a Higher-Severity Adolescent Sample
The broad within-group improvements observed across both intervention tracks are notable, particularly given the elevated clinical severity of the present middle- and high-school sample. Both SPT-SAFE and DBT-BI were associated with reductions in suicidal ideation, NSSI frequency, depressive symptoms, anxiety, and aggression, as well as improvements in self-concept, despite the relatively brief duration of the interventions and higher baseline symptom levels compared with the elementary-school sample examined in our prior retrospective study (
Kwak & Ahn, 2026).
Both intervention tracks incorporated shared structural elements, including structured safety monitoring, therapeutic engagement, and caregiver–school coordination. These elements are widely recognized as central components of adolescent self-harm intervention and school-based mental health care (
Fazel et al., 2014;
Ougrin et al., 2015;
Meza et al., 2023). These shared features may account for a substantial portion of the observed improvements across outcome domains, independent of modality-specific components.
The observed immediate pre–post improvements suggest short-term clinical viability of both intervention tracks within the school-based service context. However, because no follow-up assessments were available, the present study cannot determine whether these changes were maintained, whether post-treatment risk re-escalation occurred, or whether either intervention provided longer-term clinical utility. The absence of broad between-intervention superiority should therefore not be interpreted as evidence of clinical inefficacy; rather, it should be understood within the limits of immediate post-treatment outcomes and in light of prior evidence that multiple active interventions can produce comparable levels of meaningful change in adolescent self-harm populations (
Fox et al., 2020;
Wright-Hughes et al., 2025).
4.3. Limited Between-Intervention Differentiation and a Behavioral Regulation Signal
A nominal, uncorrected Group × Time interaction was observed for impulsiveness (BIS; partial η2 = 0.042), with the DBT-BI group showing a descriptively larger reduction than the SPT-SAFE group. However, this effect was small and did not remain statistically significant after Bonferroni correction (adjusted α = 0.007) or Benjamini–Hochberg false discovery rate correction across the seven interaction tests. Accordingly, this finding should not be interpreted as evidence of differential efficacy, domain-specific comparative efficacy, or a treatment-specific mechanism. Rather, it should be regarded as a small, uncorrected, exploratory signal for future hypothesis generation.
The four-outcome grouping used in this study was therefore an organizing structure for description only, not a framework confirmed by the present data. Theoretical links between DBT-informed skills and behavioral regulation are discussed only as background context, not as an explanation for the observed pattern (
Linehan, 1993;
Rathus & Miller, 2002). Confounding by indication also remains a plausible alternative explanation: adolescents perceived by clinicians as needing more structured behavioral stabilization may have been more likely to receive DBT-BI, and such clinical indications may also be related to short-term changes in impulsivity. Regression to the mean, measurement variability, and other unmeasured clinical differences may also have contributed to the observed BIS pattern. No other outcome showed a multiplicity-adjusted Group × Time interaction.
Taken together, the findings indicate broadly comparable improvement across most outcomes, with only a small, uncorrected exploratory signal emerging for impulsiveness.
This pattern may inform future research, but it does not establish differential treatment effects or domain-specific comparative efficacy.
4.4. Cautions and Directions for Future School-Based Intervention Research
The present findings do not provide a sufficient basis for using a domain-sensitive perspective as a framework for intervention planning. Because adolescents with suicidal ideation and NSSI present with heterogeneous clinical profiles, describing change across multiple outcomes may, at most, serve as a starting point for future hypothesis generation rather than as a basis for prescriptive or age-based treatment matching. Cross-cohort observations relative to the prior elementary-school study (
Kwak & Ahn, 2026) are noted only as an exploratory context and are not interpreted as evidence of age-specific treatment effects or confirmed developmental trajectories.
4.5. Strengths and Limitations
The present study has several notable strengths. First, it was conducted within a naturalistic school-based clinical setting, which enhances the ecological validity of the findings and their relevance to real-world intervention planning. Unlike efficacy trials conducted under highly controlled conditions, the present findings reflect outcome patterns observed under the operational constraints typical of school-based mental health services, including limited session duration, ongoing safety monitoring, and coordination with caregivers and school personnel. Second, the use of multiple validated outcome measures across distinct functional domains—including suicidal ideation, NSSI frequency, behavioral regulation, emotional distress, and self-concept—allowed for a domain-sensitive examination of intervention-associated change beyond global symptom indices. Third, the present study extends a prior retrospective observational study conducted within the same service system (
Kwak & Ahn, 2026), providing a clinically grounded basis for examining whether outcome patterns may vary across developmental and severity contexts. Fourth, both intervention tracks were implemented within a structured clinical decision-making framework with documented safety protocols, enhancing the internal consistency of the service delivery context.
The most important limitation is the retrospective, non-randomized design and the resulting risk of confounding by indication. Intervention assignment was embedded in routine clinical decision-making rather than random allocation. Although measured baseline differences were described using SMDs and supplementary baseline-adjusted analyses were conducted, these approaches cannot establish exchangeability or rule out unmeasured selection factors. In particular, adolescents assigned to DBT-BI may have been perceived as having greater behavioral dysregulation, higher observed or clinician-perceived impulsivity, greater crisis instability, or a stronger need for structured behavioral stabilization, and these same factors may have contributed to the observed reduction in impulsivity. Confounding by indication therefore remains the most plausible specific alternative explanation for the nominal BIS finding and should be considered before any causal or intervention-specific interpretation is drawn. Accordingly, between-group comparisons should be interpreted as exploratory and not as evidence of comparative efficacy. Because the final analytic sample was a treatment-engaged completer sample rather than a representative sample of all school-linked referrals or treatment-initiating adolescents, selection bias and complete-case bias cannot be ruled out. Adolescents who discontinued treatment, had incomplete post-intervention assessments, or were escalated to higher-intensity care were not retained in the analytic sample. Therefore, the present findings may systematically underestimate adverse trajectories and should not be interpreted as sufficient evidence of safety in the broader treatment-initiating population. Second, only immediate post-treatment outcomes were available; therefore, conclusions about durability, post-treatment risk re-escalation, and longer-term clinical utility cannot be drawn. In this population, immediate post-intervention scores alone cannot determine whether clinical gains were maintained or whether self-harm or suicide-related risk re-escalated after treatment ended. Future studies should include structured follow-up assessments to evaluate the durability of school-based intervention effects. Third, the study was conducted within a single provincial, school-linked service system in South Korea, which limits generalizability to other cultural, clinical, and institutional contexts. Although the center functioned as a regional referral and intervention hub receiving students from urban, suburban, and rural schools across Chungcheongnam-do, transferability to other settings may depend on the presence of comparable school referral pathways, brief intervention structures, suicide-risk monitoring procedures, caregiver–school coordination, therapist training, and supervision systems. In addition, SPT-SAFE is grounded in established sandplay therapy principles but was structured for use within a school-linked suicide and NSSI intervention context; therefore, the transportability of this structured school-linked application of sandplay therapy should be examined through independent replication in other sites and cultural contexts, preferably using prospective designs. Fourth, therapeutic processes and mechanisms of change were not directly assessed, leaving open the question of whether the observed domain-specific patterns reflect the theoretically proposed components of each intervention or nonspecific factors. Fifth, although Bonferroni and Benjamini–Hochberg false discovery rate procedures were applied as post hoc sensitivity checks, the study was not designed around a prespecified multiplicity-controlled primary outcome. Therefore, isolated nominal findings, particularly the BIS interaction, remain vulnerable to Type I error and should be interpreted cautiously.
The study may also have been underpowered to detect clinically meaningful Group × Time interaction effects across multiple outcomes. Therefore, non-significant Group × Time interactions should not be interpreted as evidence of equivalence or comparable efficacy between SPT-SAFE and DBT-BI. Such findings indicate only that robust differential effects were not detected in this retrospective sample; clinically meaningful differential effects may have been missed, reflecting the risk of Type II error. This concern should be considered alongside the risk of Type I error for the nominal BIS finding, which did not survive multiplicity correction. Future prospective studies should include a priori power calculations for interaction effects and recruit larger samples capable of evaluating differential treatment effects with adequate precision.
A further limitation concerns treatment fidelity and adherence assessment. Although manual-derived fidelity/adherence checklists were developed for both SPT-SAFE and DBT-BI and used to guide supervision and monitoring, fidelity assessment was limited by the retrospective nature of the study. The archival dataset did not include full-session recordings, validated external fidelity instruments, fidelity ratings for every session, or systematically compiled quantitative fidelity scores. Therefore, we cannot fully rule out the possibility that observed outcome patterns were influenced by therapist-related differences, nonspecific therapeutic factors, or shared safety-management procedures rather than intervention-specific components. Future prospective studies should incorporate fully manualized intervention procedures, documented therapist training, systematic supervision logs, independent fidelity ratings, and session-level adherence monitoring across a larger proportion of sessions. Another limitation concerns the modified self-concept measure. Although no newly written items were introduced and internal consistency was good in the present sample, the 30-item, 6-point Likert format has not been independently validated; therefore, self-concept findings should be interpreted only as exploratory secondary outcome data and not as equivalent to PHCSCS-2 norm-based scores or cutoffs.
Taken together, these considerations indicate that the present findings are best understood as preliminary and hypothesis-generating. Future research should employ prospective or randomized designs with adequate statistical power, incorporate follow-up assessments, and include process-level measures to clarify how developmental stage, clinical severity, and functional domains interact in shaping intervention response in adolescents with suicidal ideation and non-suicidal self-injury.