1. Introduction
Hypnotic responsiveness is commonly assessed using behavioral scales that rely on a formal hypnotic induction and index hypnotic susceptibility, defined as responsiveness to suggestions delivered following induction. In contrast, the Creative Imagination Scale (CIS) was developed to assess imaginative suggestibility without induction, focusing on individuals’ capacity for voluntary imaginative involvement in suggested experiences.
Rather than assessing hypnotic susceptibility in the induction-dependent sense, the CIS captures responsiveness to non-induction-based suggestions and thus reflects the broader construct of suggestibility (
Wilson & Barber, 1978). The CIS consists of ten standardized experiential suggestions—such as arm heaviness, imagined anesthesia, gustatory imagery, time distortion, and age regression—which respondents are instructed to visualize or experience through imagination.
Empirical studies and normative applications have demonstrated acceptable psychometric properties for the CIS, supporting its use as a valid measure of imaginative responsiveness to suggestion that is conceptually related to—but distinct from—hypnotic susceptibility (
Barber & Wilson, 1978;
Wilson & Barber, 1978).
In a series of initial investigations, normative data for the CIS were established, and the scale was shown to possess satisfactory psychometric properties, including test–retest reliability, split-half reliability, and factorial validity, supporting its coherence as a measure of imaginative responsiveness to suggestion (
Barber & Wilson, 1978;
Wilson & Barber, 1978).
Previous work has raised the question of whether imaginative suggestibility, as assessed by the CIS, is fully unitary. Factor-analytic investigations have suggested that CIS items may differentially reflect imagery-related absorption and performance-related perceptual alterations. In particular,
Hilgard et al. (
1981) reported a two-factor structure comprising an Absorption/Imagination factor and a Hypnotic Performance factor, with CIS items loading on both dimensions. At the same time, earlier descriptive and psychometric work had often treated the CIS as a largely unidimensional measure (e.g.,
Wilson & Barber, 1978).
Comparative research has further examined the construct validity of the CIS by comparing it with induction-based measures of hypnotic susceptibility. For example,
McConkey et al. (
1979) reported a modest positive correlation between CIS scores and scores on the Harvard Group Scale of Hypnotic Susceptibility, Form A (HGSHS:A; r = 0.28). Their findings indicate that, although the two measures share some variance related to suggestibility, they reflect largely distinct underlying dimensions, with the CIS primarily assessing imagery- and imagination-based processes.
Additional empirical work has reported acceptable internal consistency of the CIS and significant correlations with theoretically related constructs such as absorption, providing support for its reliability and convergent validity (
Sapp & Hitchcock, 2003). In their study, the CIS demonstrated a coefficient alpha of approximately 0.84 and showed significant correlations with scores on the Tellegen Absorption Scale (TAS).
Although classical studies have reported modest positive correlations between the CIS and the full HGSHS:A (e.g.,
McConkey et al., 1979), direct empirical correlations between the CIS and the recently developed short form HGSHS-5:G (Harvard Group Scale of Hypnotic Susceptibility—5 item short form, German version) have not yet been reported.
Riegel and colleagues proposed a 5-item short version of the HGSHS:A to reduce test time and facilitate broader use. In German samples, the HGSHS-5:G demonstrated acceptable validity, reliability, and classification agreement relative to the original HGSHS:A (
Riegel et al., 2021).
This reveals a research gap: although the HGSHS-5:G has been standardized in German as an economical measure of the motor challenge dimension of hypnotic susceptibility, a validated German version of the CIS is still lacking. Comparing both instruments in a clinical sample, such as individuals with mild to moderate depressive symptoms, may therefore provide important insights into their convergent and discriminant validity.
Although the CIS was originally developed to assess imaginative suggestibility in non-clinical contexts, examining the scale in individuals with depressive symptoms is theoretically meaningful because several CIS items rely on vivid imagination, absorption, and responsiveness to internally generated experiences. Mental imagery has been shown to evoke stronger emotional responses than verbal representation and may contribute to the maintenance of emotional disorders (
Holmes & Mathews, 2010). Furthermore, in individuals with depressive disorders, negative mental imagery induced stronger negative affect compared to healthy controls, and a positive imagery deficit was observed in the explicit measure (
Görgen et al., 2015). Given that depressive symptomatology is associated with imagery-related cognitive and affective processes, it may influence mechanisms relevant to responses on the CIS. Investigating CIS scores in individuals with mild to moderate depressive symptoms therefore provides an opportunity to examine the construct validity of the scale in a population in which mental imagery processes are relevant to depressive symptomatology. Importantly, the psychometric properties of suggestibility-related measures cannot be assumed to generalize from healthy to clinical populations, as depressive symptomatology is associated with alterations in cognitive, affective, and motivational processes that are directly relevant to imaginative engagement. Examining the CIS in a clinical sample thus allows for a more ecologically valid assessment of how imaginative suggestibility (CIS) and the motor challenge dimension of hypnotic susceptibility (HGSHS-5:G) differ in this population.
2. Materials and Methods
2.1. Participants and Procedure
Participants were recruited as part of the HypnoDeep randomized controlled trial, which examined the effects of telemedical hypnosis, progressive muscle relaxation, and routine care in individuals with mild to moderate depressive symptoms and elevated stress. Participants had clinically diagnosed depressive episodes, classified as mild according to ICD-10 criteria (International Classification of Diseases, 10th revision). Self-reported depressive symptoms at baseline, assessed with the Beck Depression Inventory–II (BDI-II), were in the mild to moderate range and were assessed as a continuous variable to characterize the sample. Perceived stress was assessed using the Perceived Stress Scale (PSS) and a visual analog scale for stress (VAS-Stress).
Recruitment was conducted in Germany via online advertisements, newsletters, and outpatient clinical networks. After an online screening for eligibility, participants were enrolled and provided written informed consent before participation.
In the original trial, participants were randomly assigned to one of 12 online groups: four received guided hypnotherapy, four progressive muscle relaxation, and four served as a wait-list control group without intervention. Questionnaires and demographic data were administered during the pre-intervention assessment phase. In total, 100 participants were enrolled and randomized in the HypnoDeep trial. Three participants were mistakenly randomized, did not complete any baseline assessments, and did not receive the allocated intervention. They were therefore excluded from all analyses. The present study is based on the remaining 97 participants, of whom 94 (97%) completed both the CIS and the HGSHS-5:G.
The final sample consisted of 69 women (71.1%), 26 men (26.8%), and 2 participants identifying as diverse (2.1%), with a mean age of 39.2 years (SD = 12.02; range = 20–68 years).
Inclusion criteria were: age between 18 and 70 years, mild to moderate depressive symptoms according to ICD-10 criteria (F32.0 or F33.0), and a self-reported stress level of ≥40 mm on a visual analog scale. Exclusion criteria included current or past psychotic disorder, post-traumatic stress disorder, personality disorder, acute suicidality, substance abuse, or severe medical illness that could interfere with participation.
2.2. Measures
2.2.1. CIS
The CIS (
Wilson & Barber, 1978) is a 10-item measure of imaginative suggestibility, designed to be administered without a formal hypnotic induction. Items tap into a range of imaginative experiences (arm heaviness, finger anesthesia, gustatory hallucinations, time distortion, age regression) and are rated on a 5-point scale according to participants’ subjective experience from 0 to 4 (0, not at all the same [experience]; 1, a little the same; 2, between a little and much the same; 3, much the same; 4, almost exactly the same), resulting in a total sum score between 0 and 40. A German CIS version based on the translation by
Bahlinger (
1995), which was used and documented in the diploma thesis by
Egbers (
1998; Imaginative Processing of Directly and Indirectly Formulated Texts Depending on Attentional Demand), was used in the present study. In the thesis, the full wording of the German CIS was provided in the appendix and the scale was used to assess imaginative ability. To our knowledge, however, no formal psychometric validation of this German version has previously been published. The translated items were administered in their original German wording as documented by
Egbers (
1998).
2.2.2. HGSHS-5:G—5-Item Short Form
The HGSHS-5:G (
Riegel et al., 2021;
Zech et al., 2024) is a recently validated German short version of the HGSHS:A (
Shor & Orne, 1962), containing five motor challenge items (arm immobility, finger lock, arm rigidity, head shaking inhibition, eye catalepsy). Items are scored dichotomously (0 = fail, 1 = pass), yielding a total score between 0 and 5.
2.3. Statistical Analyses
All analyses were conducted in R (version 4.5.1;
R Core Team, 2025) using RStudio (version 2025.05.1+513 Mariposa Orchid by Posit Software, 2025). The data were prepared, cleaned, and analyzed in a fully reproducible workflow with
tidyverse packages (
Wickham et al., 2019). Analyses were conducted in several steps corresponding to the psychometric evaluation of the German CIS and its comparison with the HGSHS-5:G. Data were imported from SPSS files originally generated using SPSS version 29 (IBM Corp.) via the
haven (
Wickham et al., 2025) and
here packages (
Müller, 2025) to ensure reproducible file paths.
Descriptive statistics: Descriptive statistics (means, standard deviations, skewness, and kurtosis) were computed for all variables with the
sjPlot (
Lüdecke, 2025a),
sjstats (
Lüdecke, 2025b) and
sjlabelled (
Lüdecke, 2022) packages. Missing values were visualized and inspected for patterns using
Amelia (
Honaker et al., 2011). For descriptive tables and summary outputs,
gtsummary (
Sjoberg et al., 2021) was used.
Item analysis: Item difficulty, discrimination, and corrected item–total correlations were calculated using
sjPlot (
Lüdecke, 2025a),
sjstats (
Lüdecke, 2025b),
performance (
Lüdecke et al., 2021a) and custom functions from the
strengejacke package (
Lüdecke, 2019). Mean inter-item correlations were inspected as an additional reliability indicator.
Reliability analysis: Internal consistency was assessed via Cronbach’s α and McDonald’s ω using the
psych package (
Revelle, 2025). For the CIS, both α and ω total were computed; for the HGSHS-5:G (dichotomous items), α was interpreted as equivalent to KR-20. Confidence intervals for α and item-deleted reliabilities were also examined.
Exploratory factor analysis (EFA): Sampling adequacy was tested with the Kaiser–Meyer–Olkin (KMO) statistic and Bartlett’s test of sphericity (
Revelle, 2025). To explore the dimensionality of the CIS and HGSHS-5:G, principal components analyses were conducted with
psych and
parameters (
Lüdecke et al., 2020). Principal component analyses with varimax rotation were run using
parameters, supplemented by maximum likelihood factor analysis on tetrachoric correlation matrices for the HGSHS-5:G.
Confirmatory factor analysis (CFA): Competing one- and two-factor models and a second-order CFA model of the CIS were evaluated using the
lavaan package (
Rosseel, 2012). CFAs were conducted using maximum likelihood estimation (ML), which is appropriate given the approximately symmetric response distributions (skewness: −0.25 to 0.37) and five response categories of the CIS. Model fit was judged according to established criteria (
Hu & Bentler, 1999), reporting χ
2 tests, CFI, TLI, RMSEA with 90% CI, and SRMR. Nested models were compared using χ
2 difference testing.
Visualization and reporting: Factor loadings, item characteristics, and reliability indices were visualized with
see (
Lüdecke et al., 2021b) and
ggplot2 (
Wickham, 2016;
Wickham et al., 2019). Results were summarized in APA-style tables using
knitr (
Xie, 2014,
2015,
2025) and
report (
Makowski et al., 2023). All analyses were fully reproducible, with package management handled through
pacman (
Rinker & Kurkiewicz, 2018) and selected GitHub packages (strengejacke/strengejacke, easystats/report, easystats/see, easystats/performance).
Correlation analysis: Correlations between ordinal CIS items and dichotomous HGSHS-5:G items were computed using mixed polychoric–tetrachoric–polyserial correlations implemented in the
psych::mixedCor() function (
Revelle, 2025). Correlations between the CIS and HGSHS-5:G total scores were computed using Pearson’s r and Spearman’s ρ, as both composite scores can be considered approximately continuous.
2.4. Sample Size Considerations
There is ongoing debate about the minimum sample size required for factor-analytic procedures. Rules of thumb typically recommend at least 5–10 participants per item (
Gorsuch, 1983), with absolute minimums of n = 100–200 for exploratory factor analysis (
MacCallum et al., 1999). For confirmatory factor analysis (CFA), many authors suggest N ≥ 200 as a general lower bound (
Kline, 2016), though the exact requirements depend strongly on factor loadings, the number of items per factor, and overall model complexity (
Wolf et al., 2013). Reliability analyses such as Cronbach’s α also require sufficient sample sizes for stable estimates, with N ≥ 100 considered adequate for exploratory work, but N ≥ 300 recommended for precise confidence intervals (
Bonett, 2002). Against this background, our sample size of approximately 100 participants can be considered sufficient for an initial validation, while acknowledging that larger replications will be required to confirm the CFA results and provide stable estimates of reliability.
4. Discussion
The present study represents the first validation of the German version of the CIS in a clinical sample of individuals with mild to moderate depressive symptoms. The study further compared the German CIS with the HGSHS-5:G to examine convergent and discriminant aspects of imaginative versus behavioral suggestibility. Establishing a validated German CIS was necessary, as no standardized measure of imaginative suggestibility currently exists for German-speaking clinical populations.
Overall, the results demonstrate that the German CIS shows good internal consistency and preliminary but limited evidence for a two-factor structure—though absolute CFA fit was suboptimal in the present sample, which should be interpreted in light of the relatively small sample size—alongside a moderate positive correlation with the HGSHS-5:G, supporting convergent validity while indicating that imaginative and behavioral suggestibility represent related but distinct constructs.
4.1. Psychometric Properties of the CIS
The German CIS demonstrated good internal consistency, comparable to or exceeding reliability estimates reported in earlier English-language studies of the CIS in experimental and clinical samples (
Wilson & Barber, 1978;
Sheehan et al., 1978;
Laidlaw & Large, 1997). Exploratory and confirmatory factor analyses indicated that a two-factor model provided a slightly better fit compared to a unidimensional structure. The factors, interpretable as Imagination/Absorption and Performance/Perceptual Alteration, are consistent with previous factor-analytic work suggesting that imaginative suggestibility encompasses multiple dimensions (
Hilgard et al., 1981;
Laurence et al., 2008). This is tentatively consistent with the view that the CIS may capture multiple dimensions of imaginative suggestibility, though the high interfactor correlation (r = 0.80) and suboptimal model fit in the present sample limit the strength of this conclusion.
Overall, these findings are broadly consistent with earlier psychometric and comparative work on the CIS, including studies reporting acceptable reliability as well as investigations of its factor structure and relationship to other suggestibility measures (e.g.,
Wilson & Barber, 1978;
Hilgard et al., 1981;
McConkey et al., 1979). The present study extends this work by providing initial validation data for a German clinical sample with depressive symptoms. The CFA results extend the EFA findings by providing preliminary support for a correlated two-factor solution as a better representation of the CIS structure than the unidimensional model. Neither model reached conventional thresholds for acceptable fit (CFI > 0.90, RMSEA < 0.08), which should be interpreted in light of the relatively small sample size. The two latent dimensions were highly correlated (r = 0.80), indicating substantial shared variance; they may therefore be better understood as correlated facets of a broader imaginative suggestibility construct than as independent dimensions. This is consistent with prior work on multidimensional structures of suggestibility scales (
McConkey et al., 1980;
Piesbergen & Peter, 2006;
Polito et al., 2014;
Zahedi et al., 2024) and suggests that a two-dimensional conceptualization warrants further investigation, though firm conclusions about subscale utility are premature pending replication in larger samples.
4.2. Psychometric Properties of the HGSHS-5:G
The HGSHS-5:G was confirmed as a reliable and unidimensional instrument, consistent with recent German validation studies (
Riegel et al., 2021;
Zech et al., 2024). Although internal consistency was only modest (α = 0.68), this level of reliability is acceptable given the brevity and heterogeneity of the five dichotomous motor challenge items. This supports the scale’s use as an efficient behavioral index of hypnotic responsiveness, even in clinical populations with mild depressive symptoms.
4.3. Convergent Validity and Distinctiveness
The CIS and the HGSHS-5:G showed a moderate positive correlation (r = 0.52), with a disattenuated correlation of approximately 0.68, indicating substantial shared variance alongside clear conceptual distinctiveness. This result is broadly consistent with earlier English-language studies comparing the CIS with the HGSHS:A. For example,
McConkey et al. (
1979) reported a modest positive correlation between CIS and HGSHS:A scores (r = 0.28), while demonstrating that the two instruments reflect largely independent underlying dimensions. Importantly, the HGSHS:A comprises a broader range of hypnotic suggestions, whereas the HGSHS-5:G focuses exclusively on motor challenge items. Differences in scale composition, sample characteristics, and clinical context may therefore partly account for the stronger correlation observed in the present study. Overall, these findings support the view that imaginative suggestibility, as assessed by the CIS, and behavioral hypnotic responsiveness under induction, as assessed by the HGSHS-5:G, are related but non-redundant constructs.
4.4. Clinical and Theoretical Implications
From a clinical perspective, these results are highly relevant. Hypnotic suggestibility has been linked to treatment outcomes in hypnotherapy (
Montgomery et al., 2011). In addition, hypnotherapy has been shown to be non-inferior to cognitive behavioral therapy in the treatment of mild to moderate depressive episodes (
Fuhr et al., 2021). The CIS may be particularly valuable as a brief, non-induction-based screening tool for imaginative abilities in patients with depression, where imagery-based interventions (e.g., guided imagery, hypnosis, or mental imagery training) have shown therapeutic potential (
Holmes et al., 2016;
Renner et al., 2017). For clinical use, the total CIS score currently represents the more robust index, as the preliminary nature of the two-factor solution and the high interfactor correlation (r = 0.80) warrant caution in interpreting subscale scores until replication in larger samples is available. The moderate convergence between CIS and HGSHS-5:G also indicates that both instruments tap into core mechanisms of suggestibility, but in complementary ways: one through imaginative absorption, the other through motor challenge performance. This suggests that the two measures may be differentially sensitive to individual differences in cognitive style and clinical symptomatology.
4.5. Limitations
A key limitation of the present study concerns the relatively modest sample size (N ≈ 94), which falls below many recommended thresholds for confirmatory factor analysis and reliability estimation (
Kline, 2016;
Wolf et al., 2013). For exploratory factor analysis specifically, the adequacy of a given sample size depends critically on the level of communality and factor overdetermination: while 100–200 cases may suffice when communalities are high (>0.60) and factors are well-defined, lower communalities and weakly determined factors increase the required sample to 300 or more (
MacCallum et al., 1999). In the present data, item communalities ranged from 0.43 to 0.67 (M = 0.55), suggesting that the EFA results in particular should be interpreted with appropriate caution. Although our sample size was sufficient to support exploratory analyses, the structural findings remain preliminary. Future research with larger and more diverse samples is needed to confirm the robustness of the two-factor structure of the CIS and to refine its psychometric properties. Moreover, because the EFA and the CFA were performed on the same dataset and the CIS comprises a small and heterogeneous set of items, the structural validation results should be interpreted with caution, as this design limits the robustness of the factor-analytic evidence.
4.6. Strengths
Despite these limitations, the present study has several notable strengths. It represents the first validation of the German CIS in a clinical sample of individuals with mild to moderate depressive symptoms, addressing a critical gap in suggestibility assessment for German-speaking populations. The study employed a rigorous psychometric approach, combining exploratory and confirmatory factor analyses and examining both internal consistency and convergent validity with the HGSHS-5:G. By integrating a non-induction-based measure of imaginative suggestibility (CIS) and a behavioral measure of hypnotic performance (HGSHS-5:G), the study provides a comprehensive and differentiated assessment of suggestibility, highlighting the multidimensional nature of the construct.