Next Article in Journal
What’s Happening in the Exam Room? A Mixed-Methods Study About the Provision of Patient-Centered Contraceptive Care for Baltimore Latine Patients
Previous Article in Journal
Potential Benefits of Complementary Therapies for Women with Breast Cancer Undergoing Oncological Treatment: A Systematic Review
Previous Article in Special Issue
Bringing Psilocybin-Assisted Therapy to Palliative Oncology: Early Lessons from Real-World Implementation
 
 
Brief Report
Peer-Review Record

Developing Methods for Observing Awe Narration in Psilocybin-Assisted Therapy

Healthcare 2026, 14(11), 1589; https://doi.org/10.3390/healthcare14111589
by Elise C. Tarbi 1,2,*, Ian Bhatia 2, Nabil Balach 2, Suzannah Buehler 2, Magdalena Demeo-Meres 2, Cailin Gramling 2,3, Tej Thambi 4, Julia Hart 2, Maija Reblin 2, Donna M. Rizzo 5, Robert Gramling 2, Manish Agrawal 6 and Emily Manetta 7
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Reviewer 3: Anonymous
Healthcare 2026, 14(11), 1589; https://doi.org/10.3390/healthcare14111589
Submission received: 9 April 2026 / Revised: 27 May 2026 / Accepted: 2 June 2026 / Published: 5 June 2026
(This article belongs to the Special Issue Psychedelic Therapy in Palliative Care)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

This brief report develops and tests a direct observation coding system for “awe narration” within psilocybin-assisted therapy (PAT) sessions. Using recordings from 32 encounters involving 8 participants, the authors operationalize awe through three features: vastness, need for accommodation, and ineffability. The paper’s core contribution is methodological: it provides an initial human-annotation framework that could later serve as training data for AI/NLP models in psychotherapy research.

  • This is a genuinely novel contribution, particularly at the intersection of:
  • Psychedelic therapy mechanisms
  • Communication science
  • Computational psychiatry

From experience, most PAT studies stop at self-report scales or outcome measures. This paper instead tries to capture what actually happens inside the session, which is exactly where the field is currently weakest.

That said, the manuscript is best understood as foundational and exploratory, not definitive. Its impact will depend heavily on whether the authors can demonstrate scalability, reliability, and clinical relevance in follow-up work.

  1. You report 246 “awe moments” from 8 participants, but the analysis treats these moments as independent observations.this is a classic issue in observational coding studies:
  • Moments are nested within sessions, and sessions within participants
  • Ignoring clustering inflates statistical confidence (especially ORs and p-values) I would recommend the following
  • Reanalyze using multilevel modeling (e.g., mixed-effects logistic regression)
    • Random intercepts for participant (minimum)
  • Or at least acknowledge this explicitly as a statistical limitation
  1. The manuscript report:
  • 58% agreement between coders
  • Confidence levels

But it do not report standard reliability metrics (e.g., Cohen’s kappa, ICC).

In coding research, this is a red flag. Percentage agreement alone is not sufficient especially with subjective constructs like awe.

  • Report Cohen’s kappa (or Krippendorff’s alpha) for:
    • Presence/absence of awe
    • Each feature (vastness, ineffability, accommodation)
  • If temporal overlaps are complex, explain how reliability was operationalized

 

  1. the manuscript are careful to state that you measure narration, not internal experience but the manuscript occasionally blurs this distinction. From experience, this is a common trap in psychotherapy research:
  • Patients may experience awe without articulating it
  • Conversely, narration may be shaped by therapist prompting, language ability, or personality

Recommendation:

  • Strengthen the conceptual boundary:
    • Add a paragraph explicitly contrasting:
      • Experienced awe vs narrated awe
  • Discuss implications for validity (construct vs observable proxy)

 

  1. You acknowledge difficulty operationalizing this and your solution (splitting into disruption + optional accommodation) is sensible.
  • Provide clearer decision rules in Table 1:
    • What counts as disruption without accommodation?
    • What linguistic markers distinguish the two?
  • Consider including negative examples (what is not accommodation)
  1. The manuscript includes:
  • Odds ratios
  • Chi-square tests

But these analyses are:

  • Based on small N (participants)
  • Potentially inflated (see clustering issue)
  • Not central to the paper’s real contribution

Recommendation:

  • Either:
    • Strengthen statistical rigor (multilevel modeling), or
    • De-emphasize inferential statistics and frame results as descriptive feasibility findings

 

  1. All participants are:
  • Cancer patients
  • Mostly older (median 59)
  • Almost entirely white

Recommendation:

  • Expand the discussion on linguistic and cultural dependency
  • Explicitly state that the coding system may not transfer without adaptation
  1. Add a Figure: Coding Workflow

A simple diagram would help:

  • Input (video/audio)
  • Coding process
  • Features identified
  • Output (annotated moments)

This would significantly improve clarity for readers outside communication science.

 

  1. Improve Table 1

Table 1 is one of the best parts of the paper, but:

  • It mixes definitions and examples without structure
  • It lacks contrast cases

I would suggested improvements:

  • Add a column: “Common misclassification”
  • Add a column: “Key linguistic markers”
  • Separate:
    • Definition
    • Example
    • Coding cues
  1. Add a Reliability Table

Include:

  • Agreement metrics
  • Confidence levels
  • Feature-level agreement
  • Minor typographical issues (e.g., “Somtimes” in reference list) should be corrected
  • The term “boundary moments” is interesting but underexplained consider a brief example
  • Clarify how start/stop times overlap were judged for agreement
  • Some sentences in the Introduction are overly dense consider tightening for readability
  • Ensure consistency between “accommodation” vs “need for accommodation” terminology

Author Response

please see attachment

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

1.      Noted that this is a brief report, however, the title should be adjusted to show that this is a prelimiarny report/brief report. From the title, the readers may not know if this is a small/large cohort study.

2.      In the abstract “backrground”, the authors state that “Measuring what actually happens during PAT in large scale studies will be an essential component of this work”. However, this study only reports the observation from 8 participants, which is a very small study.

3.      The authors used the term “mechanism of action” and “mechanistics study” thoughout the manuscript, however, the finding of the study do not show any mechanistic action of PAT. The study is purely observational with no components of mechanism of action. Thus, the authors are advised to revise the texts to adhere to outcome of the study.

4.      The sample size is so small, and what are the control groups? How do the authors ensure validity of the outcome of analysis?

5.      The justification of patients selection should be explained further. Why cancer patients with depression? Why there’s no control group for comparison? At least is there a pre- vs post PAT comparison, why not?

6.      The discussion section should be expanded. Technically, the authors merely presented 2 paragraphs of discussion, with another paragraph explaining the limitation of current study. The discussion section should relate the current findings with previous literature. Why is this study needed? Has the objectives and hypothesis been met? How do this study add  to the knowledge/application of the field? What are the potential future studies to be conducted?

7.      The current study employed manual coding of the narration of awe. To be relevant to the current trend, the authors should at least mention in the discussion that how AI can help to replace manual coding in this study. What are the future plans? Is it feasible? What are the point of cautions or risk of application?

8.      Reference #21 is a proceeding from a conference/workshop. It should be specified in the reference list.

Author Response

see attached document

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

In the abstract, the main issue lies in the lack of precision and the overly strong tone of some statements. The objective is formulated somewhat generically (“to develop a direct observation coding system”), without clearly indicating whether this is a feasibility, development, or preliminary validation study. It would be more rigorous to specify this, for example: “to develop and preliminarily evaluate the feasibility and reliability of a direct observation coding system.” In the methods, the description is vague, particularly regarding the assessment of agreement between coders, which is referred to only as “coder agreement” without explaining how it was measured. Ideally, a more established metric should be mentioned, or at least the procedure should be clarified. In the results, there is also an important interpretative issue: the fact that 42% of the moments were identified by only one coder suggests moderate or limited agreement, but this is not addressed. It could be reformulated as: “agreement between coders was moderate, with 58% of events identified by both coders.” Finally, the conclusion is too assertive for a study with such a small sample and no external validation (“awe narration is directly observable”). It would be more appropriate to soften this: “our findings suggest that awe narration may be directly observable…”. There are also minor stylistic issues to correct, such as consistent use of hyphenation in “psilocybin-assisted therapy” and “large-scale.”

In the introduction, the main problem is the lack of a clearly articulated scientific gap. The text addresses several elements - therapy mechanisms, contextual complexity, advances in artificial intelligence, but does not directly state the specific gap the study aims to fill. It would be important to include a clear sentence such as: “however, there are currently no validated methods to directly observe and quantify subjective experiences such as awe during PAT.” In addition, the section on AI and natural language processing, while relevant, appears somewhat disconnected from the central objective and may give the impression of over-justification; it should be more directly linked to the need for human-annotated data. Another important point is the absence of an explicit hypothesis. Even in an exploratory study, it would be useful to state an expectation, for example: “we hypothesize that awe narration can be reliably identified using observable linguistic and behavioral markers.” The definition of “awe” also appears relatively late; it would be preferable to introduce it earlier to better structure the theoretical reasoning. Finally, the transition to the study objective could be stronger, explicitly linking the identified problem to the proposed solution, for example: “to address this gap, we developed…”. Overall, the introduction would benefit from a more linear structure: importance of PAT, problem of poorly understood mechanisms, methodological gap, definition of awe, and finally objective and hypothesis.

The methods section contains the most significant limitations. First, the classification of the study as “cross-sectional” is not entirely appropriate, as it is essentially an observational study with a qualitative component based on content coding; it would be more accurate to describe it as an “observational qualitative study with structured coding” or as a mixed-methods design. Second, the sample is very small (n=8), and the purposeful selection is not sufficiently justified; it would be important to specify inclusion criteria and the rationale for selection (e.g., diversity, theoretical saturation, etc.). One of the most relevant methodological issues is the absence of standardized inter-rater reliability metrics, such as Cohen’s kappa, ICC, or Krippendorff’s alpha. The use of “overlap” and “confidence” is insufficient to support the validity of the coding system, and it is strongly recommended to include one of these metrics or, at least, justify its absence. Furthermore, the unit of analysis (“moments”) is not clearly defined: the criteria for determining start and end points are unclear, which may introduce significant variability. The concept of “boundary moments” is interesting but lacks clear operationalization and explanation of how it was used in the analysis. The use of coder confidence levels also raises concerns, as it is a subjective measure that may introduce bias; it should be better justified based on the literature or complemented with more objective measures. Another critical point is the change in protocol midway through the study (focusing on dosing and integration sessions after the first four participants), which should be clearly justified as a methodological decision and discussed as a limitation. Finally, the statistical description is insufficient: odds ratios are reported, but the model used (e.g., logistic regression) and its assumptions are not specified. The study is also not explicitly framed as mixed-methods, despite combining qualitative and quantitative analysis, which should be clarified.

In the results, the presentation is generally clear, but there are some issues of consistency and analytical depth. The description of agreement between coders raises important concerns that are not adequately explored. The fact that only 58% of the moments were identified by both coders suggests moderate agreement, but this is neither critically discussed nor contextualized with benchmarks from the literature. Instead, the results are presented in a relatively neutral way, when in fact this is a key element for evaluating the robustness of the coding system. It would be important to explicitly acknowledge this limitation, for example: “Although agreement between coders was moderate, this highlights the inherent ambiguity of identifying complex emotional constructs such as awe.”

In addition, there are some inconsistencies in the reported numbers (e.g., 246 total moments versus 237 in the case of ineffability), which should be clarified to avoid confusion about data handling. The presentation of results could also benefit from clearer structure, for example by separating: (1) overall frequency of events, (2) inter-coder agreement, (3) distribution of features (vastness, ineffability, accommodation), and (4) statistical analyses. From a statistical perspective, although odds ratios are reported, there is a lack of information about the model used and no control for potential confounders, which limits interpretation. The association between vastness and coder confidence is interesting but may simply reflect greater perceptual salience rather than construct validity; this possibility should be acknowledged.

Another relevant point is the decision to change the analytical focus midway through the study (effectively excluding preparation sessions after the first participants), which directly affects the results. This decision is described but not critically examined. Ideally, it should be accompanied by a stronger justification and possibly a sensitivity analysis, or at least a clear note on how it affects generalizability.

In the discussion, the article offers interesting interpretations but tends to be somewhat optimistic given the limitations of the data. Identifying “vastness” as an entry point for recognizing moments of awe is a plausible contribution and is well aligned with the theoretical literature, but it may be biased by the fact that it is the most easily identifiable feature for coders. In other words, there is a risk of circularity: what is easiest to detect becomes the defining criterion. This should be explicitly acknowledged, for example: “It is possible that vastness emerged as a key feature due to its relative salience and ease of identification by coders.”

The proposal to reformulate “need for accommodation” into two components (cognitive disruption and possible accommodation) is one of the strongest aspects of the article, as it reflects a more nuanced analysis of the data. However, this interpretation could be better supported with additional examples or some quantification of how frequently these patterns occur. The discussion could also better integrate the findings with concrete clinical implications, for example how identifying moments of awe might inform therapeutic practice or the evaluation of psilocybin interventions.

Another aspect to improve is the link between the findings and the promise of scalability via computational methods. Although this connection is mentioned, it is not directly supported by the data, which are still at an early stage. It would be more appropriate to frame this as a future possibility rather than an immediate implication.

Limitations are acknowledged but somewhat incompletely. In addition to sample size and lack of diversity, it would be important to include other relevant limitations, such as the absence of standardized inter-rater reliability metrics, the subjectivity of confidence ratings, potential bias introduced by mid-study protocol changes, and ambiguity in the definition of units of analysis. Including these would make the discussion more balanced and credible.

In the conclusion, the main issue is again the overly assertive tone given the exploratory nature of the study. The claim that awe narration is “directly observable” and that coders can categorize these moments using explicit criteria may be too strong, given the observed level of agreement and the lack of external validation. It would be more appropriate to soften these statements, for example: “Our findings suggest that awe narration may be observable using explicit criteria, although further validation is needed.”

Furthermore, the conclusion could be more informative by clearly indicating next steps for research, such as testing the system in larger and more diverse samples, assessing reliability using standardized metrics, and exploring the relationship between narrated awe and clinical outcomes. This would provide a clearer sense of scientific progression.

Author Response

see attached document

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

No further comments, Thank you!

Author Response

We remain most grateful for your thorough review.

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

The authors have addressed my previous comments.

I have no further comments.

Author Response

Thank you very much for your insights which have strengthened our manuscript.

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

The article presents an original and intellectually robust proposal, but it would benefit from several important improvements in terms of conceptual clarity, methodological rigor, and theoretical positioning. First, although the study appropriately acknowledges its exploratory nature, the manuscript occasionally formulates conclusions with a level of assertiveness that exceeds what the data can reasonably support. Given the small sample size, the demographic homogeneity of the participants, and the fact that all participants belonged to the same therapeutic context, it would be advisable to moderate some of the claims, particularly in the conclusions, by replacing categorical formulations with more cautious language compatible with a preliminary study. For example, rather than stating that “awe narration is directly observable,” it may be more rigorous to suggest that the findings “indicate” or “support the feasibility” of the structured observation of awe experiences.

Another aspect that deserves further strengthening is inter-coder reliability. Since the manuscript is fundamentally focused on developing an observational coding system, the absence of a formal inter-rater agreement metric represents perhaps its main methodological limitation. Although the authors report co-identification rates and coder confidence levels, it would be important to include standardized measures such as Cohen’s kappa, Krippendorff’s alpha, or another equivalent strategy adapted to temporally bounded events. Even acknowledging the inherent challenges involved in coding continuous discursive phenomena, the inclusion of a formal reliability metric would substantially increase the credibility of the method and reinforce its future applicability in computational and machine learning contexts.

The concept of “need for accommodation” also appears somewhat unstable throughout the manuscript. The article addresses this issue honestly and productively, yet the operationalization remains relatively diffuse, particularly because it combines two distinct processes: the experience of cognitive disruption and the subsequent attempt to integrate or make sense of that disruption. The authors’ later reformulation separating “cognitive disruption” from “accommodation” is in fact one of the manuscript’s most interesting contributions, but it could be presented more explicitly as a conceptual refinement emerging from the data themselves. Rather than appearing merely as a technical adjustment to the codebook, this distinction could be emphasized as a central theoretical outcome of the study.

The discussion section could also benefit from greater conceptual depth. The manuscript implicitly engages with several phenomenologically rich themes, including ineffability, the dissolution of ordinary cognitive structures, the limits of language, and the difficulty of symbolizing intense experiences, yet it remains framed almost exclusively within a functional and methodological register. A brief engagement with phenomenological, hermeneutic, or existential psychological literature could significantly enrich the manuscript’s theoretical density without compromising its scientific objectivity. Such an expansion would also help situate the concept of awe beyond its operational utility as an observable variable.

At the level of writing style, the manuscript is generally clear, well organized, and enjoyable to read, particularly given the complexity of the topic. Nevertheless, some sentences are excessively dense and conceptually overloaded, especially in the introduction, where multiple abstract ideas are occasionally compressed into a single long sentence. Selective simplification of the syntax would improve readability and make the text more accessible to readers from different disciplinary backgrounds. It may also be useful to reduce the repetition of certain expressions related to “complex interactions,” “mechanisms,” and “context,” thereby making the argument more direct and concise.

Finally, the article would benefit both visually and conceptually from the inclusion of a schematic figure summarizing the final coding model. A simple diagram illustrating the relationship between vastness, cognitive disruption, accommodation, and ineffability would help consolidate the manuscript’s methodological contribution and make the coding system more readily usable by other researchers. This would be especially valuable given that the manuscript positions itself not only as an empirical study, but also as a foundational framework for future automated narrative analysis in psychedelic-assisted therapy.

Author Response

  1. [Conceptual clarity] “Given the small sample size, the demographic homogeneity of the participants, and the fact that all participants belonged to the same therapeutic context, it would be advisable to moderate some of the claims, particularly in the conclusions, by replacing categorical formulations with more cautious language compatible with a preliminary study. For example, rather than stating that “awe narration is directly observable,” it may be more rigorous to suggest that the findings “indicate” or “support the feasibility” of the structured observation of awe experiences.”

Response: We have taken these suggestions and adjusted language throughout where appropriate.

  1. [Inter-coder reliability] “Although the authors report co-identification rates and coder confidence levels, it would be important to include standardized measures such as Cohen’s kappa, Krippendorff’s alpha, or another equivalent strategy adapted to temporally bounded events.”

Response: We have added Cohen’s kappa to the Results.

  1. [Theoretical positioning] “The authors’ later reformulation separating “cognitive disruption” from “accommodation” is in fact one of the manuscript’s most interesting contributions, but it could be presented more explicitly as a conceptual refinement emerging from the data themselves. Rather than appearing merely as a technical adjustment to the codebook, this distinction could be emphasized as a central theoretical outcome of the study… The discussion section could also benefit from greater conceptual depth.”

Response: Text addressing the conceptual contribution of this study has now been added to the Discussion.

  1. [Writing style] “Selective simplification of the syntax would improve readability and make the text more accessible to readers from different disciplinary backgrounds. It may also be useful to reduce the repetition of certain expressions related to “complex interactions,” “mechanisms,” and “context,” thereby making the argument more direct and concise.”

Response: We have revised for readability, especially in the Introduction.

  1. Add a figure: “A simple diagram illustrating the relationship between vastness, cognitive disruption, accommodation, and ineffability would help consolidate the manuscript’s methodological contribution and make the coding system more readily usable by other researchers.”

Response: Thank you for this helpful suggestion. Figure 2 has been added to the Results.

Author Response File: Author Response.pdf

Back to TopTop