Caffeine Expectancy Does Not Affect Interval Running Performance, Physiological Responses, or Running Kinematics in Trained Runners: A Randomized Controlled Crossover Trial
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsEvaluation of Manuscript nutrients-4402065
The topic is relevant and relatively original, as most of the literature on caffeine-related placebo effects has been conducted using time-trial or performance-testing protocols, whereas the present study investigates an interval training session. Below, I provide several comments intended to help the authors improve the quality and scientific rigor of the manuscript.
The title is broad and suggests a general investigation of caffeine expectancy during interval training. However, the main finding of the study was the absence of significant differences in overall performance, physiological responses, and kinematic variables between conditions. I recommend that the authors adopt a more specific title that better reflects the actual findings. One possible alternative is: "Caffeine expectancy does not significantly improve interval running performance, physiological responses, or running kinematics in trained runners: an exploratory placebo study."
The introduction addresses relevant concepts related to the placebo effect and caffeine expectancy; however, the overall structure lacks logical progression. In several instances, the text reads as a sequence of previously published studies rather than a coherent narrative leading the reader toward the specific knowledge gap that the present study aims to address.
Furthermore, I believe that the rationale supporting the need for this study remains insufficient. The authors state that "given the scarce scientific literature on this topic, further investigations are necessary to consolidate these outcomes", but this claim is not adequately supported by the literature presented. Although the cited studies primarily investigated time-trial protocols and specific performance tests, the literature on caffeine expectancy and placebo effects in sport is broader than suggested in the introduction, encompassing different sports, experimental protocols, competitive settings, and outcome measures. Consequently, it remains unclear whether the primary gap relates specifically to interval training, the exercise model employed, the athlete population investigated, or the outcomes assessed. I recommend that the authors define more explicitly the unresolved knowledge gap and demonstrate, based on the available literature, why investigating the placebo effect of caffeine expectancy during interval training sessions represents a meaningful advancement beyond previous studies.
The experimental design is presented too briefly, making it difficult to understand both the sequence of procedures and the purpose of each variable collected. I recommend expanding the “Experimental Design” subsection to provide a clearer description of the study flow, experimental conditions, procedural order, and, most importantly, the rationale for including each outcome measure (performance, heart rate, kinematic variables, perceived exertion, and adverse effects). In addition, ethical aspects, such as ethics committee approval and informed consent procedures, should be presented in this initial methods section rather than appearing later within the participant description.
The inclusion and exclusion criteria are relatively generic and do not adequately characterize the eligible participant profile. Considering that the study aims to investigate the placebo effect associated with caffeine expectancy, it would be important to report potentially relevant characteristics such as habitual caffeine consumption, previous use of ergogenic supplements, competitive experience, and other factors that may influence susceptibility to placebo responses.
Moreover, given the small sample size (n = 14), it is essential to present a sample size calculation justifying the number of participants recruited, including the expected effect size, significance level, and statistical power considered.
The “Procedure” subsection could also be improved through the inclusion of a schematic representation of the experimental design. A flowchart detailing participant visits, the familiarization session, condition randomization, washout periods, placebo administration, assessments performed, and outcomes collected would facilitate understanding of the protocol and improve methodological transparency.
I also suggest that the authors clarify several additional methodological aspects that may influence the interpretation of the findings: (i) the method used to generate the randomization sequence and ensure counterbalancing of conditions; (ii) whether the assessors responsible for timing performance were blinded to the experimental condition; (iii) whether the effectiveness of the experimental manipulation was assessed, that is, whether participants actually believed they had ingested caffeine; and (iv) a more detailed description of environmental testing conditions, as variables such as temperature, humidity, and wind may influence performance during outdoor running sessions.
Regarding the statistical analyses, the authors should justify the analytical approach adopted for each outcome and provide evidence that the assumptions underlying the statistical models were adequately verified. Although the use of the Shapiro–Wilk test is reported, no information is provided regarding homogeneity of variances, sphericity in repeated-measures analyses, or any corrections applied when assumptions were violated.
Multiple analyses were conducted across intervals, splits within intervals, physiological variables, kinematic variables, and adverse effects. However, no strategy to control Type I error resulting from multiple comparisons was described. Consequently, some statistically significant findings may simply reflect chance findings.
I recommend that the authors report confidence intervals for both the differences between conditions and the effect sizes, allowing a more robust interpretation of the magnitude and precision of the observed effects.
Given the crossover design and the small sample size, the authors should also consider the use of linear mixed-effects models either as a replacement for, or as a complement to, the repeated-measures ANOVAs and multiple paired comparisons currently employed.
I also recommend greater caution in the interpretation of the individual results used to support the existence of a placebo effect. Although the authors emphasize that 71.4% of participants improved performance under the placebo condition and that the mean reduction in time exceeded the Smallest Worthwhile Change (SWC) threshold, the primary outcome of the study—the total time required to complete the training session—did not differ significantly between conditions (p = 0.270). Therefore, the emphasis placed on individual analyses and the proportion of participants who improved may lead readers to infer the existence of an effect that was not confirmed by the primary analysis. These findings should be treated as exploratory analyses and should not serve as the primary basis for conclusions regarding the efficacy of caffeine expectancy during interval training. Furthermore, the authors should discuss the possibility that some of these individual responses simply reflect biological variability and measurement error, particularly given the lack of information regarding the test–retest reliability of the protocol employed.
In the Results section, lengthy text blocks describe findings that are already presented in tables and figures, creating redundancy and reducing readability. I recommend limiting the narrative to the most important findings and positioning each table or figure as close as possible to the corresponding text.
Figures 1 and 2 present partially overlapping information related to performance. The authors should evaluate whether both figures are necessary in their current form or whether some information could be moved to supplementary material to improve conciseness. Similarly, Tables 1, 2, and 3 are extensive and contain a large number of individual comparisons, many of which appear to have limited statistical or practical relevance.
The overall narrative appears to focus on the few statistically significant findings identified in secondary analyses, which may create a disproportionate impression regarding the magnitude and importance of the observed effects.
Although the authors acknowledge that no statistically significant differences were found for the primary outcome of the study (total session time), much of the discussion is devoted to justifying a potential placebo effect based on secondary analyses, isolated interval differences, or the proportion of participants classified as responders.
The discussion could also more thoroughly explore the reasons why the findings differ from those reported in previous studies on caffeine expectancy. While some earlier investigations demonstrated performance improvements in time-trial protocols and specific performance tests, the present study examined an interval training session characterized by distinct physiological and perceptual demands.
I recommend a more comprehensive discussion of the study limitations, including the small sample size, the absence of a reported sample size calculation, the inclusion of only male participants, the lack of assessment of manipulation effectiveness (i.e., whether participants truly believed they had ingested caffeine), the potential influence of habitual caffeine consumption, and the absence of adjustments for multiple statistical comparisons.
Finally, although participants reported greater feelings of energy and increased urine production under the placebo condition, these findings should be interpreted within the exploratory context of the study and should not be considered evidence of a consistent effect of caffeine expectancy on sports performance.
Author Response
The point-by-point responses to the reviewers' comments are provided. Thank you for your valuable feedback and for helping us improve our manuscript.
Reviewer 1
The topic is relevant and relatively original, as most of the literature on caffeine-related placebo effects has been conducted using time-trial or performance-testing protocols, whereas the present study investigates an interval training session. Below, I provide several comments intended to help the authors improve the quality and scientific rigor of the manuscript.
We would like to sincerely thank the reviewer for their positive assessment of our manuscript. We fully agree that most of the existing literature on caffeine-related placebo effects has relied on time-trial or performance-testing protocols, and we believe that examining this phenomenon within the context of an interval training session more representative of athletes' regular training practice adds valuable and novel insight to this field of research. We have carefully considered each of the reviewer's comments below and have revised the manuscript accordingly, as detailed in our point-by-point responses.
The title is broad and suggests a general investigation of caffeine expectancy during interval training. However, the main finding of the study was the absence of significant differences in overall performance, physiological responses, and kinematic variables between conditions. I recommend that the authors adopt a more specific title that better reflects the actual findings. One possible alternative is: "Caffeine expectancy does not significantly improve interval running performance, physiological responses, or running kinematics in trained runners: an exploratory placebo study."
We thank the reviewer for this valuable suggestion. Following the reviewer's recommendation, we have revised the title to better reflect our findings: "Caffeine expectancy does not affect interval running performance, physiological responses, or running kinematics in trained runners: an exploratory placebo study."
The introduction addresses relevant concepts related to the placebo effect and caffeine expectancy; however, the overall structure lacks logical progression. In several instances, the text reads as a sequence of previously published studies rather than a coherent narrative leading the reader toward the specific knowledge gap that the present study aims to address. Furthermore, I believe that the rationale supporting the need for this study remains insufficient. The authors state that "given the scarce scientific literature on this topic, further investigations are necessary to consolidate these outcomes", but this claim is not adequately supported by the literature presented. Although the cited studies primarily investigated time-trial protocols and specific performance tests, the literature on caffeine expectancy and placebo effects in sport is broader than suggested in the introduction, encompassing different sports, experimental protocols, competitive settings, and outcome measures. Consequently, it remains unclear whether the primary gap relates specifically to interval training, the exercise model employed, the athlete population investigated, or the outcomes assessed. I recommend that the authors define more explicitly the unresolved knowledge gap and demonstrate, based on the available literature, why investigating the placebo effect of caffeine expectancy during interval training sessions represents a meaningful advancement beyond previous studies.
We thank the reviewer for this observation. We have included a better explanation about the lack of studies about the expectancy of caffeine ingestion on Interval training as follow: “Although the available evidence supports the existence of a placebo effect derived from caffeine expectancy in endurance performance, it has been obtained exclusively from studies using simulated competitions or single-effort time-trial protocols. To date, this effect has not been examined during regular interval training sessions, which constitute a core component of endurance training programs. If caffeine-induced expectancy improves endurance performance (as measured by time trials or time-to-exhaustion tests), the absolute speed achieved during an interval training session may also increase. Time sustained at a high fraction of maximal oxygen consumption (VO2max; e.g., ≥90%) has emerged as a key metric for evaluating the effectiveness of interval training protocols (Midgley et al, 2006). Since daily training largely determines athletes' long-term performance progression, it remains unknown whether caffeine expectancy could similarly enhance performance under these ecologically relevant conditions and thereby contribute to overall competitive performance. Therefore, the present study aimed to analyze the potential placebo effect of the belief of ingesting a moderate dose of caffeine in an interval training running session in trained runners.
The experimental design is presented too briefly, making it difficult to understand both the sequence of procedures and the purpose of each variable collected. I recommend expanding the “Experimental Design” subsection to provide a clearer description of the study flow, experimental conditions, procedural order, and, most importantly, the rationale for including each outcome measure (performance, heart rate, kinematic variables, perceived exertion, and adverse effects). In addition, ethical aspects, such as ethics committee approval and informed consent procedures, should be presented in this initial methods section rather than appearing later within the participant description.
We thank the reviewer for this observation. Following this recommendation, we have expanded the "Experimental Design" subsection to more clearly describe the sequence of procedures, the order of the experimental conditions, and, in particular, the rationale for each outcome measure collected (performance, heart rate, kinematic variables, perceived exertion, and adverse effects), explaining the specific purpose of each within the study design. We have also moved the information regarding ethics committee approval and the informed consent procedure to this initial section, which previously appeared later, within the description of participants.
The inclusion and exclusion criteria are relatively generic and do not adequately characterize the eligible participant profile. Considering that the study aims to investigate the placebo effect associated with caffeine expectancy, it would be important to report potentially relevant characteristics such as habitual caffeine consumption, previous use of ergogenic supplements, competitive experience, and other factors that may influence susceptibility to placebo responses.
We thank the reviewer for this observation; we have expanded the inclusion criteria in the Methods section.
Moreover, given the small sample size (n = 14), it is essential to present a sample size calculation justifying the number of participants recruited, including the expected effect size, significance level, and statistical power considered.
We thank the reviewer for this observation. An a priori power analysis (G*Power v3.1.9.7) indicated that nine participants were required to detect a placebo effect of caffeine (effect size = 1.15; two-tailed paired t-test; 1 − β = 0.80; α = 0.05), based on the placebo versus control comparison reported by Hurst et al. (11). This calculation referred to the primary outcome (total time to complete the 5 × 1000-m session). Because the present study also included repeated physiological, perceptual, and kinematic measurements across five intervals per condition, five additional participants were recruited to increase the robustness of the repeated-measures analyses.
The “Procedure” subsection could also be improved through the inclusion of a schematic representation of the experimental design. A flowchart detailing participant visits, the familiarization session, condition randomization, washout periods, placebo administration, assessments performed, and outcomes collected would facilitate understanding of the protocol and improve methodological transparency.
We agree that a schematic representation improves the clarity of the study design. Accordingly, we have included a flowchart of the experimental protocol as Figure 1 in the revised manuscript.
I also suggest that the authors clarify several additional methodological aspects that may influence the interpretation of the findings: (i) the method used to generate the randomization sequence and ensure counterbalancing of conditions; (ii) whether the assessors responsible for timing performance were blinded to the experimental condition; (iii) whether the effectiveness of the experimental manipulation was assessed, that is, whether participants actually believed they had ingested caffeine; and (iv) a more detailed description of environmental testing conditions, as variables such as temperature, humidity, and wind may influence performance during outdoor running sessions.
Participants were randomly assigned to one of the two experimental sequences (placebo–control or control–placebo) using the random number generator available in SPSS.The study was single-blind (participants were blinded to the experimental condition and were informed that they had ingested caffeine, although a placebo was administered), while the researchers were not blinded to the condition allocation. As described in the Methods section, the times for each 1000-meter interval and each200-meter split were recorded manually by two experienced researchers, and the mean of both measurements was used for analysis. This approach was implemented to minimize human error and improve reliability. Additionally, given that the measured distance was not a short sprint, the sensitivity required for time recording is lower than in sprint tests. Manual timing is also recognized as an official and acceptable method by World Athletics (we have added a reference to the World Athletics Technical Regulations in the revised version of the manuscript). Regarding environmental conditions, we confirm that variables such as temperature, humidity, and wind were not formally measured or recorded; however, all sessions were conducted avoiding rainy conditions and under visually similar weather conditions across sessions. Given that this lack of formal environmental control constitutes a methodological limitation, we have included this aspect in the Limitations section of the manuscript, as already detailed in our response to a similar comment raised by Reviewer 2.
Regarding the statistical analyses, the authors should justify the analytical approach adopted for each outcome and provide evidence that the assumptions underlying the statistical models were adequately verified. Although the use of the Shapiro–Wilk test is reported, no information is provided regarding homogeneity of variances, sphericity in repeated-measures analyses, or any corrections applied when assumptions were violated. Multiple analyses were conducted across intervals, splits within intervals, physiological variables, kinematic variables, and adverse effects. However, no strategy to control Type I error resulting from multiple comparisons was described. Consequently, some statistically significant findings may simply reflect chance findings. I recommend that the authors report confidence intervals for both the differences between conditions and the effect sizes, allowing a more robust interpretation of the magnitude and precision of the observed effects. Given the crossover design and the small sample size, the authors should also consider the use of linear mixed-effects models either as a replacement for, or as a complement to, the repeated-measures ANOVAs and multiple paired comparisons currently employed.
We sincerely thank the reviewer for this thoughtful and constructive comment. We have substantially revised the statistical analysis following these recommendations.
First, the statistical analysis section has been rewritten to provide a clearer justification for the analytical approach used for each outcome. Overall performance (total time 5x1000) was compared using a paired t-test, whereas repeated measures outcomes (performance at each interval, pacing, heart rate, rating of perceived exertion, and running kinematics) were analyzed using linear mixed-effects models (LMMs). Side-effect ratings were analyzed using the Wilcoxon signed-rank test because these variables were measured on an ordinal scale and showed a high proportion of tied observations.
Second, the repeated-measures ANOVAs originally used in the manuscript have been replaced by LMMs, which are more appropriate for the crossover repeated-measures design of the present study. These models account for the correlation between repeated observations while providing greater flexibility than traditional repeated-measures ANOVA and avoiding the assumption of sphericity.
Third, the assumptions of the LMMs were evaluated by visual inspection of residual histograms, normal Q-Q plots, detrended Q-Q plots, and residual-versus-fitted value plots. The residuals showed an approximately normal distribution and no evidence of systematic heteroscedasticity, supporting the adequacy of the fitted models.
Finally, multiple pairwise comparisons were performed using Bonferroni-adjusted estimated marginal means, thereby controlling the Type I error rate. This revised analytical approach resulted in more conservative inferences than the original analyses. Notably, the previously reported significant differences at specific intervals and 200-m splits were no longer observed after applying the LMM framework and Bonferroni adjustment.
The manuscript has been revised accordingly in both the Methods and Results sections.
I also recommend greater caution in the interpretation of the individual results used to support the existence of a placebo effect. Although the authors emphasize that 71.4% of participants improved performance under the placebo condition and that the mean reduction in time exceeded the Smallest Worthwhile Change (SWC) threshold, the primary outcome of the study—the total time required to complete the training session—did not differ significantly between conditions (p = 0.270). Therefore, the emphasis placed on individual analyses and the proportion of participants who improved may lead readers to infer the existence of an effect that was not confirmed by the primary analysis. These findings should be treated as exploratory analyses and should not serve as the primary basis for conclusions regarding the efficacy of caffeine expectancy during interval training. Furthermore, the authors should discuss the possibility that some of these individual responses simply reflect biological variability and measurement error, particularly given the lack of information regarding the test–retest reliability of the protocol employed.
We thank the reviewer for this observation, with which we fully agree. We note that a similar comment was raised by Reviewer 2, who likewise expressed caution regarding the responder analysis based on the percentage of participants who improved their performance and on the smallest worthwhile change (SWC) threshold. In response to both observations, we have thoroughly revised the Discussion section: we have removed the SWC threshold as a criterion for classifying responders/non-responders, tempered the interpretation of the percentage of participants who improved under the placebo condition, and added an explicit acknowledgment that these individual differences may reflect normal biological variability between sessions and measurement error, rather than constituting firm evidence of a genuine expectancy effect, particularly given the absence of data on the test–retest reliability of the protocol employed. We have also reinforced in the manuscript that the primary outcome of the study—total session time, which did not differ significantly between conditions (p = 0.270)—should be regarded as the central finding of the study, while the individual-level analyses are now explicitly presented as exploratory and not as the primary basis for conclusions regarding the efficacy of caffeine expectancy. We thank both reviewers for converging on this point, which has strengthened the robustness and interpretive rigor of the revised manuscript.
In the Results section, lengthy text blocks describe findings that are already presented in tables and figures, creating redundancy and reducing readability. I recommend limiting the narrative to the most important findings and positioning each table or figure as close as possible to the corresponding text. Figures 1 and 2 present partially overlapping information related to performance. The authors should evaluate whether both figures are necessary in their current form or whether some information could be moved to supplementary material to improve conciseness. Similarly, Tables 1, 2, and 3 are extensive and contain a large number of individual comparisons, many of which appear to have limited statistical or practical relevance.
Thanks for this suggestion. We have substantially revised this section to improve clarity and conciseness by focusing the narrative on the main findings while avoiding repetition of information already presented in tables and figures. Furthermore, the Results section has been rewritten to reflect the new linear mixed-effects model analyses. The narrative now focuses on the main fixed effects and interactions identified by the LMMs, with post hoc comparisons reported only when appropriate, resulting in a more robust and streamlined presentation of the findings. In addition, the original Figure 1 has been removed, and the information it contained has been incorporated into the new Table 1. The previous Tables 2 and 3 have also been merged into this new Table 1, which we believe provides a clearer and more concise summary of the main outcomes of the study while avoiding unnecessary individual comparisons. We have also revised the placement of tables and figures so that they appear as close as possible to the corresponding text.
The overall narrative appears to focus on the few statistically significant findings identified in secondary analyses, which may create a disproportionate impression regarding the magnitude and importance of the observed effects.
We thank the reviewer for this observation, with which we fully agree. This comment is consistent with similar concerns raised by both this reviewer and Reviewer 2 regarding the emphasis placed on secondary findings in the original manuscript. As detailed in our previous responses, we have reanalyzed the repeated-measures outcomes using linear mixed-effects models, which yielded more conservative results, with the previously reported pacing differences no longer remaining significant. Accordingly, we have revised both the Discussion and the Conclusion to emphasize the consistent lack of significant effects of caffeine expectancy on performance, pacing, physiological responses, and running biomechanics, while presenting the remaining findings as exploratory. We believe these changes provide a more balanced and robust interpretation of the study results, and we thank the reviewer for helping us improve the manuscript.
Although the authors acknowledge that no statistically significant differences were found for the primary outcome of the study (total session time), much of the discussion is devoted to justifying a potential placebo effect based on secondary analyses, isolated interval differences, or the proportion of participants classified as responders.
We have substantially revised the Discussion to focus primarily on the absence of significant effects on the primary outcome.
The discussion could also more thoroughly explore the reasons why the findings differ from those reported in previous studies on caffeine expectancy. While some earlier investigations demonstrated performance improvements in time-trial protocols and specific performance tests, the present study examined an interval training session characterized by distinct physiological and perceptual demands.
We thank the reviewer for this observation. We agree that it is necessary to further explore the possible reasons underlying the differences between our findings and those of previous studies on caffeine expectancy, considering that the latter focused on time-trial protocols and specific performance tests, whereas our study employed an interval training session characterized by clearly distinct physiological and perceptual demands. As part of the overall restructuring of the Discussion section, which we are undertaking in response to the comments raised by both reviewers, we will take this observation into account and incorporate a more detailed reflection on these differences between protocols, in order to more fully contextualize our results in relation to the previous literature.
I recommend a more comprehensive discussion of the study limitations, including the small sample size, the absence of a reported sample size calculation, the inclusion of only male participants, the lack of assessment of manipulation effectiveness (i.e., whether participants truly believed they had ingested caffeine), the potential influence of habitual caffeine consumption, and the absence of adjustments for multiple statistical comparisons.
We thank the reviewer for this observation. In line with the suggestions raised by Reviewer 2, we have expanded the Limitations section to more comprehensively address the points raised: the absence of female participants, the choice of the control condition, the lack of a formal manipulation check (to confirm whether participants genuinely believed they had ingested caffeine), the potential influence of habitual caffeine consumption, and the lack of strict environmental control over weather conditions between sessions. Regarding the sample size, an a priori power analysis was performed and has now been more clearly described in the Methods section (see our response to the corresponding comment above). Furthermore, the concern regarding multiple comparisons has been addressed by reanalyzing the repeated-measures outcomes using linear mixed-effects models with Bonferroni-adjusted post hoc comparisons, thereby providing a more robust control of Type I error. Finally, as specified in the inclusion criteria, all participants were low-to-mild habitual caffeine consumers. Therefore, all participants were familiar with the effects of caffeine, making them an appropriate population in which to investigate expectancy-related responses.
Finally, although participants reported greater feelings of energy and increased urine production under the placebo condition, these findings should be interpreted within the exploratory context of the study and should not be considered evidence of a consistent effect of caffeine expectancy on sports performance.
We thank the reviewer for this observation. We agree that the original wording was too categorical in describing placebo ingestion as a "safe and effective strategy," when in fact our findings regarding side effects (greater feelings of energy and increased urine production) stem from an exploratory analysis and a small sample. We have revised this paragraph to temper this interpretation, making explicit that these findings should be understood within the exploratory nature of the study and not as evidence of a consistent or generalizable effect of caffeine expectancy, either on side effects or on sports performance.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript addresses an interesting question and is generally easy to follow. The topic is relevant because the potential contribution of expectancy effects to the ergogenic response commonly attributed to caffeine remains incompletely understood, particularly in training settings rather than laboratory-based performance tests. I also appreciated that the study was conducted with trained runners and used a workout that resembled a session athletes might realistically perform during regular training.
My main concern relates to the sample size. Fourteen participants seem rather limited for the type of conclusions being drawn and I wonder whether the study was adequately powered to detect meaningful effects. The absence of statistically significant differences across most variables may indicate a true lack of effect, but it may also reflect insufficient statistical power. For this reason, I would be cautious when interpreting both the negative findings and the apparent individual responses.
I was also not entirely persuaded by the responder analysis. The observation that some participants appeared to improve under the placebo condition is interesting, but with such a small sample, it is difficult to know how much of this reflects a genuine expectancy effect and how much may simply represent normal variation between testing sessions. Day-to-day fluctuations in performance are common in endurance athletes, and I do not think the current data allow firm conclusions regarding the existence of distinct responder and non-responder groups.
Another issue concerns the large number of split-by-split comparisons. While I understand the rationale for examining pacing behaviour in detail, performing numerous statistical tests increases the likelihood of false-positive findings. The manuscript would benefit from a brief discussion of this issue and clarification regarding whether any adjustment for multiple comparisons was considered. Some of the isolated significant findings should probably be interpreted with caution.
I also found myself wondering about the choice of the control condition. Comparing a placebo capsule with a no-ingestion condition certainly addresses an interesting practical question, but it makes it difficult to disentangle expectancy effects from the simple act of taking a capsule. A brief discussion of this limitation would strengthen the manuscript.
The physiological and perceptual findings are also worth discussing in a more balanced manner. Heart rate, perceived exertion and the measured kinematic variables were largely unchanged between conditions. These results are important because they suggest that any placebo-related effect, if present, was small. At times, however, the discussion appears to place greater emphasis on the favourable individual responses than on the overall pattern of findings.
Finally, I think the conclusions could be toned down somewhat. Most primary outcomes did not differ between conditions, and therefore, the study provides limited evidence that caffeine expectancy meaningfully improves interval-training performance in this population. The individual responses are interesting and may justify further investigation, but they should probably be treated as preliminary observations rather than firm evidence of a placebo effect. Overall, I believe the manuscript addresses a worthwhile question, but a more cautious interpretation of the results would improve its scientific rigour.
Relevant missing references
The following papers are relevant to this manuscript:
https://pubmed.ncbi.nlm.nih.gov/37871019/
Introduction when discussing caffeine-related ergogenic and perceptual responses.
https://pubmed.ncbi.nlm.nih.gov/37612360/
Introduction or Discussion when describing known performance effects of caffeine and inter-individual variability in response.
Author Response
The point-by-point responses to the reviewers' comments are provided. Thank you for your valuable feedback and for helping us improve our manuscript.
Reviewer 2
The manuscript addresses an interesting question and is generally easy to follow. The topic is relevant because the potential contribution of expectancy effects to the ergogenic response commonly attributed to caffeine remains incompletely understood, particularly in training settings rather than laboratory-based performance tests. I also appreciated that the study was conducted with trained runners and used a workout that resembled a session athletes might realistically perform during regular training.
We would like to sincerely thank the reviewer for their positive and encouraging comments regarding our manuscript. We have carefully considered each of the reviewer's comments below and have revised the manuscript accordingly, as detailed in our point-by-point responses.
My main concern relates to the sample size. Fourteen participants seem rather limited for the type of conclusions being drawn and I wonder whether the study was adequately powered to detect meaningful effects. The absence of statistically significant differences across most variables may indicate a true lack of effect, but it may also reflect insufficient statistical power. For this reason, I would be cautious when interpreting both the negative findings and the apparent individual responses.
We thank the reviewer for raising this important point regarding sample size and statistical power. We would like to clarify that an a priori power analysis was indeed conducted prior to data collection to ensure the study was adequatelypowered. Using G*Power (v3.1.9.7), we determined that a minimum of nine participants was required to detect a placebo effect of caffeine (effect size = 1.15; two-tailed paired t-test; 1 − β = 0.80; α = 0.05), based on the effect sizereported for the placebo vs. control comparison in Hurst et al. (10).
We recruited fourteen participants, exceeding this minimum estimate, in order to (i) provide a safety margin against potential dropouts or missing data across the multiple testing sessions required by the study protocol, and (ii) increase the robustness and precision of the effect size estimates for the secondary and exploratory variables, which were expected to show smaller effects than the primary placebo-effect outcome used in the power calculation. This conservative approach is consistent with common practice in exercise/nutrition intervention studies, where a priori power calculations based on a single reference effect size are typically supplemented by additional participants to safeguard statistical power against attrition and to strengthen the reliability of secondary outcome analyses.
We have amended the manuscript accordingly, and this information has now been incorporated into the Methods section to improve transparency regarding the sample size justification. The revised text now reads as follows:
“An a priori power analysis (G*Power v3.1.9.7) indicated that nine participants were required to detect a placebo effect of caffeine (effect size = 1.15; two-tailed paired t-test; 1 − β = 0.80; α = 0.05), based on the placebo versus control comparison reported by Hurst et al. (11). This calculation referred to the primary outcome (total time to complete the 5 × 1000-m session). Because the present study also included repeated physiological, perceptual, and kinematic measurements across five intervals per condition, five additional participants were recruited to increase the robustness of the repeated-measures analyses.”
I was also not entirely persuaded by the responder analysis. The observation that some participants appeared to improve under the placebo condition is interesting, but with such a small sample, it is difficult to know how much of this reflects a genuine expectancy effect and how much may simply represent normal variation between testing sessions. Day-to-day fluctuations in performance are common in endurance athletes, and I do not think the current data allow firm conclusions regarding the existence of distinct responder and non-responder groups.
We thank the reviewer for this valuable observation and agree that the responder analysis should be interpreted with caution. We acknowledge that, given the limited sample size, it is not possible to reliably distinguish genuine expectancy-driven responses from the normal day-to-day variability in performance commonly observed in endurance athletes. Accordingly, we have substantially revised the corresponding section of the Discussion. Rather than interpreting these findings as evidence of distinct responder and non-responder profiles, we now present them as descriptive observations and explicitly acknowledge that the small individual differences between conditions may simply reflect normal inter-session variability rather than a true placebo effect. We believe this revised interpretation provides a more balanced and appropriately cautious discussion of the individual responses observed in the present study.
Another issue concerns the large number of split-by-split comparisons. While I understand the rationale for examining pacing behaviour in detail, performing numerous statistical tests increases the likelihood of false-positive findings. The manuscript would benefit from a brief discussion of this issue and clarification regarding whether any adjustment for multiple comparisons was considered. Some of the isolated significant findings should probably be interpreted with caution.
We thank the reviewer for this important observation. We agree that the large number of split-by-split comparisons in the original analysis increased the risk of Type I error. As also suggested by Reviewer 1, we have re-analyzed all repeated-measures outcomes using linear mixed-effects models and performed Bonferroni-adjusted post hoc pairwise comparisons. As a result, the previously reported isolated significant differences between conditions at specific intervals and 200-m splits were no longer statistically significant. The Results and Discussion sections have been revised accordingly, and the conclusions are now based on the overall LMM results rather than on individual pairwise comparisons.
I also found myself wondering about the choice of the control condition. Comparing a placebo capsule with a no-ingestion condition certainly addresses an interesting practical question, but it makes it difficult to disentangle expectancy effects from the simple act of taking a capsule. A brief discussion of this limitation would strengthen the manuscript.
We thank the reviewer for this pertinent observation. We agree that comparing the placebo condition (an inert capsule presented as caffeine) with a no-ingestion condition, rather than with an inert capsule administered without any expectancy manipulation (i.e., a "double-dummy" or blind placebo design), does not allow us to fully disentangle the specific effect of caffeine-related expectancy from the possible effect of the mere act of ingesting a capsule (e.g., a general ritual or placebo effect unrelated specifically to caffeine). We deliberately chose this design to address a practically relevant and ecologically valid question: whether informing athletes that they are receiving caffeine (regardless of the actual substance) produces a measurable ergogenic effect compared to a "no intervention" scenario, which more closely reflects the real-world decision-making of athletes considering the use of supplements. Nevertheless, we acknowledge that this design limits our ability to isolate the specific contribution of expectancy from other non-specific effects associated with the act of ingesting a capsule itself. We have included this limitation in the Discussion section, as detailed below.
The physiological and perceptual findings are also worth discussing in a more balanced manner. Heart rate, perceived exertion and the measured kinematic variables were largely unchanged between conditions. These results are important because they suggest that any placebo-related effect, if present, was small. At times, however, the discussion appears to place greater emphasis on the favourable individual responses than on the overall pattern of findings.
We have substantially revised the Discussion to better reflect the overall pattern of results. The revised version now emphasizes the absence of significant effects of caffeine expectancy on physiological responses, perceived exertion, running biomechanics, and performance, while the individual responses are presented only as exploratory observations that should be interpreted with caution. We believe this provides a more balanced and accurate interpretation of our findings.
Finally, I think the conclusions could be toned down somewhat. Most primary outcomes did not differ between conditions, and therefore, the study provides limited evidence that caffeine expectancy meaningfully improves interval-training performance in this population. The individual responses are interesting and may justify further investigation, but they should probably be treated as preliminary observations rather than firm evidence of a placebo effect. Overall, I believe the manuscript addresses a worthwhile question, but a more cautious interpretation of the results would improve its scientific rigour.
We have revised the Conclusions section to provide a more cautious interpretation of the findings, ensuring that it accurately reflects the overall results of the study
Relevant missing references
The following papers are relevant to this manuscript:
 
https://pubmed.ncbi.nlm.nih.gov/37871019/
Introduction when discussing caffeine-related ergogenic and perceptual responses.
 
https://pubmed.ncbi.nlm.nih.gov/37612360/
Introduction or Discussion when describing known performance effects of caffeine and inter-individual variability in response
We thank the reviewer for this helpful suggestion. We have considered both references and have incorporated one of them into the Introduction, where the ergogenic effects of caffeine on sports performance are discussed.
Round 2
Reviewer 1 Report
Comments and Suggestions for Authors2nd Review of Manuscript nutrients-4402065
I would like to congratulate the authors on the effort devoted to revising the manuscript. The revised version shows substantial improvements compared with the original submission. Nevertheless, despite these important improvements, several issues still need to be addressed before the manuscript can be considered for publication.
- Lack of verification of the effectiveness of the experimental manipulation (manipulation check).
The main objective of this study is to investigate the placebo effect resulting from the expectancy of caffeine ingestion. However, the manuscript does not report whether participants actually believed they had ingested caffeine. The absence of this information considerably limits the interpretation of the findings, since a placebo effect necessarily depends on the successful induction of expectancy. If participants did not believe the information they were given, the absence of differences between conditions may simply reflect an unsuccessful experimental manipulation rather than the absence of a placebo effect. I acknowledge that this limitation can no longer be corrected experimentally; however, it should be discussed more explicitly in the Discussion section, highlighting its impact on the interpretation of the findings and the conclusions of the study. - The adoption of linear mixed-effects models represents an important methodological improvement to the manuscript. However, the presentation of the statistical results remains largely based on F and p values. I recommend that the authors report, for the main models: (a) estimated marginal means; (b) 95% confidence intervals; and (c) appropriate measures of effect size for the models employed, whenever possible.
- The Results section reports a significant condition × interval interaction for performance during the 1000-m intervals. However, none of the post hoc comparisons remained statistically significant after Bonferroni correction. This appears contradictory, as a statistically significant interaction is reported without significant pairwise differences. The authors should clarify in the Discussion that the model detected an overall interaction pattern, but that this pattern did not translate into sufficiently robust differences at specific intervals after adjustment for multiple comparisons.
- Although the authors state that all tests were conducted under similar environmental conditions, temperature, humidity, and wind speed were not objectively recorded. Because the protocol was performed outdoors, this limitation deserves greater emphasis in the Discussion, as these environmental factors may influence running performance.
- The overall quality of the writing has improved considerably. However, the manuscript still contains several minor grammatical and stylistic issues throughout the text. I recommend a final revision by a professional English language editor or a native English speaker with experience in scientific writing before publication.
Author Response
I would like to congratulate the authors on the effort devoted to revising the manuscript. The revised version shows substantial improvements compared with the original submission. Nevertheless, despite these important improvements, several issues still need to be addressed before the manuscript can be considered for publication.
RESPONSE: We sincerely thank the reviewer for the careful evaluation of our revised manuscript and for recognizing the improvements made. We appreciate the additional comments and have carefully addressed each of the remaining issues, which we believe have further strengthened the quality and clarity of the manuscript.
Lack of verification of the effectiveness of the experimental manipulation (manipulation check). The main objective of this study is to investigate the placebo effect resulting from the expectancy of caffeine ingestion. However, the manuscript does not report whether participants actually believed they had ingested caffeine. The absence of this information considerably limits the interpretation of the findings, since a placebo effect necessarily depends on the successful induction of expectancy. If participants did not believe the information they were given, the absence of differences between conditions may simply reflect an unsuccessful experimental manipulation rather than the absence of a placebo effect. I acknowledge that this limitation can no longer be corrected experimentally; however, it should be discussed more explicitly in the Discussion section, highlighting its impact on the interpretation of the findings and the conclusions of the study.
RESPONSE: We thank the reviewer for this important observation. We agree that the absence of a formal manipulation check represents a limitation of the study. Although participants' expectancy was not formally assessed using a questionnaire, their comments during and after the experimental sessions suggested that they believed they had ingested caffeine. For example, some participants reported feeling unusually active for the remainder of the day, while others asked about the brand of caffeine supposedly administered and how they could obtain it for future competitions. Nevertheless, we acknowledge that these anecdotal observations cannot replace a formal assessment of expectancy.
It should also be noted that data collection was conducted over an extended period, with participants attending the experimental sessions on different dates. Under these circumstances, formally asking participants whether they believed they had ingested caffeine could have increased the likelihood of information being shared with participants who had not yet completed the study, potentially compromising the expectancy manipulation in subsequent participants.
Accordingly, we have expanded the Discussion to explicitly acknowledge this limitation and to emphasize that the absence of a formal manipulation check should be considered when interpreting the study findings.
The adoption of linear mixed-effects models represents an important methodological improvement to the manuscript. However, the presentation of the statistical results remains largely based on F and p values. I recommend that the authors report, for the main models: (a) estimated marginal means; (b) 95% confidence intervals; and (c) appropriate measures of effect size for the models employed, whenever possible.
RESPONSE: We thank the reviewer for this valuable suggestion. Following this recommendation, we have revised the presentation of the linear mixed-effects model results. For the main effects of condition, we now report the estimated marginal means (EMMs) together with their corresponding 95% confidence intervals in the Results section. In addition, Table 1 has been expanded to include the estimated difference between EMMs (Control − Placebo) and its 95% confidence interval for each interval. We have retained the observed means ± standard deviations in the table to provide a clear descriptive summary of the data, while the additional EMM-based estimates and confidence intervals offer the model-based inference requested. We believe these changes improve both the completeness and interpretability of the statistical results.
The Results section reports a significant condition × interval interaction for performance during the 1000-m intervals. However, none of the post hoc comparisons remained statistically significant after Bonferroni correction. This appears contradictory, as a statistically significant interaction is reported without significant pairwise differences. The authors should clarify in the Discussion that the model detected an overall interaction pattern, but that this pattern did not translate into sufficiently robust differences at specific intervals after adjustment for multiple comparisons.
RESPONSE: Thanks for your suggestion. Now it reads: “Although the linear mixed model revealed a significant condition × interval interaction, together with a main effect of interval, Bonferroni-adjusted pairwise comparisons showed no significant differences between the placebo and control conditions at any of the five 1000-m intervals. This suggests that the significant interaction reflected an overall pattern across the repeated measurements rather than robust differences at any specific interval. Accordingly, caffeine expectancy did not meaningfully modify pacing behavior during the interval training session.”
Although the authors state that all tests were conducted under similar environmental conditions, temperature, humidity, and wind speed were not objectively recorded. Because the protocol was performed outdoors, this limitation deserves greater emphasis in the Discussion, as these environmental factors may influence running performance.
RESPONSE: Thank you for this comment. We have expanded the Discussion to acknowledge this limitation and its potential impact. Now it reads: Finally, although all sessions were conducted under similar weather conditions (i.e., no rain and comparable ambient temperature), scheduling them at comparable times of day and within the same season to minimize environmental variability, temperature and humidity were not formally recorded or strictly standardized, and minor day-to-day fluctuations cannot be entirely ruled out as potential confounding factors. Nevertheless, previous evidence suggests that meteorological factors, particularly ambient temperature and wind, are more likely to influence endurance running performance when conditions differ substantially or fall outside the optimal range for performance (Mantzios et al., 2022). Given the apparently similar conditions across testing sessions, it is unlikely that small variations in weather meaningfully affected the present findings. In addition, the exercise protocol consisted of five 1,000-m intervals separated by 2-min recovery periods, which may have facilitated partial heat dissipation and limited the progressive accumulation of thermal strain, potentially further reducing the influence of modest environmental differences on performance. However, this interpretation should be considered with caution, as the effects of small variations in meteorological conditions during intermittent endurance running protocols have not been directly investigated.
The overall quality of the writing has improved considerably. However, the manuscript still contains several minor grammatical and stylistic issues throughout the text. I recommend a final revision by a professional English language editor or a native English speaker with experience in scientific writing before publication.
RESPONSE: Thank you for your recommendation. We have carefully revised the entire manuscript to further improve the grammar, style, and overall readability. The text has been thoroughly proofread, and minor linguistic issues have been corrected throughout the manuscript.
Reviewer 2 Report
Comments and Suggestions for AuthorsGeneral comments
I have no further concerns about this manuscript. The authors adequately addressed all the issues I raised.
Author Response
We sincerely thank the reviewer for the careful evaluation of our manuscript and for the constructive comments provided throughout the review process. We appreciate your positive assessment of the revised version and are pleased that our revisions have satisfactorily addressed your concerns.
Round 3
Reviewer 1 Report
Comments and Suggestions for Authors3rd Review of Manuscript nutrients-4402065
I would like to congratulate the authors on the revisions made to the manuscript. The manuscript has improved substantially compared with the previous version, and the issues raised during the review process have been addressed satisfactorily. I recommend acceptance of the manuscript in its current form.

