Review Reports
- Giulia Barresi 1,*,
- Karine Maria Porpino Viana 2 and
- Daniela Bulgarelli 3
- et al.
Reviewer 1: Anonymous Reviewer 2: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsOverall, I enjoyed reading the paper. It addresses an important preregistered replication question in peer action coordination, and the null finding is potentially valuable—especially given the strong emphasis on emotion understanding in earlier work.
The design is clear and the write-up is generally readable. However, it seems that there is several technical/statistical issues and a few reporting inconsistencies. Several comments is noted below which may strengthen the manuscript considerably.
- Unit of analysis / non-independence of dyadic data
In the cooperative scenario, the result reflects the combined performance of both children (they both contribute to the shared board result). If each child is analyzed as a separate data point in regressions and t-tests, the calculated standard errors and p-values might be skewed. I suggest clarification - whether the analyses were conducted at the dyad level (N≈54) or at the individual level (N=108). If analyses were performed at the individual level, it might be beneficial to redo them using either (a) dyads as the unit of analysis, or (b) a mixed model that includes a random intercept for dyads (and potentially for school/class as well).
- The matching procedure might diminish the effects being assessed
pairs were matched based on their individual performance in the labyrinth and their understanding of emotions. However, understanding emotions is one of the key predictors (H1). Matching on a predictor can limit its variance and lower the likelihood of identifying a relationship with cooperative performance. At least, this should be explicitly discussed as a potential reason for the null result. Ideally, it would be beneficial to report the remaining variance in emotion understanding after matching (for example, distribution by dyad) and to consider conducting sensitivity analyses (models that exclude matching on emotion understanding, if possible, or approaches comparing dyad-mean versus dyad-difference).
- The scoring system for the Labyrinth Ball Game clarification
The "average score achieved for each level" is aggregated across levels and then divided by 39 if yes then it would be beneficial to clarify:
-what exactly constitutes the “score” for each trial or level (distance traveled, number of holes navigated, or time taken?) and how is “progress” defined.
How are multiple trials for each level combined (is it the average of up to 5 trials or the best trial?).
How are incomplete levels treated, and is there any ceiling effects in the easier levels?
If scoring relied on video analysis, please include any reliability checks (even if it’s just a small sample that was double-coded).
- ANT inhibitory control index: computation and quality control
It would be helpful to clarify whether RTs were computed on correct trials only (typical practice) and whether participants with very low accuracy were excluded.
Since your inhibitory index is an RT difference score (incongruent − congruent), and some children showed negative values (range includes −28 ms) – what is the interpretation of those values (noise? strategy?).
Think about including descriptive statistics for reporting accuracy, as well as examining speed–accuracy tradeoffs (e.g., a straightforward correlation between inhibitory RT-cost and accuracy, or conducting an “inverse efficiency” sensitivity analysis).
- Power analysis wording
The “post hoc sensitivity power analysis” is not very informative for inference and may be misleading (it depends on the observed/sample assumptions). Since the study is preregistered, perhaps it would be better to emphasize rationale for sample size planning, or at least discuss detectable effects (possibility that true effects are smaller than “medium”)
- Age × gender interaction
The age×gender effect is interesting but whether this was really preregistered(?). If not, it should be labeled as exploratory.
Also, it would be beneficial to report the interaction coefficient, SE, CI, and simple slopes (age effect within boys and within girls), not only AIC/ΔR² and LOESS curves.
Moreover, if dyads are same-gender “gender effects” are effectively dyad-gender differences; this should be interpreted explicitly.
- Reporting values – consistency
table 2: ΔR² for M4 appears to be 0.08 not 0.8 (since R² goes from .30 to .38) (typo?)
consistnecy in naming model „M0“ vs „MBaseline“
Table 1 uses “ρ” but the text uses “r”; it should be specified whether Pearson or Spearman was used
Author Response
Please see file in attachment
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThis study delves into the complex interplay between socio-cognitive abilities and real-time social interaction, revealing nuanced associations influenced by age and gender. However, to further elevate the quality and impact of the manuscript, several revisions are recommended.
1.The introduction provides a comprehensive review of relevant literature but could be more focused. Specifically, when discussing prior studies on emotion understanding and inhibitory control, include only the most directly relevant and recent findings to strengthen the rationale for the current study.
2.The hypotheses (H1, H2, H3) are clearly stated, but consider providing more rationale for why inhibitory control might moderate the relationship between emotion understanding and peer action coordination.
3.In the materials and methods section, provide more details on how participants were recruited from the four primary schools. Include information on whether the schools were randomly selected or if there were specific criteria for their inclusion.
4.The sample size criteria require reference to the literature.
5.Relevant statistical symbols should be italicized (e.g., F, r, p).
6.Figure 4 is mentioned but not described in detail in the results section; provide a brief interpretation of what the LOESS curves depict.
7.All figures in this study (use TIFF images with minimum 300 dpi).
8.When discussing the null findings for emotion understanding and inhibitory control, delve deeper into potential reasons for these results. Consider discussing whether the tasks used were sensitive enough to detect individual differences in these constructs.
9.The observed gender differences in peer action coordination are significant and warrant a more detailed discussion. Explore potential sociocultural, biological, and developmental explanations for why boys outperformed girls in this task.
10.Given that the study was conducted in Northern Italy, consider whether similar findings would be observed in other parts of the world with different cultural norms and values.
11.The suggestions for future research are valuable. Expand on how mixed-gender dyads might be studied differently, including potential challenges and benefits of such an approach.
Author Response
Please see the attachment
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for Authors- The reporting of the dyad-level analysis must align with the presented results.
Most impportant recommendation - you mention that the analyses were conducted again using dyads as the unit of analysis; however, many of the reported degrees of freedom/test statistics appear to still align with individual-level sample sizes (e.g., F(3,104); t-tests with dfs around 102–105). Please confirm the analytical dataset and re-export all model results to ensure that N, df, SEs, and p-values are consistently representative of dyad-level analyses (or, if individual-level rows were maintained, use a suitable mixed model with dyad as a random effect).
- Power analysis phrasing
PLease eliminate or further soften any lingering phrases related to “post hoc sensitivity power analysis confirmed…” and maintain the focus on the rationale for the a priori/preregistered sample size, as well as the notion that actual effects might be smaller than what is achievable with this sample.
- Clarification on labyrinth scoring and coding
3.The explanation of the scoring is significantly more understandable. There is one final point: the manuscript mentions coding from the video and settling disputes through discussion. Please include a short clarification regarding the number of coders and (if feasible) a minimal reliability statement (such as a small double-coded subset/agreement) or clearly explain why reliability assessment was not conducted.
Final review and proofreading
4.please conduct a last editorial check to eliminate any duplicate sentences or formatting issues and to guarantee uniformity in symbols and terminology throughout the text and tables (e.g., r vs ρ; naming of models; formatting of tables).
Author Response
Comment 1: Alignment of dyad-level analyses and reported statistics
Response:
Thank you for highlighting this critical issue. We confirmed the analytical dataset and re-ran all analyses using dyads as the unit of analysis (N = 54 dyads). All model results were re-exported to ensure that sample sizes, degrees of freedom, standard errors, and p-values consistently reflect dyad-level analyses. Any remaining individual-level degrees of freedom (e.g., F(3,104), t-tests with df ≈ 100) were corrected throughout the manuscript. The Methods and Results sections now explicitly state that all cooperative performance analyses were conducted at the dyad level.
Comment 2: Power analysis phrasing
Response:
We revised the wording to eliminate references to post hoc sensitivity power analyses. The manuscript now focuses exclusively on the a priori, preregistered rationale for sample size selection, while explicitly acknowledging that true effects may be smaller than those detectable with the present sample. These changes appear in the Methods (Participants/Sample Size) and Discussion (Limitations).
Comment 3: Clarification on Labyrinth Ball Game scoring and coding
Response:
We expanded the Methods section to clarify the scoring procedure and coding process. Specifically, we now report that six trained coders coded all performances from video recordings using standardized guidelines and practiced on pilot data (not included in the final sample). Because experimental videos were not cross-coded, formal inter-rater reliability statistics were not calculated; this is now explicitly stated. Any ambiguities during training were resolved through discussion prior to coding the experimental data.
Comment 4: Final editorial check and consistency
Response:
We conducted a full editorial review to eliminate redundancies and ensure consistency in terminology, symbols, and formatting across the manuscript and tables. This included standardizing statistical notation (e.g., Pearson’s r), harmonizing model labels, correcting typographical and formatting inconsistencies, and ensuring alignment between reported statistics and tables.