Next Article in Journal
Nephron Flow Lab: Vibe Coding a Visual Nephrology Learning App
Previous Article in Journal
Beyond the Archive: Designing a Digital Showroom for Exploring the Living Heritage of Nakhon Si Thammarat Brocade
Previous Article in Special Issue
A Mixed Reality Tool with Automatic Speech Recognition for 3D CAD Based Visualization and Automatic Dimension Generation in the Industry 5.0 Shipyard
 
 
Article
Peer-Review Record

Synchronized Listening-While-Reading in EFL Vocabulary Learning: Immediate Gains and Exploratory Heart-Rate Patterns

Multimodal Technol. Interact. 2026, 10(9), 92; https://doi.org/10.3390/mti10090092
by Heeseong Ahn 1, Nahkyoung Han 2, Jong-su Park 2, Young Seok Oh 2, Hubert H. Pak 2,* and Chungwan Lim 3,*
Reviewer 1: Anonymous
Reviewer 2:
Multimodal Technol. Interact. 2026, 10(9), 92; https://doi.org/10.3390/mti10090092
Submission received: 30 March 2026 / Revised: 2 September 2026 / Accepted: 3 September 2026 / Published: 7 September 2026

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

This study investigates the impact of multimodal input on vocabulary acquisition among Korean university EFL learners. By challenging the traditional "Cognitive Load Reduction" hypothesis, the authors propose a "Balanced Efficiency" model.

1.While the abstract mentions delayed retention, the results section focuses heavily on immediate post-test data. A more detailed statistical analysis of the delayed post-test scores would strengthen the arguments regarding long-term lexical retention.

2. With a total of 40 participants (20 per group), the sample size is relatively small for a between-subjects design. The authors should briefly acknowledge this as a limitation regarding the generalizability of the findings to broader EFL populations.

3. While average heart rate is a good indicator, including Heart Rate Variability  analysis could provide deeper insights into the autonomic nervous system's regulation of cognitive load and emotional state.

4.The discussion could be expanded to suggest how developers of Computer-Assisted Language Learning systems might "engineer" digital materials to trigger optimal engagement levels based on this model.

5.Some related works may be missing.[1]ADP: Graph Adaptive Pooling based on Edge Understanding with Graph Pooling Information Bottleneck [2] It Takes Two: Multi-frequency Perception with Complementary Fusion Network for Complex Scene Segmentation

Author Response

Please see the attachment.

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

The manuscript addresses a pertinent question at the intersection of multimodal learning, second language acquisition, and psychophysiological measurement: whether synchronized audio-plus-text input enhances vocabulary learning by reducing cognitive load, increasing engagement, or achieving a balance of both. 

The behavioral findings are promising; the text-with-audio group demonstrates substantial pre/post improvement, whereas the text-only group does not. However, the methodological foundation is insufficient to support the central mechanistic claim. In its current form, the study presents an interesting pilot experiment, but the evidence does not yet substantiate the proposed “balanced efficiency” model as an empirically validated explanatory framework.

Primary Methodological Concerns

The most serious issue is construct validity. The paper argues that higher task heart rate in the multimodal condition reflects “engagement” rather than “overload,” yet engagement is never measured directly. The study records heart rate only, without heart-rate variability, subjective workload, self-reported engagement, comprehension monitoring, or behavioral attention measures. Heart rate is a highly nonspecific autonomic signal. It can reflect cognitive effort, emotional arousal, novelty, stress, attentional mobilization, or combinations of these. As a result, the manuscript currently infers a latent mechanism from a single physiological proxy. That is too weak a basis for the proposed biopsychological interpretation. At most, the data support a tentative interpretation of differential task activation, not a validated model of engagement-driven learning.

A second major weakness is the lack of procedural specification. Essential experimental details are either missing or insufficiently described, including the exact duration of the intervention, passage characteristics, the number and distribution of target words, narration speed, synchronization logic, learner control over pacing, and the device or interface configuration in the multimodal condition. The interaction design is central, yet the manuscript provides only a general description of the multimodal system. While Fig. 1 serves as a basic flowchart, it does not address the absence of interface-level and protocol-level detail. Without this information, replication is challenging, and the mechanism underlying the observed effect remains unclear.

A third concern relates to the vocabulary assessment. The manuscript reports that vocabulary knowledge was measured using a 20-item researcher-developed test, scored on a five-point scale, yet also describes the items as multiple-choice with four alternatives. This scoring logic is unclear and requires clarification. Additionally, while the manuscript states that Cronbach’s alpha was examined, the actual reliability coefficient is not provided. There is no evidence of item validation, content alignment, or difficulty calibration. Given that the same construct is assessed immediately before and after a brief intervention, practice effects are a legitimate concern. The manuscript should specify whether the pretest and posttest were identical, parallel, or counterbalanced. As currently written, the validity of the primary learning outcome is not adequately established. 

Statistical Analysis Concerns

The statistical strategy is suboptimal for the study design. The manuscript relies on separate paired t-tests within each condition and then narratively compares patterns across groups, which does not provide the strongest test of the research question. In a two-group pretest/posttest design, the key inferential quantity is the time by condition interaction, which should be tested using a mixed ANOVA, an ANCOVA with the pretest as a covariate, or, preferably, a mixed-effects model. The current approach reduces the rigor of the argument and may lead to misleading emphasis on within-group significance differences.

Additionally, the manuscript indicates that a between-group post hoc comparison was conducted, but the actual inferential results are not clearly reported in the results section. This omission is significant. If the behavioral effect is central, the manuscript should report the between-group post hoc contrast, confidence intervals, and, ideally, the interaction effect. The same recommendation applies to physiological data.

Evaluation of Results and Interpretation

The behavioral results are promising, whereas the physiological results are considerably weaker than the discussion suggests. The heart-rate differences cited in support of the central theoretical reframing are only marginal: mean task heart rate and ΔHR approach significance but do not reach conventional thresholds, and peak heart rate, peak elevation, and recovery time do not differ meaningfully between groups. Thus, while the behavioral evidence is strong relative to the sample, the physiological evidence remains suggestive rather than confirmatory. The manuscript should avoid presenting the “balanced efficiency” account as empirically demonstrated. A more defensible conclusion is that the findings are consistent with an engagement-related interpretation, but do not yet clearly distinguish among alternative explanations.

The proposed model in Fig. 2 is conceptually appealing, but currently serves as a speculative synthesis rather than a tested model. The study did not directly measure engagement, did not distinguish between extraneous and germane load, and did not employ multimethod triangulation. Therefore, Figure 2 should be presented as a provisional conceptual framework derived from pilot data, rather than as a validated explanatory model.

Another notable issue is the mention of delayed retention in the abstract and limitations, without adequate evidence supporting it or supplementation in the developed results. If delayed retention is a substantive contribution, it should be reported transparently. If not, it should be omitted from the abstract or clearly identified as exploratory ...

Author Response

Please see the attachment.

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

Although the quality of the introduction is acceptable, the bibliography is outdated. No single reference from the last 6 years is presented. Unfortunately, the implication is that what was acceptable 6 years ago, after the advances of the application in listening in EFL has changed dramatically. This is not only an issues for this publication but it will be in any other one. As a result, the introduction needs to be totally revised and the advances in these 6 years, of which chatbots use in listening training are not the least, is absolutaely necessary. The starting point is adequate and so is the convenience sample but beyond the  principle of this research based on Cognitive Load Theory (CLT). It is necessary to observe that newer theories in EFL listening research., are currently more predominant since 2019 (when the last references are cited). Being that the case, it is necessary to revise other theories the field which have shifted strongly toward metacognitive approaches, strategy-based instruction, mobile‑assisted listening, and sociocultural scaffolding, all of which appear more prominent in recent empirical studies than CLT, and all of them should, at least, be considered in the literature review if not in the experimental section. The only exception is the reference to 2021 (one total in these six last years) and the sentnce "However, recent research in learning sciences suggests that this explanation may be too simple. Many scholars now distinguish between cognitive load and learner engagement as related but separate aspects of the learning process [11–13]." as an example of how outdated the paper is. Additionally, some parts relate to emotional aspects which are not really addressed in the research.

The research questions and objectives are also unclear and should be shaped to really be connected to the data collection and analysis. Additionally, it is had to see what interventions were produced over the experimental section. They are mentioned but not justified. The question the author/s should answer is what has happened from the pre-test to the post test, under what conditions and especially how long has this process been.

The statistical analysis is potentially adequate but without the previous information and a sound context, it is hard to see its research value. So are the conclusions that are both tentative and limited.

The author/s are recommended to check these references:

Ahmadi Safa, M., & Motaghi, F. (2021). Cognitive vs. metacognitive scaffolding strategies and EFL learners’ listening comprehension development. Language Teaching Research, 1–24.

Jiang, D. (2024). Cognitive Load Theory and Foreign Language Listening Comprehension. Springer.

Qasserrasi, L. (2025). Systematic review of effective teaching listening practices in ESL/EFL settings. European Journal of English Language Teaching, 9(6). 

Tan, S., Abd Samad, A., & Ismail, L. (2024). The influence of meta-cognitive listening strategies on listening performance in the MALL: The mediation effect of learning style and self-efficacy. SAGE Open, 1–16.

Author Response

Please see the attachment.

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

Accept

Author Response

Thank you very much for your positive evaluation and for your constructive comments throughout the review process. We sincerely appreciate your time and careful consideration of our manuscript.

Reviewer 2 Report

Comments and Suggestions for Authors

The authors have thoroughly addressed all eight first-round comments. The two structural issues I raised (inferring a latent mechanism from a single nonspecific proxy and testing a between-group question with within-group tests) are now fully resolved. The behavioural analysis is statistically sound and internally consistent. Figure 2 has been revised from a speculative causal diagram to an accurate inferential map.

What remains is: the heart-rate instrument is identified only as a "validated wearable device," with no model, sensing modality or validation citation; "recovery time" is reported without an operational return-to-baseline criterion; the physiological contrasts lack confidence intervals; and the ethics waiver, justified by voluntary and anonymous participation, does not cover continuous physiological recording linked to individual test scores.

None of these concerns requires new data collection. I therefore recommend minor revision, provided the four points above are addressed.

Author Response

Please see the attached point-by-point response.

Author Response File: Author Response.docx

Back to TopTop