1. Introduction
Vocabulary knowledge plays an essential role in second language (L2) reading comprehension and overall language proficiency, particularly in English as a Foreign Language (EFL) contexts where reading functions as a primary source of linguistic input [
1,
2]. Lexical knowledge not only predicts reading comprehension but also supports listening comprehension, academic discourse participation, and long-term language attainment [
3,
4]. From a usage-based perspective, repeated exposure and attention to language input allows connections between word forms and meanings to become stronger over time [
5,
6]. Therefore, instructional methods that influence how learners process language input may have significant effects on vocabulary development, even when the amount of input remains similar.
At the same time, digital learning environments have significantly changed how learners encounter language input. Contemporary reading platforms now combine written texts with additional features such as synchronized audio narration, adjustable pacing, glosses, and adaptive learning tools. Such multimodal systems raise a central theoretical and design question: do they indeed enhance learning primarily by reducing cognitive load, by amplifying learner engagement, or by simultaneously calibrating both processes [
7,
8,
9]? Although multimedia learning theory provides useful principles for combining visual and auditory information, second language acquisition (SLA) research has not yet fully explained how these principles influence vocabulary learning in realistic classroom settings.
A common assumption among educators is that listening while reading makes reading easier for language learners. Audio support may assist learners in recognizing pronunciation patterns, mapping written forms to sounds, and identifying prosodic cues in the text and thereby lower processing strain. According to Cognitive Load Theory, lowering unnecessary processing demands, often referred to as extraneous load, allows learners to allocate more cognitive resources to meaningful learning processes such as comprehension and schema formation [
10]. Within L2 contexts, this may facilitate lexical encoding by reallocating cognitive resources toward semantic integration rather than surface-level decoding.
Recent developments in EFL listening research, however, suggest that successful listening cannot be explained solely in terms of reducing cognitive processing demands. Contemporary work increasingly emphasizes the role of metacognitive regulation, strategic listening, scaffolding, and digitally mediated learning environments. Metacognitive scaffolding, for example, can promote learners’ monitoring, evaluation, and gradual self-regulation of listening processes [
11], while research in mobile-assisted language learning (MALL) similarly indicates that metacognitive listening strategy use is positively associated with listening performance, with learning style and self-efficacy playing mediating roles [
12]. Recent reviews further characterize effective listening as an active process involving strategically coordinated pre-, while-, and post-listening activities and interactions between top-down and bottom-up processing [
13]. At the same time, recent applications of Cognitive Load Theory to foreign-language listening demonstrate that the effectiveness of read-only, listen-only, and read-and-listen formats may vary according to learner expertise [
14]. Beyond these developments, recent advances in conversational AI and chatbot-assisted language learning illustrate a broader shift toward interactive and learner-responsive digital environments [
15,
16].
Taken together, these developments suggest a broader shift from viewing effective listening instruction primarily as a problem of reducing cognitive load toward understanding it as a dynamic process involving the regulation of cognitive resources, strategic attention, and learner engagement within increasingly multimodal digital environments. The present study builds on this broader perspective while focusing specifically on synchronized listening-while-reading and its behavioral and physiological correlates.
Many scholars now distinguish between cognitive load and learner engagement as related but separate aspects of the learning process [
17,
18,
19]. Engagement-based frameworks often conceptualize learning as involving cognitive effort, behavioral participation, and emotional involvement [
17,
18,
19,
20,
21]. Under this perspective, effective instruction may reduce certain processing difficulties while at the same time increasing sustained attention, motivation, and task involvement. In such cases, learning activities might not appear less demanding from a physiological perspective; instead, they may produce higher levels of attentional activation while still supporting learning.
This distinction carries methodological and theoretical significance for SLA research incorporating psychophysiological measures. Physiological indices such as heart rate (HR) and heart rate variability (HRV) are sensitive to integrated autonomic activation reflecting both cognitive effort and affective arousal [
22,
23]. Elevated HR may signal increased processing demand, heightened attentional mobilization, emotional engagement, or a combination thereof. Without theoretical grounding and behavioral triangulation, interpreting HR elevation as unequivocal evidence of “higher load” risks conflating adaptive engagement with detrimental overload. In digital language learning contexts, such misinterpretation could lead to overly simplified instructional designs aimed at minimizing activation rather than optimizing learning.
Recent SLA discussions call for greater integration of process measures with outcome data to strengthen mechanistic explanations [
24,
25]. Technologies such as eye-tracking and pupillometry have been used to examine attention during language processing [
24], but comparatively few studies have examined autonomic indicators such as HR or HRV in vocabulary learning contexts. Even fewer have explicitly modeled the interplay between multimodal input, cognitive load, engagement, and physiological regulation. As a result, the field lacks a coherent biopsychological account of how digital listening–reading environments reshape lexical processing states.
The present study attempts to address this issue by considering listening-while-reading as a mechanism that may balance cognitive efficiency and learner engagement. When written text is combined with synchronized audio, decoding difficulties may decrease, which can improve processing efficiency. At the same time, the presence of audio may maintain learner attention and encourage sustained task involvement, potentially leading to higher levels of physiological activation. Rather than assuming that effective multimodal learning must reduce physiological arousal, this study investigates whether vocabulary learning is associated with reduced strain, increased engagement-related activation, or a balanced combination of both.
Using experimental data comparing text-only and listening-while-reading conditions among Korean university EFL learners, we investigate two core questions: (a) whether multimodal listening-while-reading produces superior immediate vocabulary learning and (b) whether heart rate patterns during task performance support a cognitive-load reduction explanation, an engagement-based explanation, or a balanced interpretation involving both processes. By integrating behavioral outcomes with physiological measures, this study contributes to a theoretically grounded model of cognitive–affective regulation in digital L2 reading environments.
Accordingly, the present study has three main purposes:
To compare vocabulary learning outcomes under text-only and text-with-audio listening-while-reading conditions.
To examine heart-rate dynamics during task performance and recovery as indices of task-related physiological activation.
To evaluate whether multimodal digital reading operates under a reduced-strain profile, an engagement-amplification profile, or a calibrated “balanced efficiency” profile.
By examining listening-while-reading within a biopsychological framework that integrates cognitive load theory, learner engagement, and autonomic regulation, this study seeks to advance a more mechanistic understanding of digital language learning in EFL contexts.
3. Methods
3.1. Participants and Design
Participants were 40 Korean university students enrolled in required English courses at a large metropolitan university. All participants were native speakers of Korean and had received a minimum of six years of prior formal English instruction through secondary education. Their proficiency level corresponded approximately to lower-intermediate to intermediate based on institutional placement criteria. None reported hearing impairments or cardiovascular conditions that would affect physiological measurement.
Participants were randomly assigned to one of two experimental conditions:
The study employed a between-subjects pretest–posttest experimental design with physiological monitoring during the learning phase (see
Figure 1). Random assignment was conducted to minimize systematic group differences. Pretest vocabulary equivalence was statistically verified prior to outcome analyses.
The learning task consisted of a controlled digital reading activity in which target lexical items were embedded in expository text. Both experimental groups received identical written materials and target lexical items under standardized exposure conditions. The principal experimental manipulation was the presence or absence of synchronized auditory narration: the text-with-audio group listened to narration while viewing the corresponding written text, whereas the text-only group viewed the same written material without auditory support. Detailed specifications of the learning materials, auditory presentation, synchronization procedure, learner pacing, and interface configuration are provided in
Section 3.2.
Physiological data were recorded continuously during:
This design enabled simultaneous examination of behavioral learning outcomes and task-related physiological activation patterns.
3.2. Learning Materials and Experimental Procedure
The experimental learning task consisted of a controlled digital reading activity designed for Korean university-level EFL learners at approximately lower-intermediate to intermediate proficiency. Both experimental conditions were presented with identical written learning materials containing the target lexical items. The principal experimental manipulation was the mode of input presentation: participants in the text-only condition received the written material without auditory support, whereas participants in the text-with-audio condition received the same written material accompanied by corresponding auditory narration. The written linguistic content and exposure duration were held constant across conditions to minimize differences unrelated to the modality manipulation.
The learning material consisted of one expository passage of approximately 250–300 words and contained approximately 20 target lexical items corresponding to the vocabulary content assessed in the study. The target items were embedded within the passage rather than presented as an isolated vocabulary list, allowing lexical processing to occur within a meaningful reading context. The same passage and target lexical content were used in both experimental conditions.
In the text-with-audio condition, the written passage was accompanied by continuous English narration corresponding directly to the linguistic content presented visually. The complete written passage remained visible during the learning task while the narration proceeded in the same sequential order as the written text. No additional definitions, glosses, translations, or explanatory information were introduced through the auditory channel. Thus, the auditory component provided a corresponding phonological representation of the written material rather than additional semantic content.
The auditory narration lasted approximately 2 min and was delivered at an average rate of approximately 125–150 words per minute, a presentation rate intended to permit simultaneous processing of the auditory and orthographic information without introducing unnecessarily rapid auditory presentation. Temporal coordination between the two modalities was maintained throughout the task so that the auditory narration corresponded sequentially to the written material being processed.
To maintain standardized exposure conditions, participants were not permitted to pause, replay, skip, or modify the playback speed during the experimental task. This restriction was intended to minimize individual variation in exposure duration and repeated access to particular lexical items. Participants in the text-only condition were likewise exposed to the written material for the standardized task period, thereby maintaining comparable temporal conditions between groups.
The learning materials were presented individually on a laptop or desktop computer using a standardized digital display. Participants assigned to the text-with-audio condition received the auditory narration through individual headphones, whereas auditory output was absent in the text-only condition. The experimental task was conducted individually in a quiet indoor environment to minimize task-irrelevant auditory and visual distraction. The visual presentation of the written material was maintained consistently across the two conditions.
As illustrated in
Figure 1, participants completed the vocabulary pretest before undertaking their assigned learning condition, followed by the immediate vocabulary posttest. Physiological monitoring incorporated a resting baseline period, continuous heart-rate recording during the 120-s task window, and a post-task recovery period. The same task duration was applied across conditions, allowing physiological responses to be compared under standardized temporal conditions.
The experimental configuration was therefore designed to isolate the instructional contrast between text-only reading and synchronized listening-while-reading while maintaining equivalent written linguistic content and exposure conditions. Importantly, the experimentally manipulated variable was input modality rather than cognitive load or learner engagement. Cognitive load and engagement were not directly manipulated or measured in the present experiment. Accordingly, these constructs are treated as potential theoretical mechanisms for interpreting the observed behavioral and physiological patterns rather than as experimentally established mediators.
3.3. Measures
3.3.1. Vocabulary Knowledge
Vocabulary knowledge was assessed using a 20-item researcher-developed multiple-choice test designed to measure lexical form–meaning knowledge associated with the target vocabulary presented in the experimental learning materials. Each item contained four response alternatives and was scored dichotomously as correct or incorrect. Each correct response was assigned 5 points, yielding a total possible score ranging from 0 to 100. The same scoring procedure was applied consistently at pretest and posttest.
The same 20-item assessment was administered immediately before and after the learning task to permit direct comparison of vocabulary performance across measurement occasions. The assessment focused on the lexical content targeted during the experimental task and was administered under the same testing conditions for both the text-only and text-with-audio groups.
Because repeated administration of identical items may introduce item familiarity or practice effects, pretest-to-posttest improvement was not interpreted solely on the basis of separate within-group changes. Instead, the primary inferential analysis examined whether the magnitude of change differed between the two instructional conditions through the Condition × Time interaction. Because both groups completed the same assessment at the same measurement occasions, any general test–retest familiarity would have been shared across conditions; nevertheless, the possibility of practice effects cannot be excluded and is acknowledged as a methodological limitation. Future studies should employ independently validated parallel forms to distinguish learning-related improvement more clearly from repeated-testing effects. Because the vocabulary measure was researcher-developed specifically for the target lexical items used in the experiment, its psychometric properties were not independently validated in a separate sample. Future research should therefore employ independently validated vocabulary measures or parallel test forms to distinguish learning-related improvement more clearly from repeated-testing effects.
3.3.2. Physiological Measures
Heart rate (HR) was continuously recorded using a Polar H10 chest-strap heart-rate sensor (Polar Electro Oy, Kempele, Finland), which uses electrical cardiac sensing to detect beat-to-beat R–R intervals and derive continuous HR measurements. Data were collected during three phases:
Baseline phase: resting HR recorded prior to task onset
Task phase: 120-s reading activity
Recovery phase: post-task resting period
Heart-rate values were sampled at 10-s intervals from 0 to 120 s, resulting in 13 time points per participant. From these time-series data, four derived indices were calculated:
- 1.
Mean Task Heart Rate (MTHR):
The average HR across the 120-s task window.
- 2.
Mean Task Elevation Over Baseline (ΔHR):
MTHR minus baseline HR, indexing task-related activation.
- 3.
Peak Task Heart Rate:
The highest HR value recorded during the task window.
- 4.
Recovery Time:
The elapsed time from task termination until HR first returned to within ±5 bpm of the participant’s individual pre-task baseline HR.
Baseline normalization procedures were applied to account for individual variability in resting heart rate. Physiological indices were interpreted within a cognitive–affective regulation framework rather than as direct proxies for cognitive load. HR was not treated as a direct measure of engagement.
3.4. Data Analysis
3.4.1. Preliminary Analyses
Group equivalence at pretest was evaluated using an independent-samples t-test. Assumptions of normality and homogeneity of variance were examined prior to inferential testing. Descriptive statistics were computed for all behavioral and physiological measures. Outliers were screened using standard deviation criteria and visual inspection of boxplots.
3.4.2. Behavioral Analyses
Vocabulary learning outcomes were analyzed using a 2 × 2 mixed-design analysis, with Condition (text-only vs. text-with-audio) as the between-subjects factor and Time (pretest vs. posttest) as the within-subjects factor. The primary inferential test was the Condition × Time interaction, which directly evaluated whether the magnitude of vocabulary change differed between the two instructional conditions.
To facilitate interpretation of the interaction, individual vocabulary gain scores were defined as posttest minus pretest scores, and the magnitude of improvement was compared between conditions. Within-condition pretest-to-posttest changes were additionally examined using paired-samples t-tests as supplementary analyses.
Effect sizes were reported alongside significance tests. Partial eta squared (ηp
2) was reported for the mixed-design interaction, and Cohen’s d was reported for between-condition differences in gain scores. Ninety-five percent confidence intervals were reported for the between-condition gain difference. Effect sizes were interpreted according to established conventions [
28] while also considering field-specific benchmarks in quantitative applied linguistics [
29].
Assumptions of normality and homogeneity of variance were examined prior to inferential testing. Statistical significance was evaluated at an alpha level of p < 0.05.
3.4.3. Physiological Analyses
Physiological indices (MTHR, ΔHR, peak task HR, peak elevation above baseline, and recovery time) were compared between conditions using independent-samples t-tests. Cohen’s d and 95% confidence intervals for the between-condition mean differences were reported to quantify the magnitude and precision of these comparisons. Because heart rate is a nonspecific indicator of autonomic activation, elevated HR was not interpreted as a direct measure of engagement, cognitive load, or stress. Physiological findings were therefore interpreted conservatively as indices of task-related autonomic activation and considered alongside, but analytically distinct from, behavioral learning outcomes.
All analyses were conducted using standard statistical software. Statistical significance was set at α = 0.05.
5. Discussion
5.1. Interpretation of Behavioral and Physiological Findings
The present study examined whether synchronized listening-while-reading influences vocabulary learning and task-related physiological activation in digital EFL reading. The clearest empirical finding was behavioral: learners in the text-with-audio condition demonstrated substantially greater vocabulary improvement than those in the text-only condition. Because the two groups showed comparable pretest performance and were exposed to the same textual content for a standardized period, the observed difference suggests that the addition of synchronized auditory input altered the conditions under which lexical information was processed and encoded.
The vocabulary findings are broadly consistent with Dual Coding Theory [
26] and multimedia learning principles [
7,
8]. Coordinated auditory and written input may provide complementary phonological and orthographic information, thereby supporting the formation of more differentiated lexical representations. From a usage-based perspective [
5,
6], synchronized presentation may also increase the salience of lexical forms and strengthen associations among orthographic form, phonological representation, and meaning. The present behavioral findings are therefore compatible with the possibility that listening-while-reading supported more effective lexical encoding than text-only exposure.
The physiological findings, however, require a more cautious interpretation. Mean task heart rate and mean elevation above baseline were moderately higher in the text-with-audio condition than in the text-only condition. Importantly, these between-group differences did not reach the conventional threshold for statistical significance (MTHR: p = 0.069, d = 0.59; ΔHR: p = 0.053, d = 0.63). Accordingly, these physiological results should be regarded as suggestive patterns rather than confirmatory evidence of a specific psychological process. In addition, peak task heart rate, peak elevation, and post-task recovery time showed only small and nonsignificant group differences. Taken together, the physiological data indicate a tendency toward greater sustained task-related activation in the text-with-audio condition but do not establish the psychological source of that activation.
This distinction is particularly important for the construct validity of the present study. Heart rate reflects integrated autonomic activity and may vary as a function of cognitive effort, attentional mobilization, emotional arousal, motivational involvement, stress, or combinations of these processes [
22,
23]. Consequently, HR cannot be treated as a direct or specific measure of learner engagement. Likewise, elevated HR cannot by itself establish either cognitive overload or productive engagement. The present study therefore interprets HR conservatively as an index of task-related physiological activation, rather than as a direct proxy for engagement or cognitive load.
The joint pattern of behavioral and physiological findings is nevertheless theoretically informative. The text-with-audio group showed markedly stronger vocabulary learning while also tending to exhibit moderately greater task-related physiological activation. Thus, improved learning did not occur alongside a clear reduction in physiological activation. This observation complicates a simple reduced-strain account in which more effective multimodal learning should necessarily be accompanied by lower autonomic activation. At the same time, the present data do not permit the stronger conclusion that the increased activation represented engagement. Rather, the results indicate that enhanced vocabulary learning can co-occur with moderate task-related physiological activation.
This more conservative interpretation also affects the proposed concept of balanced efficiency. In the present study, balanced efficiency should not be understood as an empirically demonstrated psychological state in which cognitive load and engagement were directly measured and shown to be optimally balanced. Neither construct was directly assessed with a validated self-report or behavioral instrument. Instead, balanced efficiency is proposed here as a provisional interpretive framework for the observed coexistence of improved behavioral performance and non-reduced physiological activation. Within this framework, multimodal input may redistribute processing demands rather than simply minimize them: synchronized audio may reduce some decoding-related demands while other cognitive, attentional, or affective processes remain active.
Accordingly, the present findings do not demonstrate that listening-while-reading increases engagement, nor do they demonstrate that physiological activation itself causes improved vocabulary learning. Rather, they show that synchronized text-audio input was associated with substantially greater vocabulary gains and a tendency toward moderately greater task-related autonomic activation. Establishing whether this activation specifically reflects attentional engagement, germane cognitive processing, affective arousal, or another mechanism will require direct measurement of these constructs in future research.
5.2. Implications for SLA Theory
The present findings have implications for theoretical approaches to second language acquisition that conceptualize learning as emerging from interactions among input characteristics, cognitive processing, attention, and affective-regulatory processes. Robinson’s [
30] Cognition Hypothesis and Schumann’s [
31] neurobiological perspective, for example, emphasize that language learning is influenced not only by linguistic input but also by the processing conditions under which that input is encountered. The present results are compatible with such integrative perspectives because the addition of synchronized auditory input was associated with both altered learning outcomes and a different pattern of task-related physiological activation.
The findings also invite a cautious reconsideration of how Cognitive Load Theory [
9,
10] is applied to multimodal L2 learning. CLT emphasizes the importance of minimizing unnecessary or extraneous processing demands so that limited cognitive resources can be allocated to learning-relevant processes. Listening-while-reading may plausibly reduce some decoding demands by providing phonological information simultaneously with written forms. However, the present physiological results do not provide direct evidence that cognitive load itself was reduced. Indeed, HR did not decrease in the text-with-audio condition. Because HR is not a specific measure of cognitive load, this finding should not be interpreted as evidence against CLT; rather, it illustrates the difficulty of inferring cognitive load from autonomic activation alone.
A potentially useful interpretation is that multimodal input changes the allocation, rather than simply the overall amount, of processing effort. Synchronized auditory support may reduce uncertainty associated with phonological decoding while simultaneously recruiting other processes related to temporal coordination, lexical integration, attention, or affective response. Under such circumstances, improved learning would not necessarily be accompanied by lower overall physiological activation. This interpretation remains theoretical, but it offers a more nuanced account of why the text-with-audio condition could produce strong vocabulary gains without a corresponding reduction in HR.
The findings similarly relate to engagement-based perspectives in SLA [
17,
20,
21], but the distinction between theoretical relevance and empirical measurement is critical. Engagement is generally conceptualized as a multidimensional construct involving cognitive, behavioral, and affective components. Because these dimensions were not directly measured in the present study, the HR findings cannot verify increased engagement. Rather, engagement provides one possible theoretical lens through which the co-occurrence of improved learning and physiological activation may be examined. Future studies should directly assess cognitive, behavioral, and affective engagement to determine whether these dimensions mediate or moderate the effects of multimodal listening-while-reading.
From a usage-based perspective [
5,
6], the behavioral findings may indicate that synchronized auditory and written input increased the salience of lexical forms and supported repeated activation of form–meaning relationships. The strong vocabulary gains observed in the text-with-audio condition are compatible with this account. Nevertheless, the present experiment does not identify the specific cognitive mechanism responsible for these gains. Possible explanations include improved phonological–orthographic mapping, increased attentional allocation, deeper lexical processing, or combinations of these mechanisms.
Methodologically, the study therefore illustrates both the value and the limitations of incorporating psychophysiological measures into SLA research. Physiological data can provide information about changes in bodily activation that may not be captured by behavioral outcomes alone, but their interpretation requires triangulation with measures that more directly assess the psychological constructs of interest. HR should consequently be treated as a complementary process-level measure rather than as a standalone indicator of cognitive load or engagement. Future research combining HR and HRV with validated workload scales, behavioral engagement measures, eye-tracking, pupillometry, and other process measures would allow for stronger inferences regarding the mechanisms underlying multimodal language learning.
Overall, the present findings support a provisional balanced-efficiency account rather than confirming a fully specified mechanistic model. The behavioral results provide strong evidence that synchronized listening-while-reading benefited vocabulary learning in the present sample, whereas the physiological results suggest that this benefit did not require a reduction in task-related autonomic activation. Whether this pattern reflects productive engagement, redistribution of cognitive resources, increased attentional mobilization, or another cognitive–affective mechanism remains an empirical question for future research.
5.3. Implications for Educational Practice
The present findings also have implications for the design of digitally mediated language-learning environments. Most importantly, the behavioral results suggest that synchronized auditory support can be a useful instructional feature for vocabulary learning when audio and written text are temporally coordinated. Rather than viewing multimodal design solely as a means of making learning easier, instructional designers may consider how different modalities distribute information and support complementary aspects of lexical processing.
For digital reading platforms, synchronized narration may assist learners by providing phonological information while written forms remain visually available. Such coordination may be particularly useful for learners who have difficulty mapping orthographic forms onto pronunciation or maintaining fluent processing of unfamiliar lexical items. However, the present findings should not be interpreted as showing that greater physiological activation is itself pedagogically desirable. The study did not experimentally manipulate physiological arousal, nor did it establish an optimal level of HR for learning.
Similarly, educators should not use physiological activation alone as an indicator that learners are either effectively engaged or cognitively overloaded. A learner displaying greater autonomic activation may be investing cognitive effort, attending closely, experiencing emotional arousal, encountering difficulty, or responding to other task characteristics. Consequently, instructional decisions should be based primarily on learning performance and direct indicators of learner experience rather than on physiological activation in isolation.
For developers of adaptive or AI-supported learning environments, the present findings underscore the potential value of combining multiple sources of information when evaluating learner states. Behavioral indicators such as response accuracy, time-on-task, rereading behavior, and interaction patterns could be considered together with subjective workload or engagement measures and, where appropriate, physiological signals. Such multimodal assessment may eventually support more precise adaptation of instructional difficulty and modality presentation. However, physiological measures should be regarded as complementary signals whose psychological meaning must be validated rather than assumed.
The practical implication of the present study is therefore not that digital learning systems should seek to increase or decrease physiological activation. Instead, the findings suggest that successful multimodal learning can occur even when improved performance is not accompanied by reduced physiological activation. For educators and designers, the more relevant goal is to create conditions that facilitate effective lexical processing while monitoring whether learners achieve meaningful learning outcomes without evidence of excessive difficulty or disengagement.
In this respect, synchronized listening-while-reading appears promising as an instructional approach for digital EFL vocabulary learning. Nevertheless, its effectiveness should be evaluated across different proficiency levels, task types, text complexities, and learning durations before broader pedagogical recommendations are made.
For clarity, the present study distinguishes among three levels of inference. First, the study directly demonstrates differences in behavioral vocabulary outcomes between the two instructional conditions. Second, it directly measures task-related heart-rate activation, although the principal between-group HR differences were moderate in magnitude and did not reach conventional statistical significance. Third, constructs such as cognitive load, engagement, attentional mobilization, and affective arousal represent theoretical interpretations rather than directly measured psychological variables. The proposed balanced-efficiency framework should therefore be understood as a provisional account integrating the observed behavioral and physiological patterns, rather than as a validated causal model of the psychological mechanisms underlying listening-while-reading.
5.4. Practical Implications for Multimodal CALL Design
The findings have practical implications for multimodal CALL design. The behavioral results suggest that temporally synchronized auditory and textual input can support immediate vocabulary learning more effectively than text-only presentation under the conditions examined here. Thus, effective multimodal design may depend less on simply adding multiple input channels than on coordinating auditory and textual information to support integrated lexical processing. Future CALL systems could extend this principle through learner-controlled, adaptive, or personalized features such as narration pacing, text highlighting, and audio–text synchronization tailored to learners’ processing needs. More broadly, recent computational research has highlighted the value of adaptive information aggregation and complementary fusion for preserving and integrating task-relevant information across complex information streams [
32]. Although these approaches were developed outside the domain of language learning, they provide a useful computational analogy for future adaptive multimodal CALL systems in which auditory and textual information may be coordinated dynamically according to learners’ processing needs. However, because the physiological findings of the present study were suggestive rather than confirmatory, physiological activation should not be used alone to infer cognitive load or engagement or to determine adaptive system responses. Future adaptive CALL research should therefore integrate behavioral performance with direct measures of learner experience and complementary physiological indicators.
6. Limitations and Future Research
Several limitations of the present study should be acknowledged.
First, the sample size (
n = 20 per condition) limits statistical precision, particularly for detecting smaller physiological differences. The relatively small sample, drawn from Korean university EFL learners, also limits the generalizability of the findings to broader EFL populations and instructional contexts. Although large behavioral effect sizes were observed for vocabulary outcomes [
29], the magnitude and stability of these effects require replication in larger and more heterogeneous samples. Future research should employ larger samples to test potential moderating variables such as proficiency level, working memory capacity, or language anxiety. Such analyses could clarify whether the behavioral and physiological patterns interpreted within the provisional balanced-efficiency framework generalize across different learner populations and educational contexts.
Second, heart-rate measures were used as indicators of task-related physiological activation. While HR reflects overall autonomic activity, it cannot clearly distinguish cognitive effort from attentional engagement, affective arousal, or stress-related responses. In addition, affective states were not directly assessed using validated self-report or emotion-specific measures; therefore, references to affective processes in the present study should be understood as theoretical interpretations rather than empirically differentiated emotional responses. The present study also did not include heart-rate variability (HRV), which could provide complementary information about autonomic regulation during learning. Future research should therefore combine HR and HRV with direct measures of cognitive load and engagement, such as subjective workload and behavioral indicators, to more clearly characterize cognitive–affective states during multimodal learning. Such multimethod assessment would help determine whether physiological activation is associated with cognitive effort, attentional engagement, affective arousal, or excessive strain without attributing it to any single psychological process.
Third, the same vocabulary assessment was administered at pretest and immediate posttest, which may have introduced some degree of item familiarity or practice effect. Because both experimental conditions completed the same assessment at the same measurement occasions, any general test–retest familiarity would have been shared across conditions; however, its influence on the observed vocabulary gains cannot be completely excluded. Future studies should employ independently validated parallel test forms to more clearly distinguish learning-related improvement from repeated-testing effects.
Fourth, the present study assessed vocabulary learning immediately after the experimental task and did not include a delayed posttest. Accordingly, the findings should be interpreted as evidence of immediate vocabulary learning rather than long-term lexical retention. Future research should incorporate delayed assessments across multiple time intervals to determine whether the observed advantage of synchronized listening-while-reading persists over time. Such longitudinal designs would help distinguish short-term performance gains from durable lexical learning.
Finally, expanding this research across diverse linguistic backgrounds and instructional contexts would enhance generalizability. Cross-linguistic comparisons could determine whether phonological transparency or orthographic distance moderates the benefits of synchronized audio. The present findings may inform theoretical models of second language acquisition that conceptualize learning as emerging from interactions among input properties, attentional allocation, affective processes, and cognitive regulation [
30,
31]. Accordingly,
Figure 2 is presented as a provisional conceptual framework intended to generate testable hypotheses for future research rather than as an empirically validated model.
7. Conclusions
The present study examined the role of multimodal listening-while-reading in digital EFL vocabulary learning by combining behavioral learning outcomes with physiological indicators of task-related autonomic activation. The results provide evidence that synchronized auditory and textual input was associated with substantially greater immediate vocabulary gains than text-only reading under the present experimental conditions. Learners in the listening-while-reading condition demonstrated substantial vocabulary gains, suggesting that coordinated multimodal input may support lexical encoding and the integration of phonological and orthographic information.
However, the physiological findings suggest that the relationship between multimodal learning and physiological activation may be more complex than would be predicted by a simple reduced-activation account. Instead of showing lower physiological activation, learners in the listening-while-reading condition showed numerically higher mean task heart rate and elevation from baseline, with moderate effect sizes; however, these between-condition differences did not reach conventional statistical significance. Peak activation and recovery measures likewise showed no significant between-condition differences. Accordingly, the physiological findings should be regarded as suggestive rather than confirmatory. Because heart rate is a nonspecific index of autonomic activation, the observed pattern cannot establish whether the numerical differences in activation reflected attentional engagement, cognitive effort, affective arousal, or other processes.
Based on these findings, this study proposes a provisional balanced-efficiency framework for interpreting multimodal digital listening-while-reading. Rather than assuming that effective learning necessarily requires reduced physiological activation, the framework considers the possibility that successful multimodal learning may coexist with sustained task-related activation. In listening–reading contexts, synchronized audio may support the coordination of phonological and orthographic information and alter the distribution of processing demands; however, these mechanisms were not directly measured in the present study. Accordingly, balanced efficiency should be regarded as a hypothesis-generating account rather than a validated causal mechanism.
From a theoretical perspective, the findings raise the possibility that effective second-language vocabulary learning cannot be explained solely in terms of minimizing cognitive or physiological demands. Instead, the behavioral and physiological patterns observed here provide a basis for future research examining how cognitive effort, attentional engagement, and multimodal processing may jointly contribute to learning. Because cognitive load and engagement were not directly measured, the present study does not establish how these processes were calibrated or whether the observed physiological activation reflected productive attentional mobilization.
More broadly, this study contributes to an emerging direction in SLA research that integrates behavioral outcomes with process-level indicators. By examining vocabulary gains alongside task-related physiological activation, the study illustrates the potential value of combining behavioral and physiological evidence while maintaining a clear distinction between directly observed outcomes and theoretically inferred mechanisms. Future research incorporating direct measures of cognitive load and engagement, complementary physiological indices such as HRV, delayed posttests, larger samples, and diverse learner populations will be necessary to evaluate the proposed framework and determine whether the observed immediate learning advantage persists over time.
In conclusion, the present findings provide evidence of greater immediate vocabulary learning in the listening-while-reading condition while showing only suggestive differences in task-related physiological activation. This pattern is compatible with—but does not directly demonstrate—the proposed provisional balanced-efficiency framework.