1. Introduction
In Spain, the topic of audiovisual accessibility, researched within the framework of Audiovisual Translation Studies (TAV), has aroused increasing interest in recent years. This effort has driven the development of techniques such as audio description (hereinafter, AD), as a service consisting of the set of techniques and tools that, applied to audiovisual communication, have the objective of compensating for the lack of perception of visual messages by people with visual impairment, transforming these messages into auditory messages [
1].
While AD has consolidated itself as a fundamental tool for inclusion, research in this field has developed a predominantly linguistic focus. Initial studies have concentrated on textual aspects, such as the type of language used to describe the audiovisual product, the described content, or the synchronization of the AD with the visual elements [
2,
3,
4,
5,
6]. Subsequent studies expanded this scope by assessing essential facets of AD quality and user experience, such as user satisfaction, levels of presence or narrative immersion [
7], comprehension [
8], and memory recall [
9]. More recent work has further explored emotional responses to AD through both self-reported measures and physiological data [
5,
10]. However, only a limited number of studies have examined the role of the narrator’s voice [
11,
12,
13], despite the fact that vocal parameters strongly influence the emotional reception of the description.
The narrator’s voice has traditionally been treated as a neutral conduit for information [
11,
13]. This minimalist conception overlooks its fundamental role in shaping the reception and emotional resonance of the described content. Recent studies in psycholinguistics, affective neuroscience, and media psychology suggest that vocal features such as pitch, tone, timbre, rhythm, and intonation can profoundly influence how narratives are processed cognitively and emotionally [
14,
15]. These vocal dimensions do not act independently of the message; rather, they modulate its interpretation, affecting comprehension, memory, empathy, and enjoyment. Essentially, the listener’s brain does not process the words alone; it responds to the sound of the voice itself as a social and emotional signal [
14].
This gap becomes particularly relevant in the context of the performing arts, where AD serves not only an informative but also an esthetic and emotional function [
4,
5,
7]. In contemporary dance, for instance, where visual content is often abstract and multisensory, mediation strategies must go beyond mere description to offer an immersive and emotionally resonant experience. In such contexts, the narrator’s voice becomes a central expressive channel, capable of shaping tone, atmosphere, and the audience’s emotional engagement.
The present study aims to address this gap by examining how narrator gender (female vs. male) influences the reception of poetic AD in contemporary dance. Through a multimethod approach that integrates subjective self-report data with physiological markers, such as electrodermal activity and heart rate variability, we investigate how different vocal profiles modulate the audience’s cognitive and emotional experience. Ultimately, this research seeks to provide empirical evidence to optimize vocal selection in the creation of accessible and inclusive artistic media, advocating for a more personalized and emotionally resonant approach to AD.
2. The Role of Narrator Voice in AD
The pleasantness of the voice as a function of gender is a difficult concept to define from an objective standpoint, given that it is intrinsically linked to the subjective preference of everyone receiving the stimulus. Nonetheless, there is a growing recognition that the voice is not merely a technical characteristic but a central semiotic and affective component of the experience [
15].
In this regard, recent research in the field of AD has indicated that certain vocal profiles, conventionally associated with female voices, are often rated as more pleasant, expressive, and emotionally engaging in a narrative context, whereas those perceived as male may be associated with neutrality or authority [
13,
15]. Rather than being inherent gender traits, these differences may stem from specific acoustic parameters, such as higher fundamental frequency, varied intonation patterns, and specific timbre, that have been shown to influence listener reception [
16,
17]. Moreover, neurocognitive findings suggest that these vocal qualities inherent to female voices may potentially modulate the mental load required for comprehension, thereby enhancing processing efficiency depending on the complexity of the task [
15]. These patterns of brain activation are often associated with social signals such as trust and emotional salience, which are deeply influenced by the listener’s expectations and the communicative context [
18].
This emphasis on the affective role of the voice is particularly critical in the context of the performing arts such as dance and other non-verbal art forms, where meaning is often abstract, metaphorical, and embodied. In these arts forms, AD functions not only as an informative tool but as an esthetic and emotional bridge between the performance and the audience. The narrator’s voice becomes part of the expressive fabric, contributing to the overall tone and atmosphere of the piece [
4]. Studies on vocal expression in dance [
6,
7,
14,
15] suggest that a voice with slow tempo, warm timbre, and dynamic prosodic modulation may enhance emotional engagement and the sense of presence among audiences, particularly those with visual impairments. In contrast, monotonous or robotic interpretations tend to reduce immersion and affective resonance, thereby diminishing the esthetic experience. These findings highlight the relevance of prosody as a structural component in meaning construction and in shaping the emotional and narrative framing of contemporary dance, contributing to expressive coherence and semantic depth within the artistic experience.
Therefore, measuring the effectiveness of this expressive delivery requires robust metrics for deep engagement. A crucial concept for measuring the effectiveness of expressive AD, particularly in artistic contexts, is narrative transportation [
5,
6], a mechanism describing the process by which an individual becomes immersed in a narrative. It represents a feeling of being completely absorbed or carried away by the story, often resulting in a deep cognitive and emotional engagement with the narrative world. Vocal elements, such as prosody and affective delivery, are integral to achieving transportation, as they contribute to the narrative’s credibility, emotional texture, and rhythm. High transportation is associated with greater levels of enjoyment and a more profound esthetic experience. Consequently, transportation serves as an essential subjective measure for assessing the success of a creative vocal approach in AD [
6,
13,
19]. This approach ensures that the evaluation of AD goes beyond mere information delivery to encompass the full esthetic richness of the performance.
3. The Study
3.1. Aim, Variables and Hypothesis
Our aim was to evaluate audience reception of AD in contemporary dance, focusing on the narrator’s voice (male vs. female).
To explore reception, the construct was operationalized by defining a ‘good’ AD through specific criteria: it should be easy to process, requiring minimal cognitive effort; it must be emotionally engaging and enjoyable, aiming to evoke feelings and enhance immersion and satisfaction for individuals with visual impairments; and it must be functional, aiding in the recall of the described content.
The following variables were selected in order to obtain empirical evidence to test the hypotheses:
Cognitive effort was assessed using both subjective and objective measures. Subjectively, the effort perceived by participants was measured using self-reports to capture their personal assessment of cognitive load. Objectively, cardiac deceleration was monitored, with a slower heart rate serving as an indicator of greater cognitive load.
Emotional salience and pleasure were also assessed using both subjective and objective measures. Subjectively, participants completed self-reports on valence and arousal to capture their emotional responses, as well as on their sense of immersion and enjoyment of the performance to measure emotional engagement. Objectively, electrodermal activity (EDA) was measured, with higher levels indicating increased arousal, and high-frequency heart rate variability (HRV-HF) was monitored, with higher HRV suggesting the body’s response to emotional stimuli and the need to regulate arousal.
Derived from these variables, five hypotheses were stated, assuming that the gender of the voice acts as a proxy for a cluster of acoustic and prosodic features:
H1. Compared to AD voiced by a male, AD voiced by a female is expected to be associated with less cognitive effort.
H1a. Compared to AD voiced by a male, AD voiced by a female is expected to be associated with less subjective cognitive effort measured with self-report.
H1b. Compared to AD voiced by a male, AD voiced by a female is expected to be associated with less objective cognitive effort measured with heart rate.
H2. Compared to AD voiced by a male, AD voiced by a female is expected to be associated with a higher emotional impact.
H2a. Compared to AD voiced by a male, AD voiced by a female is expected to be perceived as more pleasant—measured with self-reported valence, arousal, enjoyment and transportation.
H2b. Compared to AD voiced by a male, AD voiced by a female is expected to be associated with higher emotional salience—measured with phasic skin conductance and heart rate variability.
H3. Compared to AD voiced by a male, AD voiced by a female is expected to be linked to higher cognitive engagement—measured with recall accuracy.
Hypotheses H1a and H1b, which concern cognitive effort, are based on evidence that certain vocal profiles conventionally associated with female prosody are often perceived as more intelligible or pleasant, potentially reducing the mental load required for comprehension [
15]. This is particularly relevant in dance AD, where the high level of abstraction and the predominance of non-verbal expression pose significant challenges for audiences with visual impairments in constructing meaning from the performance.
The affective hypotheses (H2a and H2b) reflect findings from emotional prosody research, which indicate that specific vocal qualities, such as those often present in female voices, tend to elicit stronger empathetic responses and are more effective in conveying emotional subtleties [
16]. This aligns with the idea that poetic AD delivered by a female voice might enhance both the emotional valence and arousal experienced by the listener, as well as support emotional regulation, as inferred from physiological data (e.g., HRV and EDA measures).
Finally, hypothesis H3, addressing cognitive engagement via recall performance, is rooted in theories of narrative transportation. These theories suggest that vocal qualities influence immersion in a narrative and the memorability of its content [
6,
19]. A more engaging voice may not only support affective involvement but also facilitate the encoding and retrieval of information.
3.2. Sample
Our sample included 33 participants, all of whom had less than 10% residual vision and no other known disabilities. The group consisted of 13 men and 20 women, with ages ranging from 16 to 80 years. While the broad age range may be considered a limitation, it is representative of the practical difficulties inherent in recruiting participants from minority communities, where sample accessibility is often constrained. Participant recruitment was carried out with the support of ONCE, the Spanish National Organization for the Blind, in the cities of Murcia and Granada. The study was conducted in compliance with the ethical guidelines of the University of Murcia’s Ethics Committee.
The sample intentionally covered a wide age range (16–80 years) to better reflect the demographic diversity of the visual impairment community and the variety of real-world uses and preferences within it. We acknowledge that age can be associated with variability in cognitive processing, emotional responding, and familiarity with contemporary dance. However, age was not used as a static control variable because restricting the age range would have reduced the ecological validity of the study and limited the representativeness of the target community. Importantly, the study used a within-subjects design: each participant experienced both vocal conditions and therefore served as their own control. As a result, the main comparisons are based on intra-individual differences in reception across conditions, which reduces the risk that between-participant factors, such as age, account for the observed effects.
3.3. Materials
The materials for the study consisted of two ten-minute excerpts extracted from two different full-length contemporary dance performances (hereinafter A and B). From each dance performance, a continuous segment of ten minutes of choreography was selected to serve as the basis for the analysis. Each ten-minute excerpt was then divided into two equal parts (A1, A2; B1, B2): the first five minutes (A1 and B1) were audio described using a male voice, and the second five minutes (A2 and B2) using a female voice. The ADs followed a poetic and expressive style, in line with recent proposals that advocate for more interpretive strategies in dance AD to better capture its symbolic, emotional and esthetic dimensions [
5].
The scripts were created by the research team, reviewed by professional dancers to ensure artistic coherence, and recorded using high-quality synthetic voices provided by Aptent, a specialized company in media accessibility. The specific voices were selected to represent prototypical gendered vocal qualities, ensuring a clear perceptual distinction for the participants. Both voices were generated using the same synthesis technology to ensure consistent audio quality, speech rate, and prosodic neutrality. The female voice was characterized by a higher fundamental frequency and a brighter timbre, while the male voice featured a lower pitch and a deeper resonance. This controlled selection ensured that the observed differences in reception were primarily attributable to the perceived gender of the voice rather than discrepancies in synthesis quality or delivery speed.
To elaborate poetic AD, we used similes, descriptions of emotional state of the performers, terms belonging to a poetic register, and rhythmic repetitions.
Table 1 (below) presents an example of each parameter.
A similar number of tokens of each category was included in the two clips. We also maintained a consistent total number of words in each AD to ensure that the descriptions are balanced and concise. To guarantee that both versions were of similar difficulty, we run several readability tests on the AD scripts with the online tool
https://legible.es/ (accessed on 21 April 2023). Specifically, we employed the Fernández Huerta index [
20], a standard for the Spanish language based on the average number of syllables per word and the mean number of words per sentence. This approach helped ensure that both poetic ADs were stylistically rich yet comparable in structure and complexity.
Table 2 (below) presents the number of occurrences for each parameter.
A within-subjects design was implemented to control inter-subject variability. To guarantee a balanced and unbiased presentation of the AD clips, we used a counterbalanced design for the order in which participants (P) experienced the different AD versions. As shown in
Table 3, the first participant first listened to the two ADs recorded with a male voice (A1 followed by B1), and then to the two ADs with a female voice in the reverse order (B2 followed by A1). The subsequent participant listened to the same ADs in a different sequence, and so on, ensuring that all possible combinations were evenly represented across the sample.
3.4. Instruments and Variable Measurements
To obtain a comprehensive understanding of how participants perceived and processed the AD of contemporary dance performances, a broad array of assessment instruments was employed. The experimental protocol and the instruments used for data collection in this study follow the multimethod framework previously established in our study on poetic versus neutral AD [
5]. These instruments were specifically selected to capture both the subjective experience of the participants and their objective physiological responses, allowing for a richer and more triangulated interpretation of the data.
To assess cognitive load, we employed the Mental-Effort Rating Scale developed by Paas [
21]. This instrument consists of a single-item measure, rated on a 9-point Likert scale, which reflects participants’ self-perceived mental effort during task engagement. Despite its simplicity, the scale is widely validated in cognitive psychology and instructional research and offers a direct insight into participants’ internal evaluations of task difficulty.
Emotional responses, particularly valence (positive vs. negative feelings) and arousal (level of physiological activation), were evaluated using the tactile version of the Self-Assessment Manikin (SAM) developed by Iturregui Gallardo [
22]. This instrument is especially suited for populations with visual impairments, as it provides accessible tactile representations of emotional states. The tactile SAM consists of two items, one representing valence and the other arousal, and offers a user-friendly and inclusive tool for capturing core affective dimensions.
To measure participants’ narrative engagement, we used the Narrative Transportation Questionnaire developed by Green and Brock [
19], which comprises twelve items. This scale evaluates the extent to which individuals feel mentally and emotionally “transported” into a story or performance, thus providing insight into the immersive qualities of the AD. In parallel, enjoyment was assessed through the Aesthetic Experience Questionnaire [
23], a ten-item instrument designed to capture subjective appreciation and satisfaction with artistic experiences.
To evaluate the functional usefulness of the AD (that is, its ability to facilitate understanding and retention of content) participants completed a recall test composed of four items. This task was designed to objectively measure how much descriptive information had been internalized and remembered, thereby linking the AD’s stylistic and emotional effects to its informational utility.
In addition to these self-report instruments, we also recorded physiological data to provide objective, continuous markers of participants’ internal states. These data were gathered using the Shimmer 3 GSR+ sensor, a versatile and non-invasive wearable device capable of monitoring two critical physiological indicators: electrodermal activity (EDA) and high-frequency heart rate variability (HF-HRV).
Electrodermal activity, also known as skin conductance, is a widely recognized measure of sympathetic nervous system activation. It captures changes in skin moisture resulting from activity in the sweat glands, which in turn reflect fluctuations in emotional arousal [
24]. States such as excitement, surprise, or stress are associated with increased skin conductance. Through EDA analysis, we were able to infer both the intensity and temporal immediacy of participants’ emotional engagement with the AD and the performance [
22]. This method is particularly advantageous when working with participants with visual disabilities, as traditional visual-based physiological measures, such as pupil dilation, are not applicable in this population [
25].
To capture the parasympathetic branch of autonomic activity, complementary to EDA, we measured HF-HRV, a biomarker associated with emotional self-regulation and positive affect [
26,
27]. HF-HRV reflects the organism’s capacity to return to a calm and regulated state following emotional or cognitive stimulation. Commonly referred to as the “rest and digest” response, higher HF-HRV values are associated with greater emotional resilience, psychological flexibility, and social attunement [
28]. These measurements were acquired through photoplethysmography (PPG) sensors, and frequency-domain analysis was used to isolate the high-frequency components specific to parasympathetic activity [
29].
Cardiac deceleration is a robust and ecologically valid operationalization of cognitive effort as phasic heart rate slowing indexes the parasympathetic-mediated orienting response and sustained attentional allocation required for processing complex dynamic stimuli [
30,
31]. Research consistently demonstrates that these cardiac decelerations track moment-to-moment cognitive resource allocation to salient narrative and structural events in video media [
32,
33], providing a granular measure of information processing that is distinct from generalized arousal [
34]. Furthermore, the synchronization of these cardiac dynamics across viewers predicts comprehension and memory retention [
35], validating cardiac deceleration as a reliable metric for capturing the fluctuating attentional demands involved in interpreting multimodal content [
36,
37].
Finally, a semi-structured retrospective interview was conducted at the end of the experimental session to complement the quantitative data with qualitative insights. During this interview, participants were asked to explicitly state their preference between the male and female narrator voices and to provide a brief justification for their choice. This approach allowed us to capture the subjective reasoning behind their preferences, such as perceived naturalness, clarity, or comfort, and provided a basis for triangulating the explicit feedback with the psychometric and physiological markers previously described.
By combining psychophysiological measures with subjective self-reports, this study adopted a multimethod approach to better understand how poetic AD is experienced cognitively and emotionally by populations with visual impairments. This integration enabled a richer exploration of how stylistic and vocal elements of the AD influence both the conscious perceptions and the unconscious bodily reactions of the audience during the esthetic encounter with contemporary dance.
In
Table 4, the description and values of the study variables are shown.
3.5. Procedure
The experimental sessions were conducted at the ONCE headquarters in the Spanish cities of Murcia and Granada. At both locations, the data collection took place in a quiet room, isolated from external noise and visual distractions, and equipped with only a table and two chairs to maintain a controlled and minimalistic environment.
Upon arrival, participants were welcomed and formally registered. They then received a clear and detailed explanation of the study’s objectives, procedures, potential risks, and expected benefits. In accordance with ethical research standards, participants were asked to read and sign an informed consent form, confirming their voluntary participation and understanding of their rights and the nature of the study.
Following this initial step, physiological monitoring devices were fitted. The Shimmer4 GSR+ sensor was used to collect both electrodermal activity (EDA) and photoplethysmography (PPG) data. Electrodes for EDA measurement were placed on the index and middle fingers of the non-dominant hand, following standardized protocols to ensure signal accuracy and minimize motion artifacts. The EDA signal was sampled at 128 Hz using Consensys software (version 1.6.0) and subsequently decomposed into tonic and phasic components using Continuous Decomposition Analysis [
38]. Each segment of analysis corresponded to a five-minute AD video clip.
Tonic EDA reflects slow-changing levels of baseline sympathetic nervous system activity and can be influenced by factors such as ambient temperature, individual stress levels, and physiological differences.
Phasic EDA, in contrast, captures short-lived, stimulus-driven fluctuations in skin conductivity, offering a marker of immediate arousal and emotional reactivity in response to events presented in the A dance performance.
The photoplethysmography signal was also sampled at 128 Hz using the Shimmer4 sensor, placed on the tip of the index finger of the non-dominant hand. The PPG signal was processed using Kubios HRV software version 2 [
39]. After removing motion artifacts and detrending the signal (first-order component), a Fourier transform was applied to calculate high-frequency heart rate variability (HF-HRV). HF-HRV served as a marker of parasympathetic nervous system activity, specifically the body’s ability to recover from stimulation and return to a resting state. The unit of analysis for both EDA and HRV measures was the 5 min duration of each video clip.
To establish a physiological baseline, participants first engaged in a five-minute relaxation task, which involved passively listening to soothing instrumental music. During this period, physiological signals were recorded continuously. After the relaxation phase, data collection was paused, and participants completed the Self-Assessment Manikin (SAM) to evaluate their baseline affective state.
The experimental phase then commenced. Each participant listened to 10 min of dance audio described by a male voice (5 min taken from dance piece A and 5 min from dance piece B) and another 10 min audio described by a female voice (5 min taken from dance piece A and 5 min from dance piece B). They listened to the first 10 min AD dance clip, while physiological data continued to be recorded. At the end of this segment, data collection was paused again, and participants completed a series of self-report questionnaires assessing cognitive effort, emotional response, enjoyment, and narrative engagement.
This process was repeated for the second 10 min AD clip. Once again, Shimmer devices recorded physiological data during the listening phase, followed by the final set of self-report measures after the clip ended. The order of the clips was counterbalanced across participants to control for order effects. The complete sequence of steps is summarized in
Table 5, which outlines the full study protocol in detail.
4. Results
To examine the impact of narrator voice gender on participants’ reception of AD in contemporary dance, a series of fixed-effects panel regression models were employed. A summary of descriptive statistics for all measured variables can be found in
Table 6. The analytical framework defined the fundamental unit of analysis as individual trials, each comprising a 5 min AD segment. Given that four separate trials were collected from every participant, treating these observations as independent was inappropriate. Consequently, we employed panel regression models, specifying that the standard errors be clustered at the participant level (using the xtreg command in STATA 13).
All the self-reported measures of SAMValence, SAMArousal, Transportation, and Enjoyment were significantly and positively correlated. This indicates that, the more positive the valence and the more intense the reported experience, the greater the enjoyment and the sense of transportation or immersion. Taken together, these measures reflected the overall level of positivity that a person expressed towards the clip. In contrast, self-reported Cognitive Effort was not related to other survey measures, indicating that effort was experienced as a separate, unrelated phenomenon. This lack of statistical relationship is evidenced in
Table 7. When examining the row or column for Cognitive Effort, the correlations with SAMArousal, SAMValence, Transportation, and Enjoyment are found to be not statistically significant.
Beyond the regression models, a specific analysis of narrator preference was conducted to address potential gender-based patterns among the participants, which included 20 women and 13 men. While the general preference for the female voice was 64%, representing 21 out of 33 participants, a distinct divergence emerged when segmenting by the participants’ gender. Among male participants, 77% (10 out of 13) expressed a preference for the female narrator voice. In contrast, female participants showed a more balanced distribution, with 55% (11 out of 20) indicating a preference for the female voice and 45% (9 out of 20) opting for the male voice. This suggests that while the overall statistical preference favors the female narrator, this is more heavily driven by the strong consensus among male listeners, whereas female listeners demonstrate a much more divided stance regarding vocal gender.
We also conducted correlations with physiological data to interpret and determine whether participants in this study interpreted their arousal positively or negatively.
Phasic EDA was negatively correlated with all the self-reported measures (except cognitive effort), and Tonic EDA was negatively correlated with Transportation and Enjoyment. This suggests that objective physiological arousal during the audio-described clips in this study was experienced as “negative” arousal, which could be related to the difficulty of the experience or frustration. In contrast to EDA, Heart Rate was positively correlated with Transportation and SAMArousal, indicating that Heart Rate might have indexed the “positive” component of physiological arousal in this task. Finally, correlations revealed that high frequency HRV was negatively correlated with SAMArousal, SAMValence and Transportation. Because EDA is a result of sympathetic autonomic activity, while HRV is a result of parasympathetic activity, it appears that the activation of both branches of the autonomic nervous system during the clips was experienced “negatively”. This is consistent with the notion that higher HRV is indicative of the need to engage in active emotional self-regulation. Consequently, it appears that participants in this study rated clips lower if they experienced more physiological components of emotions and had to engage in self-regulation as a result.
Taken together, these correlations successfully integrate the disparate variables, allowing for a comprehensive understanding of the participants’ reception process. This integrated data approach is the essential foundation for the subsequent hypothesis testing presented in Section Hypothesis Testing.
Collectively, these patterns highlight the interactions between different variables, giving a more integrated view of participants’ responses. These findings provide the basis for the hypothesis testing presented in the next Section Hypothesis Testing.
Hypothesis Testing
We tested the hypotheses regarding the gender of the voice in AD (Hypotheses H1a, H1b, H2a, H2b, and H3). The results of the panel regression analyses for these hypotheses are reported in
Table 8.
H1a. Compared to AD voiced by a male, AD voiced by a female causes less subjective cognitive effort, measured with self-report.
H1b. Compared to AD voiced by a male, AD voiced by a female causes less objective cognitive effort, measured with heart rate.
H2a. Compared to AD voiced by a male, AD voiced by a female is perceived as more pleasant, measured with self-reported valence, arousal, enjoyment and transportation.
H2b. Compared to AD voiced by a male, AD voiced by a female is more emotionally salient, measured with phasic skin conductance and heart rate variability.
H3. Compared to AD voiced by a male, AD voiced by a female is more cognitively engaging, measured with recall accuracy.
The results provided support for Hypothesis H1a by showing that self-reported Cognitive Effort decreases when the voice is female, with a significant negative coefficient (B = −0.80, p < 0.01). However, no effect of the voice gender was found for Heart Rate during the clips, giving no support for Hypothesis H1b.
Consistently with the drop in perceived Cognitive Effort, participants found female voices more positive (B = 0.46, p < 0.01) and more arousing (B = 0.45, p < 0.01). However, the female voice did not affect Enjoyment or Transportation. This gave only partial support to Hypothesis H2a, showing that female voices are perceived as more pleasant.
Furthermore, female voices elicited significantly lower Phasic EDA (B = −3.84, p < 0.05). This result, in the context of the other findings, seems to suggest that female voices were less frustrating to participants. As such, we have not found support for Hypothesis H2b, which predicted that female voices would be more emotionally salient. On the contrary, it appears that the female voice was soothing. Finally, no significant effect was found for Hypothesis H3 (Cognitive Engagement, Recall Accuracy).
Despite the mixed findings regarding affective responses, the results collected in the retrospective interviews showed a significant preference: 64% of participants preferred the female voice, compared to 36% who preferred the male voice. This clear preference aligns with the quantitative data showing reduced cognitive effort and higher self-reported pleasantness for the female voice. Reasons often cited included the female voice being perceived as more “natural” and more “relaxing”, which corroborates the lower self-reported effort (H1a) and reduced negative physiological arousal (H2b) observed in the regression analysis.
Regarding the synthetic voice, a vast majority of participants (95%) recognized the non-human nature of the narration. Only a minor proportion (20%) indicated that the synthetic quality detracted from their overall enjoyment of the experience.
Results of panel regressions analyses for testing hypotheses about the influence of voice gender (male vs. female) on cognitive effort and enjoyment of AD clips.
5. Discussion
Our results did not fully support the initial hypotheses in their entirety. It was confirmed that female voices reduce self-reported cognitive effort (H1a), measured using the Mental-Effort Rating Scale, and increase positive valence and subjective arousal (H2a), assessed via the Self-Assessment Manikin (SAM). However, at the physiological and narrative immersion level, the impact was limited or non-existent. This finding suggests that while female prosody is perceived as inherently clearer and more pleasant, facilitating mental processing, this advantage does not automatically translate into greater narrative immersion (evaluated using the Transportation Questionnaire) or esthetic enjoyment (measured with the Aesthetic Experience Questionnaire).
This partial result in H2a leads us to the same crucial distinction as in the creative AD study [
5]: pleasantness and immersion are not synonymous with esthetic pleasure or deep narrative engagement. Participants may find the female voice more comforting and easier to listen to (as indicated by the reduction in cognitive effort and positive valence), but this is not sufficient to increase their connection with an abstract and non-linear work of art such as contemporary dance. Although vocal condition influenced valence and arousal ratings, we did not observe corresponding differences in enjoyment or transportation, suggesting that changes in basic affective appraisal did not translate into broader experiential outcomes in this dataset.
The female voice proved to be less mentally demanding (H1a), which can be attributed to the greater clarity and warmth of female prosody, a finding consistent with previous studies suggesting easier processing of female voices in narrative contexts [
10]. Nevertheless, no significant differences were observed in heart rate (H1b), the physiological indicator of cognitive load. This suggests that the physiological measures used may have been insufficient to capture subtle variations in the cognitive load imposed by voice gender, given the significant demands of the complex, multimodal, and abstract nature of the dance.
The results also qualify our expectations regarding H2b. Relative to the male voice condition, the female voice was associated with lower phasic EDA, suggesting reduced sympathetic activation during reception. At the same time, the male voice condition showed comparatively higher phasic EDA, which may reflect greater autonomic engagement while listening. Importantly, phasic EDA and HRV are valence-ambiguous: they can index multiple processes such as orienting, sustained attention, task effort, emotion regulation, or increased cognitive demand, and they cannot be used to label an experience as “positive” or “negative” in isolation. For this reason, we interpret the physiological differences cautiously and in relation to the converging qualitative evidence. In the retrospective interviews, participants more often characterized the male voice as “impersonal” or “taxing,” whereas the female voice was more frequently described as easier to follow and more comfortable. Taken together, this pattern is consistent with the female voice reducing the effort or tension associated with processing abstract dance AD, and with the male voice eliciting higher activation potentially linked to increased attentional or cognitive demands. This interpretation remains tentative and does not equate autonomic activation with negative affect; rather, it treats EDA/HRV as correlates of engagement and processing load that gain meaning only when triangulated with self-report.
Despite limited findings in affective and physiological responses, retrospective interviews revealed a clear preference for the female voice: 64% of participants preferred this voice compared to 36% who opted for the male voice. This discrepancy between objective reception and stated preference is fundamental. It is plausible that this general trend is driven, in part, by familiarity, given that female voices conventionally predominate in various narrative and care contexts and are perceived as the standard or most comfortable option. It may also be related to the gender of the listener, given that, according to retrospective interviews, 10 of the 13 men (77%) and 11 of the 20 female participants (55%) preferred the female voice. This marked preference, justified by participants as more familiar, natural, and relaxing, corroborates the quantitative evidence of reduced cognitive effort (H1a) and the calming physiological effect observed through the reduction in phasic electrodermal activity (EDA). AD users, who already invest significant mental effort in image construction, prioritize the voice that facilitates this task and induces a state of lower negative arousal, even if that voice does not maximize their narrative immersion.
Our results suggest that, in the context of dance AD, factors such as vocal comfort and decreased cognitive effort are more relevant variables for overall user satisfaction than mere emotional activation (arousal). Consequently, to optimize the user experience, a multimethod approach that integrates explicit user feedback, the artistic intent of poetic AD, and empirical evidence is needed to optimize the experience.
6. Conclusions
Our work is pioneering in analyzing the effects of narrator gender in AD applied to a highly abstract art form such as dance, and highlights the need to develop personalized AD strategies that consider not only linguistic content, but also prosodic and vocal characteristics aimed at minimizing cognitive load. The results suggest that, in this context, factors such as vocal comfort and decreased cognitive effort are more relevant variables affecting overall user satisfaction than mere emotional activation.
Notably, our findings also reveal that voice preference is influenced by the gender of the listener. While the female voice emerged as the overall preferred option, largely due to its impact on reducing cognitive load and its strong appeal among male participants, female listeners showed a more nuanced and divided preference. This suggests that a one-size-fits-all approach to voice selection may be insufficient for artistic content, advocating for the implementation of user-selectable AD features where audiences can choose a narrator that aligns with their individual perceptual comfort.
In this sense, the empirical findings have immediate pedagogical and professional ramifications: in the field of AD practice, the evidence strongly supports the selection of female voices for abstract content that is cognitively demanding, favoring a vocal style characterized by low arousal and high clarity, while, at the training level, the results underscore the need to formalize the teaching of vocal performance and prosody, treating the narrator’s voice as a critical variable in the quality of accessibility. Ultimately, by reducing the cognitive barrier to entry through carefully informed vocal choices, this research directly contributes to fostering greater acceptance and attendance of contemporary dance among audiences with visual impairments. Despite the contributions of this study, several limitations should be considered when interpreting our results. The recruitment of participants with visual impairments poses significant logistical and ethical challenges. Recruitment and retention often require tailored strategies, extended timelines, accessibility considerations, and institutional collaboration. Within this context, achieving larger sample sizes or implementing more complex physiological protocols remains particularly demanding.
A particularly relevant methodological limitation is that the emotional valence of the dance stimuli was not thematically analyzed or pre-rated in a time-segmented manner. This is crucial because the inherent affective response to abstract art is dynamic and heterogeneous, meaning the valence can shift even within a single piece. The affective congruence between the dance content at any given segment and the AD delivery may have influenced participants’ segmented responses. Future research should therefore implement granular, segment-by-segment analysis of stimulus valence to fully control for this complex variable.
Another critical methodological consideration is the utilization of synthetic voices. While this choice granted maximal control over linguistic content and ensured consistency across stimuli, it inherently limits the expressive range of the narration. Additionally, all ADs in this study employed a poetic style, which may limit the generalizability of the results to more neutral or informational forms of AD.
Future research should therefore prioritize comparative studies involving both synthetic and natural human voices to explore their respective affordances and implications for audience engagement, particularly in poetic and artistic AD contexts. Moreover, interaction effects could not be tested. Future research should adopt a factorial design to examine whether the interplay between vocal tone and narrative style produces additive or interactive effects.
Author Contributions
Conceptualization, M.L.R., A.M.R.L. and M.R.C.; methodology, M.L.R., A.M.R.L. and M.R.C.; formal analysis, K.R.; investigation, M.L.R., A.M.R.L. and M.R.C.; resources, M.L.R.; data curation, K.R.; writing, M.L.R.; review and editing, A.M.R.L. and M.R.C.; supervision, A.M.R.L. and M.R.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research and the APC were funded by Fundación Séneca (Región de Murcia), through the project “ADance: the emotional and cognitive reception of the audio description of contemporary dance” (grant number 22028/PI/2).
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki, and approved by the Ethics Committee of the University of Murcia (protocol code: 4184/2023, date of approval: 21 February 2023).
Informed Consent Statement
Informed consent was obtained from all participants involved in the study.
Data Availability Statement
The data presented in this study are available on request from the corresponding authors. The data are not publicly available due to ethical and privacy restrictions.
Acknowledgments
The authors would like to thank the University of Murcia and the participants who volunteered for this study. Special thanks are also extended to Aptent for providing the technical support and the synthetic voices used in the experimental stimuli. During the preparation of this manuscript, the authors used AI for the purposes of language editing, stylistic refinement, and translation assistance to ensure academic clarity in English. The authors have reviewed and edited the output and take full responsibility for the content of this publication.
Conflicts of Interest
The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.
Disability Language/Terminology Positionality Statement
In this paper, the authors have primarily adopted person-first language (e.g., “participants with visual impairment”). This choice is grounded in the legal and theoretical framework of the UN Convention on the Rights of Persons with Disabilities and is congruent with the disciplinary context of Media Accessibility and Translation Studies in Spain. This terminology aims to emphasize the individuality and personhood of the participants, which aligns with the ethical and human-centric approach of the ADance project.
Abbreviations
The following abbreviations are used in this manuscript:
| AD | Audio Description |
| SAM | Self-Assessment Manikin |
| TAV | Audiovisual Translation Studies |
| EDA | Electrodermal Activity |
| HRV | Heart Rate Variability |
References
- UNE 153020:2005; Audiodescripción para Personas con Discapacidad Visual. Requisitos para la Audiodescripción y Elaboración de Audioguías. AENOR—Asociación Española de Normalización y Certificación: Madrid, Spain, 2005.
- Udo, J.-P.; Fels, D. “Suit the Action to the Word, the Word to the Action”: An Unconventional Approach to Describing Shakespeare’s Hamlet. J. Vis. Impair. Blind. 2009, 103, 178–183. [Google Scholar] [CrossRef] [Scilit]
- Ramos, M.; Rojo, A. Analysing the AD process: Creativity, expertise and quality. JoSTrans 2020, 34, 212–232. [Google Scholar] [CrossRef] [Scilit]
- Luján Rubio, M.; Rojo López, A.M.; Ramos Caro, M. Audiodescripción de danza: Un análisis del proceso basado en entrevistas con profesionales. MonTI 2025, 17, 430–464. [Google Scholar] [CrossRef] [Scilit]
- Ramos Caro, M.; Luján Rubio, M.; Rudnicki, K.; Rojo López, A.M. Dancing with words: The emotional reception of creative Audio Description in contemporary dance. Pozn. Stud. Contemp. Linguist. 2025, 61, 571–595. [Google Scholar] [CrossRef] [Scilit]
- Walczak, A.; Fryer, L. Creative description: The impact of audio description style on presence in visually impaired audiences. Br. J. Vis. Impair. 2017, 35, 6–17. [Google Scholar] [CrossRef] [Scilit]
- Fryer, L.; Freeman, J. Can you feel what I’m saying? The impact of verbal information on emotion elicitation and presence in people with a visual impairment. In Proceedings of the International Society for Presence Research Annual Conference—ISPR 2014, Vienna, Austria, 17–19 March 2014; pp. 99–107. Available online: https://matthewlombard.com/ISPR/Proceedings/2014/Fryer_Freeman.pdf (accessed on 10 February 2026).
- Cabeza-Cáceres, C. Audiodescripció i Recepció: Efecte de la Velocitat de Narració, L’entonació i L’explicitació en la Comprensió Fílmica. Ph.D. Thesis, Universitat Autònoma de Barcelona, Barcelona, Spain, 2013. [Google Scholar]
- Fresno Cañada, N. Is a picture worth a thousand words? The role of memory in audio description. Across Lang. Cult. 2014, 15, 111–129. [Google Scholar] [CrossRef] [Scilit]
- Iturregui-Gallardo, G.; Matamala, A. Audio subtitling: Dubbing and voice-over effects and their impact on user experience. Perspectives 2021, 29, 64–83. [Google Scholar] [CrossRef] [Scilit]
- Machuca, M.J.; Matamala, A. Neutral voices in audio descriptions. Babel 2022, 68, 668–696. [Google Scholar] [CrossRef] [Scilit]
- Fernández Alonso, I.; Machuca, M.J.; Matamala, A. Voces neutras y alteración tonal. Circ. Linguist. Apl. Comun. 2024, 97, 121–134. [Google Scholar] [CrossRef] [Scilit]
- Iglesias Fernández, E.; Martínez Martínez, S.; Chica Núñez, A.J. Cross-fertilization between Reception Studies in Audio Description and Interpreting Quality Assessment: The Role of the Describer’s Voice. In Audiovisual Translation in a Global Context; Baños Piñero, R., Díaz-Cintas, J., Eds.; Palgrave Macmillan: London, UK, 2015; pp. 72–95. [Google Scholar]
- Belin, P.; Fecteau, S.; Bédard, C. Thinking the voice: Neural correlates of voice perception. Trends Cogn. Sci. 2004, 8, 129–135. [Google Scholar] [CrossRef] [Scilit]
- Machuca Ayuso, M.J.; Ríos, A. La agradabilidad de las voces de los audiodescriptores: Estudio acústico y perceptivo. Bol. Acad. Peru. Lengu. 2022, 71, 215–252. [Google Scholar] [CrossRef] [Scilit]
- Sauter, D.A.; Eisner, F.; Ekman, P.; Scott, S.K. Cross-cultural recognition of basic emotions through nonverbal emotional vocalizations. Proc. Natl. Acad. Sci. USA 2010, 107, 2408–2412. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Scherer, K.R. Vocal communication of emotion: A review of research paradigms. Speech Commun. 2003, 40, 227–256. [Google Scholar] [CrossRef] [Scilit]
- McAleer, P.; Todorov, A.; Belin, P. How Do You Say ‘Hello’? Personality Impressions from Brief Novel Voices. PLoS ONE 2014, 9, e90779. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Green, M.C.; Brock, T.C. The role of transportation in the persuasiveness of public narratives. J. Personal. Soc. Psychol. 2000, 79, 701–721. [Google Scholar] [CrossRef]
- Fernández Huerta, J. Medidas sencillas de lecturabilidad. Consigna 1959, 214, 29–32. [Google Scholar]
- Paas, F. Training strategies for attaining transfer of problem-solving skill in statistics: A cognitive-load approach. J. Educ. Psychol. 1992, 84, 429–434. [Google Scholar] [CrossRef]
- Iturregui-Gallardo, G.; Méndez-Ulrich, J.L. Towards the Creation of a Tactile Version of the Self-Assessment Manikin (T-SAM) for the Emotional Assessment of Visually Impaired People. Int. J. Disabil. Dev. Educ. 2020, 67, 657–674. [Google Scholar] [CrossRef] [Scilit]
- Barnés Castaño, C. Bases Cognitivas del Acceso al Conocimiento en Audiodescripción Museística: Una Aproximación Experimental; Servicio de Publicaciones Universidad de Granada: Granada, Spain, 2024. [Google Scholar]
- Ravaja, N.; Kallinen, K. Emotional effects of startling background music during reading news reports: The moderating influence of dispositional BIS and BAS sensitivities. Scand. J. Psychol. 2004, 45, 231–238. [Google Scholar] [CrossRef] [Scilit]
- Mahanama, B.; Jayawardana, Y.; Rengarajan, S.; Jayawardena, G.; Chukoskie, L.; Snider, J.; Jayarathna, S. Eye Movement and Pupil Measures: A Review. Front. Comput. Sci. 2022, 3, 733531. [Google Scholar] [CrossRef] [Scilit]
- Pham, T.; Lau, Z.J.; Chen, S.H.A.; Makowski, D. Heart Rate Variability in Psychology: A Review of HRV Indices and an Analysis Tutorial. Sensors 2021, 21, 3998. [Google Scholar] [CrossRef] [Scilit]
- Geisler, F.C.M.; Vennewald, N.; Kubiak, T.; Weber, H. The impact of heart rate variability on subjective well-being is mediated by emotion regulation. Personal. Individ. Differ. 2010, 49, 723–728. [Google Scholar] [CrossRef] [Scilit]
- Mather, M.; Thayer, J.F. How heart rate variability affects emotion regulation brain networks. Curr. Opin. Behav. Sci. 2018, 19, 98–104. [Google Scholar] [CrossRef] [Scilit]
- Min, J.; Koenig, J.; Nashiro, K.; Yoo, H.J.; Cho, C.; Thayer, J.F.; Mather, M. Resting heart rate variability is associated with neural adaptation when repeatedly exposed to emotional stimuli. Neuropsychologia 2024, 196, 108819. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Richards, J.E.; Casey, B.J. Heart Rate Variability During Attention Phases in Young Infants. Psychophysiology 1991, 28, 43–53. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Thorson, E.; Lang, A. The Effects of Television Videographics and Lecture Familiarity on Adult Cardiac Orienting Responses and Memory. Commun. Res. 1992, 19, 346–369. [Google Scholar] [CrossRef] [Scilit]
- Bente, G.; Kryston, K.; Jahn, N.T.; Schmälzle, R. Building blocks of suspense: Subjective and physiological effects of narrative content and film music. Humanit. Soc. Sci. Commun. 2022, 9, 449. [Google Scholar] [CrossRef] [Scilit]
- Fischer, L.M.; Cummins, R.G.; Gilliam, K.C.; Baker, M.; Burris, S.; Irlbeck, E. Examining the Critical Moments in Information Processing of Water Conservation Videos within Young Farmers and Ranchers: A Psychophysiological Analysis. J. Agric. Educ. 2018, 59, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Keene, J.R.; Rasmussen, E.E.; Berke, C.K.; Densley, R.L.; Loof, T.; Adams, R.B.; Humma, G.H.; Marshall, A. The effect of plot explicit, educational explicit, and implicit inference information and coviewing on children’s internal and external cognitive processing. J. Appl. Commun. Res. 2019, 47, 153–174. [Google Scholar] [CrossRef] [Scilit]
- Madsen, J.; Parra, L.C. Cognitive processing of a common stimulus synchronizes brains, hearts, and eyes. Proc. Natl. Acad. Sci. USA Nexus 2022, 1, pgac020. [Google Scholar] [CrossRef] [Scilit]
- Hartnett, N.; Bellman, S.; Beal, V.; Kennedy, R.; Charron, C.; Varan, D. How to accurately measure attention to video advertising. Int. J. Advert. 2025, 44, 184–207. [Google Scholar] [CrossRef] [Scilit]
- Park, B.; Bailey, R.L. Application of Information Introduced to Dynamic Message Processing and Enjoyment. J. Media Psychol. 2018, 30, 196–206. [Google Scholar] [CrossRef] [Scilit]
- Benedek, M.; Kaernbach, C. A continuous measure of phasic electrodermal activity. J. Neurosci. Methods 2010, 190, 80–91. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Niskanen, J.-P.; Tarvainen, M.P.; Ranta-Aho, P.O.; Karjalainen, P.A. Software for advanced HRV analysis. Comput. Methods Programs Biomed. 2004, 76, 73–81. [Google Scholar] [CrossRef] [PubMed]
Table 1.
Poetic resources: some examples used in the ADs.
Table 1.
Poetic resources: some examples used in the ADs.
| Parameters | Example |
|---|
| Simile | Gira, cabizbaja, como si fuera la bailarina de una caja de música. [She turns, head bowed, as if she were a ballerina in a music box.] |
| Emotional state | Carlos la rodea, pasa por delante y se va hacia atrás, solitario y pausado. [Carlos circles her, walks past her and goes to the back, alone and unhurried.] |
| Poetic language | Sus ojos siguen sus movimientos con una intensidad que parece traspasar el espacio. [Her eyes follow his movements with an intensity that seemsto pierce through space.] |
| Repetition | Deja caer los brazos. Los sube. Caen. Suben. Caen. Suben. [He drops her arms. She raises them. They fall. They rise. They fall. They rise. They fall. They rise.] |
Table 2.
Number of poetic resources and readability parameters.
Table 2.
Number of poetic resources and readability parameters.
| Parameters | A1 Male | A2 Female | B1 Male | B2 Female |
|---|
| Simile | 13 | 14 | 13 | 8 |
| Emotional state | 6 | 5 | 4 | 3 |
| Poetic language | 10 | 14 | 12 | 13 |
| Repetition | 7 | 8 | 8 | 10 |
| Total words | 750 | 725 | 721 | 719 |
| Readability analysis | 93.04 (very easy) | 94.26 (very easy) | 86.87 (very easy) | 88.79 (very easy) |
Table 3.
Stimuli presentation.
Table 3.
Stimuli presentation.
| | Counterbalanced Presentation Order of AD Versions Across Participants |
|---|
| P1 | Male A1 | Male B1 | Female B2 | Female A2 |
| P2 | Female B2 | Female A2 | Male A1 | Male B1 |
| P3 | Male A1 | Male B1 | Female A2 | Female B2 |
| P4 | Female A2 | Female B2 | Male B1 | Male A1 |
| Etc. | Male B1 | Male A1 | Female A2 | Female B2 |
Table 4.
Description and values of the study variables (adapted from Ref. [
5]).
Table 4.
Description and values of the study variables (adapted from Ref. [
5]).
| Variable Name | Description | Values |
|---|
| Age | Age of the participant | Years |
| Sex | Sex of the participant | 0—male, 1—female |
| Impairment | Degree of visual impairment of the participant | 0—partially sighted (10% vision), 1—total sight loss |
| Time | The position of a given trial (clip) in the order of the experiment. This variable approximates the fatigue of participants as time passes. | 1–6 |
| Voice Gender | The gender of the voice used in the audio description of the clip | 0—male, 1—female |
| Voice Gender | The gender of the voice used in the audio description of the clip | 0—male, 1—female |
| SAM Arousal | Self-reported arousal measured with the Self-Assessment Manikin [10] | 1–9 |
| SAM Valence | Self-reported emotional valence measured with the Self-Assessment Manikin [22] | 1–9 |
| Transportation | Self-reported feelings of transportation measured with the narrative transportation questionnaire by Green and Brock [19] | 1–7 |
| Enjoyment | Self-reported feelings of enjoyment measured with the self-reported questionnaire by Barnés [23] | 1–7 |
| Cognitive Effort | Self-reported cognitive effort measured with the mental effort rating scale by Paas [21] | 1–9 |
| Engagement | The number of accurate responses in the ad hoc recall task | 0–4 |
| Phasic EDA | Phasic EDA measured with the time integral of the phasic driver during the AD clips according to Benedek & Karnbach [38] | muS*s |
| Tonic EDA | Average tonic EDA level during the AD clips according to Benedek & Karnbach [38] | muS |
| HR | Heart rate during the AD clips | Bpm |
| HRV | High-frequency heart rate variability (0.15–0.4 Hz)measured obtained with a Fourier transform of the IBI wave during the AD clip. | ms2 |
| Baseline Phasic EDA | Phasic EDA during the baseline measurement that preceded the relevant AD clip. This control variable allows to account for individual differences in general levels of EDA. | muS*s |
| Baseline Tonic EDA | Tonic EDA during the baseline measurement that preceded the relevant AD clip. This control variable allows to account for individual differences in general levels of EDA. | muS |
| Baseline HR | Heart rate during the baseline measurement that preceded the relevant AD clip. This control variable allows to account for individual differences in general levels of heart rate. | Bpm |
| Baseline HRV | Heart rate variability during the baseline measurement that preceded the relevant AD clip. This control variable allows to account for individual differences in general levels of HRV. | ms2 |
Table 5.
Study protocol.
| 1 | Registration and informed consent |
| 2 | Shimmer placement |
| 3 | Start recording: relaxing task (music 5 min) |
| 4 | Stop recording: SAM questionnaire |
| 5 | Start recording: 10 min dance (male or female) |
| 6 | Stop recording: questionnaires |
| 7 | Start recording: 10 min dance (male or female) |
| 8 | Stop recording: questionnaires |
Table 6.
Descriptive statistics of the study variables.
Table 6.
Descriptive statistics of the study variables.
| Dependent Variable | Male | Female |
|---|
| SAMArousal | 6.58 ± 0.25 | 7 ± 0.24 |
| SAMValence | 6.48 ± 0.28 | 6.93 ± 0.22 |
| Transportation | 44.18 ± 1.03 | 44.30 ± 1.05 |
| Enjoyment | 55.15 ± 1.04 | 55.15 ± 0.99 |
| Cognitive Effort | 4.69 ± 0.29 | 3.91 ± 0.25 |
| Phasic EDA | 15.54 ± 2.18 | 11.26 ± 2.04 |
| Tonic EDA | 2.39 ± 22 | 2.24 ± 22 |
| Heart rate | 77.98 ± 1.71 | 77.91 ± 1.73 |
| HRV | 562.66 ± 148.32 | 453.56 ± 88.37 |
Table 7.
Correlations between the study variables.
Table 7.
Correlations between the study variables.
| | | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
|---|
| 1 | SAMArousal | - | | | | | | | | |
| 2 | SAMValence | 0.83 * | - | | | | | | | |
| 3 | Transportation | 0.46 * | 0.43 * | - | | | | | | |
| 4 | Enjoyment | 0.53 * | 0.57 * | 0.50 * | - | | | | | |
| 5 | Cognitive Effort | −0.11 | −0.11 | −0.01 | −0.11 | - | | | | |
| 6 | Phasic EDA | −0.18 * | −0.15 * | −0.27 * | −0.21 * | 0.10 | - | | | |
| 7 | Tonic EDA | −0.07 | −0.00 | −0.21 * | −0.20 * | 0.15 * | 0.67 * | - | | |
| 8 | Heart Rate | 0.15 * | 0.10 | 0.31 * | 0.03 | −0.02 | 0.13 | 0.13 | - | |
| 9 | HRV | −0.13 * | −0.14 * | −0.15 * | −0.08 | 0.03 | −0.07 | −0.19 * | −0.58 * | - |
Table 8.
Results of panel regression analyses for testing hypotheses.
Table 8.
Results of panel regression analyses for testing hypotheses.
Dependent: Independent: | H1a Cognitive Effort Self-Report | H1b Cognitive Effort Heart Rate | H2a SAMValence | H2a SAMArousal | H2a Enjoyment | H2a Transportation | H2b Phasic EDA | H2b Tonic EDA | H2b HRV | H3 Engagement (Recall) |
|---|
| Age | 0.01 (0.02) | −0.14 (0.09) | 0.01 (0.02) | −0.00 (0.01) | 0.03 (0.08) | 0.01 (0.08) | 0.07 (0.10) | 0.00 (0.01) | 6.22 (2.56) * | −0.01 (0.01) |
| Sex | 0.70 (0.74) | 3.54 (2.78) | 0.15 (0.73) | 0.01 (0.69) | 1.89 (2.88) | −0.48 (2.89) | −0.67 (4.11) | −0.31 (0.55) | −287.3 (105.2) ** | −0.13 (0.22) |
| Impairment | 1.07 (0.84) | 2.81 (3.30) | −0.84 (0.83) | −10.20 (0.78) | −3.15 (3.27) | −3.61 (3.28) | −2.14 (5.26) | −0.81 (0.66) | −31.44 (143.1) | −0.04 (0.19) |
| Time | −0.09 (0.15) | −1.06 (0.40) ** | 0.02 (0.13) | 0.19 (0.13) | 0.51 (0.56) | −0.62 (0.61) | −0.88 (1.0) | 0.06 (0.12) | 68.59 (61.66) * | −0.19 (0.06) |
| Voice Gender | −0.80 (0.20) ** | −0.19 (0.45) | 0.46 (0.15) ** | 0.45 (0.15) ** | 0.08 (0.70) | 0.03 (0.80) | −3.84 (1.57) * | 0.03 (0.17) | 29.81 (0.09) | −0.13 (0.09) |
| Baseline HR | | 0.70 (0.18) ** | | | | | | | | |
| Baseline Phasic EDA | | | | | | | 1.02 (0.28) ** | | | |
| Baseline Tonic EDA | | | | | | | | 0.79 (0.20) ** | | |
| Baseline HRV | | | | | | | | | 0.25 (0.09) ** | |
| N | 132 | 61 | 132 | 132 | 132 | 132 | 62 | 62 | 61 | 131 |
| Wald Chi | 19.13 ** | 67.89 ** | 11.17 * | 11.57 * | 2.74 | 2.56 | 25.82 ** | 25.71 ** | 30.11 ** | 13.96 * |
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |