1. Introduction
In recent years, with the development of intelligent vehicles and smart cockpits, the Head-up Display (HUD) has been increasingly adopted in automobiles and is regarded as a key component of future automotive human–machine interface (HMI) [
1,
2]. Automotive HUDs can present critical driving information such as vehicle speed and navigation cues within the driver’s immediate forward field of view [
3], thereby reducing head-down glances during information acquisition; however, achieving reliable legibility requires careful visual design (e.g., luminance/contrast and text format) and remains a key technical challenge in vehicle-mounted AR-HUD commercialization [
4]. Such HUD designs may facilitate information acquisition and potentially support driving safety and user experience under appropriate task and environmental conditions. Among the various types of information displayed on HUDs, navigation—as a high-frequency function that directly influences driving decisions—its information presentation mode plays a crucial role in drivers’ attention allocation and operational responses.
It is worth noting that HUDs do not necessarily lead to safer navigation task performance. Drivers’ cognitive resources are limited; if a HUD presents excessive information or adopts an unreasonable organization mode, it may occupy additional cognitive resources and induce attention distraction [
5]. Navigation and other content displayed on HUDs may also interfere with the primary driving task, resulting in visual and cognitive distraction [
6], and inappropriate interface design may even have adverse effects on driving safety. Relevant studies have also indicated that while HUD information can facilitate hazard prompts and responses, as the visual complexity of displayed content increases, drivers’ cognitive load may rise accordingly, thereby impairing information processing efficiency [
7]. In navigation scenarios, early research suggested that HUDs used to provide concise route guidance do not significantly distract drivers or reduce navigation task performance [
8], but other studies have found that while HUD navigation improves certain performance indicators, it may induce specific negative behaviors [
9]. These inconsistent findings indicate that the presentation mode of HUD navigation information still needs to be optimized through more refined design and stricter empirical evaluation. However, beyond these empirical inconsistencies, there is still a lack of a clear theoretical explanation regarding how specific interface design factors shape drivers’ cognitive load and visual attention allocation during navigation tasks.
From the perspective of cognitive mechanisms, driving is a typical multi-task information processing process, in which drivers are required to continuously complete links such as environmental perception, information comprehension, decision-making, and control execution in a dynamic traffic environment. When external information load or task demands exceed the capacity of working memory, the cognitive system will experience overload, manifesting as attention distraction, slowed reaction, and decreased control ability [
10]. Therefore, the key to the design of HUD navigation interfaces lies not in increasing the amount of information, but in reducing unnecessary extraneous cognitive load through appropriate presentation methods, enabling drivers to quickly and accurately extract navigation cues while maintaining attention to the road environment. From this perspective, HUD interface design can be understood as a process of optimizing the distribution of visual attention and managing cognitive load within the driver’s limited processing capacity.
Existing studies suggest that the display position of HUD navigation information can influence drivers’ cognitive processing [
8], and presenting information in the upper versus lower visual field may yield different effects. A lower-field display typically overlays navigation symbols near the roadway ahead of the driver and is considered beneficial for reducing gaze shifts and maintaining continuous attention to the road surface [
11]. However, virtual elements may occlude or overlap with real-world targets (e.g., vehicles and pedestrians), resulting in visual clutter and additional cognitive demands [
12]. The upper display places navigation information in the sky or the upper area of the visual field, which can avoid occluding road targets to a certain extent [
13] and may reduce interference caused by background complexity. Nevertheless, prompts continuously located in the peripheral visual field may also weaken the immediate perceptibility of information or induce attention distraction [
14]. In addition, different drivers have inconsistent preferences for the location of HUD information presentation. Existing studies have indicated that drivers often hope HUD information is positioned outside the foveal visual field to facilitate quick scanning and acquisition [
15]. Overall, there is still a lack of quantitative comparative evidence based on subjective and objective indicators in controlled driving tasks regarding which display position—upper or lower—is more conducive to reducing cognitive load and improving navigation efficiency. More importantly, existing studies rarely examine how spatial layout interacts with drivers’ attentional strategies during real-time navigation tasks.
Regarding navigation text on HUDs, prior studies indicate that the format of textual presentation can affect attentional resource allocation and performance in visual search tasks [
16]. When graphic cues are sufficient to convey meaning, adding text may introduce redundancy, unnecessarily consuming attentional resources, increasing cognitive load, and impairing task performance [
17,
18]. In driving contexts, dense textual content may further draw drivers’ attention toward the display and away from monitoring critical road cues [
19,
20]. By contrast, when graphic symbols are semantically ambiguous or insufficiently familiar, brief supplemental text can improve comprehension accuracy and shorten response time, thereby enhancing information extraction efficiency and reducing interpretive ambiguity [
21]. Since most existing conclusions are derived from different tasks and contexts, and systematic comparative studies targeting HUD navigation scenarios remain limited, there is still a need to further verify whether navigation text information should be presented in HUDs and whether text information produces different cognitive effects under different display positions. Thus, how textual cues interact with spatial display location to influence drivers’ cognitive load and visual attention allocation remains an open question in HUD interface research.
Based on the aforementioned inconsistencies in existing research, this study focuses on two key variables in the design of automotive HUD navigation information, exploring the mechanisms through which different navigation information display positions (upper vs. lower visual field) and navigation text presentation modes (text present vs. absent) influence drivers’ cognitive load and navigation performance. The study will verify these relationships through simulated driving experiments combined with eye-tracking technology, simultaneously collecting task performance data, subjective cognitive load ratings, and eye-tracking metrics in controlled driving tasks to achieve a multi-dimensional characterization of cognitive load and information processing processes. Compared with existing studies that mostly examine a single interface factor in isolation, this study jointly incorporates display position and text information into a unified experimental framework. By integrating behavioral performance, subjective workload, and eye-movement evidence, this research aims to provide a more comprehensive understanding of how HUD interface design influences drivers’ cognitive load regulation and visual attention allocation during navigation tasks. It aims to provide more targeted empirical evidence, offer data support for the optimization of HUD navigation information presentation modes, and deliver an actionable research basis for enhancing driving safety and interactive experience.
Guided by the aforementioned theories and research progress, this study formulates testable predictions regarding the effects of the key variables and proposes the following hypotheses:
H1. Different HUD navigation information display positions have a significant impact on drivers’ cognitive load and navigation efficiency.
H2. Different HUD navigation text information presentation modes have a significant impact on drivers’ cognitive load and navigation efficiency.
H3. There is a significant interaction effect between HUD navigation information display position and text information presentation mode, i.e., their combinations will exert different effects on drivers’ cognitive load and navigation efficiency.
3. Experimental Procedures
3.1. Practice Session
To enable participants to quickly familiarize themselves with the simulated driving system and HUD interface operations, a 5 min practice driving session was arranged prior to the formal experiment. The practice scenario was similar to that of the formal experiment but with a different route, so as to avoid participants’ performance in the subsequent experiment being affected by route memorization. The HUD navigation prompts presented during the practice session covered all combinations of experimental conditions (including upper/lower display positions and with/Without text information), ensuring that participants understood all types of prompt formats and task requirements. After the practice session, the experimenter confirmed that participants had no further questions before initiating the formal experiment.
3.2. Calibration and Instructions
Before the start of the formal experiment, the experimenter first explained the task requirements and precautions to participants in detail, and displayed brief written instructions on the screen. For example: “Please drive in accordance with the HUD navigation prompts displayed on the windshield. When a turn instruction appears, please verbally report the turning direction as soon as possible.” The experimenter also emphasized the above requirements verbally. The verbal-report task was adopted as a minimally intrusive response method to isolate perceptual-cognitive processing without altering steering behavior. This design allowed us to isolate the perceptual and attentional effects of HUD variables while maintaining continuous driving as the primary task.
After participants clearly understood the task, eye tracker calibration was conducted. During the calibration process, participants were required to maintain a proper sitting posture and keep their heads stable, and fixate on the calibration points that appeared sequentially on the screen to complete the standard nine-point calibration, so as to ensure the accurate collection of eye movement data.
3.3. Formal Experimental Procedure
Participants initiated the simulated driving task and drove from the starting point to the destination in accordance with the navigation guidance provided by the HUD. During driving, participants were required to abide by virtual traffic rules (e.g., stopping at red lights) and avoid vehicles and pedestrians, so as to simulate the scenario of real urban road driving. Each driving task was set in a single-lane urban road scenario and lasted approximately 2–4 min.
Multiple navigation turn prompts appeared during the driving process (with 6 fixed turn instructions in each driving task). Whenever a turn prompt with arrow symbols (with or without accompanying text) appeared on the HUD displayed on the windshield, participants were required to verbally report the turning direction as quickly and accurately as possible (e.g., responding promptly with “turn left” or “turn right”). The experimenter measured the participants’ reaction time through frame-by-frame analysis of synchronized video and audio recordings and judged the accuracy of their verbal reports accordingly.
Participants continued driving until they reached the predetermined destination, thus completing the route task. Upon the completion of each driving task, the system automatically recorded the total driving time. Immediately after the task, participants filled in the subjective cognitive load scale (PAAS) to report their perceived subjective cognitive load during the task. Meanwhile, the objective behavioral data and eye movement data generated during the experiment were synchronously recorded for subsequent analysis. Each participant was required to complete driving tasks under all four HUD interface conditions. After finishing the task under one condition, participants could take an appropriate break to prevent fatigue from affecting their performance in subsequent tasks.
After the entire experiment was completed, the collected data were subjected to statistical analysis. IBM SPSS Statistics 26.0 software was used to conduct a repeated measures analysis of variance (ANOVA) on each indicator, so as to examine whether the main effects and interaction effects of the two independent variables (navigation information display position and text information presence) were significant. The significance level was set at α = 0.05. Assumptions of normality and sphericity were examined prior to analysis, and effect sizes (partial η2) were reported to indicate the magnitude of observed effects. Normality was assessed using the Shapiro–Wilk test, and the sphericity assumption was examined using Mauchly’s test. When the sphericity assumption was violated, the Greenhouse–Geisser correction was applied.
5. Discussion
5.1. Display Position Mode
5.1.1. Effects on Task Performance
The results of average reaction time revealed a significant main effect of display position, with the average reaction time for upper-field display being significantly longer than that for lower-field display. This indicates that lower-field display can shorten the reaction time to navigation information during driving, suggesting more efficient information acquisition and processing. Existing studies have suggested that drivers’ gaze is mostly concentrated on the road area; thus, overlaying HUD information near the road makes it easier to capture, which may also reduce mental stress and the cost of information acquisition [
31,
32]. In contrast, upper-field display overlays information on the sky area, while traditional Head-up Displays (HUDs) are located inside the cockpit—both of which may increase the cost of gaze deviation and gaze switching to a certain extent.
In terms of accuracy rate and total driving time, there were no significant main effects of display position. This suggests that different display positions do not significantly alter participants’ judgment accuracy of turning instructions, and tasks can be completed with high accuracy across all conditions. This result may be related to the characteristics of the experimental task. For instance, the direction judgment task in this study did not set strict time limits, making the accuracy rate insensitive to differences between conditions. Meanwhile, total driving time is a relatively macroscopic indicator, which may not be sufficient to reflect the subtle effects caused by differences in display position.
5.1.2. Subjective Scale Scores
Cognitive load scores can reflect participants’ subjective cognitive perceptions during task completion. Combining subjective and objective indicators helps to understand and analyze the results of other metrics. The results showed that the cognitive load score for upper-field display was significantly higher than that for lower-field display. This indicates that subjectively, participants perceived that the lower-field display mode consumed fewer cognitive resources and induced lower cognitive load during task performance. It should be noted that PAAS primarily reflects drivers’ overall perceived mental effort and task difficulty; therefore, the observed differences mainly represent global workload perception rather than detailed multidimensional workload components.
5.1.3. Effects on Eye Movement Metrics
The Mean Fixation Duration was significantly longer in the lower-field display than in the upper-field display. Fixation duration can index both the ease of information processing and the attentional allocation strategy shaped by interface design. When considered alongside the shorter reaction times and lower PAAS scores observed for the lower-field display in this study, the longer fixations are more plausibly interpreted as reflecting increased willingness to verify the navigation cue because it is positioned closer to the driving line of sight, rather than increased processing difficulty. At the same time, this difference may partly reflect drivers’ natural tendency to maintain gaze near the forward roadway region, since the lower-field HUD spatially aligns more closely with habitual road-centered gaze patterns. Prior work has noted that cognitive load and perceptual load may influence fixation duration in opposite directions, such that fixation duration should not be treated as a direct proxy for workload magnitude [
24]. Accordingly, fixation duration here is interpreted primarily as an indicator of attention allocation rather than a linear measure of cognitive load intensity.
Both fixation count and fixation count proportion were significantly higher in the lower-field display condition, indicating more frequent HUD checking behavior and suggesting that lower-field cues are more likely to fall within the driver’s attentional window during driving. This pattern is consistent with the tendency for drivers’ gaze to be concentrated near the roadway region and converges with the finding of shorter reaction times under the lower-field display. Fixation count and fixation count proportion can thus provide complementary evidence for drivers’ attention allocation to the information region and task intent [
33].
No significant difference was found in Mean Pupil Diameter between the two display positions. Although pupil size is widely used as an index of processing effort, it is also susceptible to external factors such as luminance; consequently, differences may not emerge under tasks that impose only moderate workload levels [
34]. In the present study, the navigation prompts were intentionally concise and comparable to most commercially available HUD navigation designs, providing no global route overview and only delivering guidance at specific moments [
35]. Moreover, both display positions were configured to keep the cues near the forward field of view, which likely reduced processing demands and may have attenuated position-related effects on pupil diameter.
Combining the results of task performance, subjective scores, and eye movement metrics, lower-field display demonstrated higher response efficiency, lower subjective load, and was accompanied by more frequent HUD-related fixation behaviors in the context of this study. It should be noted that these findings primarily reflect improvements in navigation information processing efficiency rather than direct evidence of enhanced driving safety, since safety-related navigation task performance indicators such as lane keeping or driving errors were not included in the present study. Meanwhile, the potential advantage of upper-field display lies in less occlusion; however, it also carries the risk of failing to capture attention promptly or incurring the cost of attentional allocation. Existing studies have also shown inconsistencies regarding the advantages and disadvantages of different display positions. For instance, some studies have indicated that overlaying navigation paths on the road may overlap with real objects such as vehicles and pedestrians, resulting in visual clutter and increased cognitive load, whereas displaying navigation information in the sky area is more reasonable and may facilitate turning maneuvers [
36]. Meanwhile, other studies have pointed out that continuously highlighting paths in the peripheral visual field may divert drivers’ attention [
35]. Therefore, against the backdrop of potential increases in HUD information volume in future applications, display position layout still needs to strike a balance between information accessibility and the risk of visual interference [
37].
5.2. Text Display Mode
Taken together, task performance (accuracy, average reaction time, and total driving time) and subjective ratings showed no significant main effect of text presentation. However, this absence of statistical significance should be interpreted cautiously. The relatively moderate task complexity, short driving duration, and sample size may have limited the sensitivity to detect subtle text-related effects. This suggests that, under the controlled task complexity and information amount in the present experiment, the presence versus absence of text did not measurably change participants’ overall perceived workload or their macro-level performance. Notably, the mean PAAS scores across conditions were approximately 5 on the 9-point scale, corresponding to a moderate level of perceived cognitive load. This indicates that the task imposed a manageable but non-trivial demand, which may have constrained the observable impact of textual manipulation on global performance indicators.
By contrast, the eye-tracking metrics indicated that Mean Fixation Duration, Fixation Count, and Fixation Count Proportion were all significantly higher when text was present than when it was absent. Importantly, fixation-based measures are not determined solely by processing difficulty; they are also shaped by information-search and retrieval strategies [
24]. In this study, textual elements provided additional details (e.g., road names and distance-to-turn cues), requiring participants to read and map the text onto the driving context. The resulting increase in fixation duration and more frequent glances toward the HUD therefore likely reflects a shift toward a real-time visual checking (or confirmation) strategy—that is, a task-driven information-sampling process in which drivers repeatedly consult the display to confirm and update their situational understanding—rather than a clear increase in subjectively experienced workload [
38].
Nevertheless, adding text inevitably increases the amount and complexity of information that must be processed while driving. Higher text density can elevate the risk of visual clutter and distraction, particularly in time-varying driving tasks where additional reading demands may compete with the attentional resources needed for road monitoring [
38]. When a HUD presents excessive information—especially in text form—it may also induce cognitive capture [
20], whereby drivers’ responses to external hazards are delayed or attenuated during interface processing, potentially compromising safety. Consistent with this account, Recarte et al. reported that engaging in cognitively demanding activities such as text processing or complex conversations can reduce event detection performance and decrease attention to the road center [
39]. More broadly, different forms of in-vehicle information interaction have been shown to alter drivers’ behavior and attention allocation, thereby increasing distraction risk [
40], which aligns with the gaze-pattern differences observed here.
Overall, the present findings indicate that text presence may have limited impact on global performance and subjective workload under relatively constrained navigation demands, yet it systematically changes drivers’ gaze behavior. It remains plausible that under higher workload conditions—such as more complex traffic environments, higher traffic density, multi-lane navigation tasks, denser text presentation, or longer exposure durations—the additional processing demands introduced by text could translate into measurable increases in subjective workload and observable performance decrements. It should be noted that the textual manipulation reflected realistic HUD navigation prompts with relatively limited information density; more substantial text variations or prolonged driving exposure may yield stronger workload effects.
5.3. Analysis of the Interaction Effect Between Display Position and Text Presentation
A significant interaction between display position and text presentation was observed for Mean Fixation Duration. With respect to display position, the lower-field display yielded significantly longer Mean Fixation Duration than the upper-field display regardless of whether text was present, consistent with the main-effect pattern. This suggests that navigation cues placed closer to the roadway region are more likely to be incorporated into drivers’ visual attention and to receive sustained verification.
The interaction further clarified how text influenced fixation behavior across positions. For the upper-field display, Mean Fixation Duration did not differ significantly between the text-present and text-absent conditions. Given that upper-field cues elicited fewer fixations overall and were less likely to attract attention, it is plausible that participants engaged with the upper-field navigation information only intermittently; consequently, adding text did not produce a measurable change in fixation duration at this position. In contrast, for the lower-field display, Mean Fixation Duration was significantly longer when text was present than when it was absent. A likely explanation is that text introduces additional informational elements (e.g., road name and distance), increasing the amount and complexity of content that must be read and integrated, thereby prolonging gaze dwell time on the HUD and encouraging more thorough on-the-fly checking.
These findings have practical implications for HUD navigation design. Prior research has suggested that excessive textual content on HUDs can reduce drivers’ attention allocation to the route and surrounding traffic environment, thereby undermining situational awareness and driving safety [
41]. Overly dense text may also interfere with the acquisition of critical road cues and increase potential risk [
42]. Accordingly, when navigation prompts are placed in locations that are likely to receive sustained attention (e.g., the lower-field display), particular care should be taken in controlling the amount, length, and presentation of text to avoid unnecessary reading demands and attentional capture. When text is required, concise wording, clear visual hierarchy, and rapidly legible layout should be prioritized to preserve informational value while minimizing interference with the primary driving task.
6. Conclusions
Based on a simulated driving platform, this study systematically examined how HUD navigation information display position and navigation text presentation influence navigation task performance and workload-related measures. The results indicate that display position exerts a more direct and robust effect on both information acquisition efficiency and subjective experience. Specifically, compared with the upper-field display, the lower-field display significantly shortened reaction time to navigation prompts and reduced PAAS scores, suggesting that placing cues closer to the roadway region facilitates rapid information capture and decision-making while supporting sustained road monitoring.
In contrast, under the controlled task difficulty and information amount used in this study, text presence did not produce significant differences in macro-level task performance or subjective workload. However, text presentation showed a consistent modulation of fixation behavior within the HUD AOI and interacted with display position. This pattern implies that the role of text is more likely expressed through online visual checking and verification processes (e.g., changes in dwell and re-checking strategies) rather than necessarily translating into measurable differences in overall performance or subjective ratings. Overall, the lower-field display was associated with shorter reaction times and lower perceived workload in the present task, suggesting relatively more efficient navigation information acquisition under the experimental conditions. Moreover, within the lower-field display, adopting text-free graphical prompts helped reduce unnecessary visual dwell time on the HUD and potential interference, thereby maintaining high navigation efficiency while improving interface simplicity and controllability.
Despite the contributions of this study, several limitations should be acknowledged. First, the participant sample consisted only of young university students with driving experience, which may limit the generalizability of the findings to other driver populations, particularly older or highly experienced drivers. Second, the experiment was conducted in a fixed-base driving simulator with predefined scenarios, which may limit ecological validity compared with real-world driving conditions. The relatively short and repeated driving tasks may also have introduced potential learning effects and may not fully capture cognitive load and fatigue changes that occur during prolonged driving. Third, this study examined only two levels of display position and text presentation, while other interface parameters (e.g., layout geometry, typography, and visual density) were held constant.
On the basis of these findings, we recommend prioritizing the placement of critical HUD navigation information in regions closer to the driver’s roadway gaze channel, while carefully assessing the necessity and limiting the amount of text to reduce distraction risk and optimize usability. Future work should validate these conclusions under more complex road scenarios, higher information density, and more representative driver populations, and further evaluate optimal layout strategies under dynamic traffic conditions and richer cue designs (e.g., multi-level prompts and three-dimensional markers) using higher-fidelity simulations or real-world driving tests.