Next Article in Journal
Cognitive Mechanisms of Referential Ambiguity Resolution in L2 Russian by Chinese Learners: Evidence from Eye-Tracking
Previous Article in Journal
Lexical Elaboration and Intentional Vocabulary Acquisition: An Eye-Tracking Study
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Characterizing Visual Neurosurgical Expertise in Brain MRI Visualization Using Eye-Tracking and 3D Fractal Dimension Analysis

Computational NeuroSurgery (CNS) Lab, Macquarie Medical School, Faculty of Medicine, Health and Human Sciences, Macquarie University, 75 Talavera Road, Sydney, NSW 2109, Australia
*
Authors to whom correspondence should be addressed.
J. Eye Mov. Res. 2026, 19(3), 62; https://doi.org/10.3390/jemr19030062
Submission received: 27 March 2026 / Revised: 14 May 2026 / Accepted: 26 May 2026 / Published: 2 June 2026

Abstract

Eye-tracking has been utilized to characterize visual behavior in medical image visualization and interpretation, yet neurosurgeons remain underrepresented. Characterizing neurosurgery-specific visual expertise is important for understanding expert search strategies, informing training, and developing computational models. This study examined gaze behavior in naïve observers (Np = 29), neurosurgery registrars (Np = 16), and consultant neurosurgeons (Np = 24), viewing normal (Np = 20) and pathological (Np = 19) brain MR images under a free-viewing paradigm. To capture expertise-related characteristics, we analyzed two features at each fixation location: (i) fixation duration, reflecting temporal allocation of visual attention, and (ii) three-dimensional fractal dimension (3DFD) around each fixation location, quantifying local structural complexity. To assess pathological-type effects, we grouped similar pathologies into five stimulus groups. Linear mixed-effects modelling revealed systematic expertise-related differences, with experts exhibiting longer fixation durations in pathological stimulus groups and pathology-type-dependent complexity sampling. Combined fixation duration and 3DFD features captured complementary aspects of visual expertise, improving Random Forest classifier’s accuracy (>93%) compared to individual features, for all five stimulus groups. These findings highlight neurosurgery-specific markers of visual expertise and demonstrate that combining behavioral and image-derived features could underpin computational models and training tools that emulate expert-level strategies in neurosurgical image interpretation. Future work should evaluate its applicability to other medical domains.

1. Introduction

Understanding how humans develop and apply visual expertise is a fundamental research question in cognitive neuroscience, psychology, and applied fields such as education, human–computer interaction, and medical training. Visual expertise governs how individuals process complex visual environments, make decisions, and respond to stimuli. Understanding the mechanisms underlying visual expertise is, therefore, critical for advancing theories of perception and designing effective training programs in some emerging applications, such as medical image visualization and interpretation.
To investigate the mechanisms underlying visual expertise, researchers have employed a range of methodologies, including functional magnetic resonance imaging (fMRI) [1], electroencephalography (EEG) [2], magnetoencephalography (MEG) [1], and eye-tracking [3]. Among these methodologies, eye-tracking offers a low-cost, portable and computationally simpler tool to study the visual behavior underlying expertise. In a typical eye-tracking study, participants are presented with stimuli, and their eye gaze patterns are recorded using an eye-tracking device. Here, eye tracking is employed as a sensing modality in which gaze signals are acquired at high temporal resolution and subsequently analyzed to characterize visual behavior during image viewing.
A recent focus has emerged in eye-tracking research where the goal is to objectively characterize physicians’ visual expertise. Prior work demonstrates that medical experts possess distinctive perceptual and cognitive strategies for distinguishing normal from pathological images [4,5,6,7,8,9,10]. Translating this domain-specific expertise into computational models could enable the development of “expert-like” computer systems for multiple applications [11,12]. It could enable an automated identification of abnormal medical images [13,14,15], real-time feedback for trainees during image interpretation, and objective evaluation of training progress in medical education.
Importantly, the “expert-like” computer systems must be tailored to the specific domain of expertise. For example, radiologists develop visual strategies through extensive experience interpreting diverse imaging modalities, whereas neurosurgeons acquire expertise shaped by their neurosurgical training, neuroanatomical familiarity, and case-specific decision-making demands. Because these cognitive and perceptual processes differ across medical specialties, an “expert-like” computer system must be informed by features that accurately capture the visual behavior unique to the target expert group. Consequently, identifying neurosurgery-specific markers of visual expertise requires dedicated data collection from neurosurgeons themselves, along with analytical methods capable of converting their distinct gaze behaviors into reliable, computationally interpretable features.
Previous research has examined eye-tracking data from medical experts from different domains. Many of the medical experts were radiologists [4,5,6,7,8,9,16] presented with stimuli such as chest radiographs [6,7,17], breast mammographs [4,5,8,9], or brain magnetic resonance images (MRIs) [16,18]. Few studies have reported eye-tracking data collected from surgeons. For example, Li et al. [19] compared gaze behaviors of two cardiac surgeons, the medical experts, and 12 novices who were medical residents and pre-residency medical students. The eye-tracking data were collected by performing a simulated surgical procedure. Khan et al. [20] analyzed eye-tracking data from 16 expert surgeons and novices during laparoscopic cholecystectomy operations using both live and video stimuli. Other similar studies include 14 laparoscopic surgeons wearing eye trackers while using a surgical simulator [21]. These studies consistently report significant differences in visual behavior between experts and novices, highlighting the role of visual expertise in guiding search strategies.
Despite growing interest in eye-tracking research, in our review of prior eye-tracking research, we found that there have been very few attempts to characterize visual expertise in neurosurgeons. Dalveren et al. [22] noted that this gap could be due to the difficulty in recruiting neurosurgeons for experimental studies. Consequently, prior work has been limited to small samples, such as 13 neurosurgeons performing computer-based simulated surgical tasks [22] and 9 neurosurgeons performing cutting and suturing tasks under a surgical microscope [23]. Notably, none of these studies have employed free-viewing tasks as stimuli. Our earlier work by Suman et al. in 2021 [10] used free-view tasks as stimuli but included only 4 neurosurgeons among 31 experts. Therefore, the conclusions in the Suman et al. 2021 [10] study were biased towards the visual strategies of radiologists. Since developing a bias-free domain-specific “expert-like” computer system of neurosurgeons would require data collected explicitly from this group, a focused study to identify the features characterizing their visual attention is necessary.
Prior studies have investigated a wide range of eye-tracking features to distinguish naïve observers from medical experts, aiming to capture markers of visual expertise. These features fall broadly into three methodological categories. The first category is the spatial event-level features, such as fixation duration [4,9,18,19,23], dwell time [5,9,18,24], number of returns to a region [5,9,18], and the spatial distribution of gaze [19], which collectively reflect how observers allocate visual attention across an image. The second category comprises time-series features, including saccade duration [9,18,19,23], blink duration [19], shifts of visual focus between experts and novices [19], mean latency and mean peak velocity of saccades [24], the Hurst exponent derived from R/S and DFA analyses [10], and two-dimensional fractal dimension (FD) calculated from the eye-tracking time-series data [10]. All these features characterize the temporal structure and long-range dependencies in gaze behavior. The third category is the temporal sequence-based scanpath features, such as ScanMatch [3,16], MultiMatch [25] and SoftMatch [26], methods that quantify the similarity of fixation sequences. Together, these approaches provide complementary insights into the spatial, temporal, and sequential dynamics that underpin expert visual search strategies.
Numerous studies have established that gaze behaviors reflect underlying cognitive processes such as attention allocation, perceptual load, and information processing. Longer fixation durations typically signal deeper attentional engagement with relevant visual information, whereas higher saccade rates indicate more exploratory search behavior [27]. Reduced blink rate and shorter saccade amplitudes similarly correspond to increased attentional demand. These cognitive interpretations provide important context for examining expertise-related differences in visual behavior, as they clarify how fixation- and saccade-based metrics serve as proxies for the strategies experts use when interpreting complex medical images [28].
Despite recent advances, eye-tracking studies remain comparatively limited in scope and methodological diversity, and important aspects of the visual information guiding gaze behavior are poorly understood. A notable limitation of existing work is that image content is typically treated as an implicit background factor rather than a quantifiable variable. As a result, most studies focus on behavioral metrics such as fixation duration, fixation count, or scanpath patterns, without explicitly characterizing the structural properties of the image regions being fixated. Consequently, it remains unclear how expertise-related gaze differences reflect systematic interactions between visual attention and underlying image complexity.
Fractal dimension (FD) has been used in eye-tracking research as a quantitative descriptor of scanpath complexity, with two-dimensional FD capturing irregularity and self-similarity in gaze behavior during medical image interpretation [10]. However, these studies have focused on the trajectory of visual search rather than the structural geometric complexity of the image itself. Therefore, the research question of how brain MRI spatial complexity interacts with where observers look remains unexplored. Spatial complexity is critical in neurosurgical image visualization and interpretation because pathological entities such as tumors, cerebrovascular malformations, and inflammatory changes often exhibit heterogeneous, irregular structures that challenge perceptual tasks and diagnostic reasoning [29,30,31].
Despite growing interest in expert–novice differences, no previous work has jointly modelled fixation behavior and local image complexity of the underlying image regions, a gap that remains largely unexplored in visual expertise characterization. Addressing this gap could reveal whether experts allocate gaze to diagnostically informative regions rather than being driven by surface-level complexity, offering insights into mechanisms of visual expertise. Such insights could, in turn, support developing computational models for training and decision support in many applications, including neurosurgery-specific models.
The main goal of this work is to design a focused study to characterize the visual expertise of neurosurgeons without any bias from the data of other medical experts. Participants from three groups of naïve observers, neurosurgery registrars, and consultant neurosurgeons viewed normal and pathological brain MR images under a free-viewing paradigm. To capture expertise-driven gaze behavior, we jointly modelled two features at each fixation location: (i) fixation duration, which reflects temporal allocation of visual attention, and (ii) three-dimensional fractal dimension (3DFD) of the underlying image region, which quantifies local structural complexity. Fixation duration reflects extended, task-relevant attentional engagement, while 3DFD at fixation locations captures the local structural heterogeneity of sampled regions. Together, these features denote complementary indicators of the mechanisms hypothesized to underline expert viewing in neurosurgical MRI. This integrated approach enables us to examine how experts distribute attention relative to image complexity, providing neurosurgery-specific markers of visual expertise that can inform future “expert-like” computational models. Note that our selected features fall within the first eye-tracking features category, i.e., the spatial event-level features category.
To explicitly test how fixation behavior and local image structural complexity contribute to neurosurgery-specific visual expertise, we adopted an integrated analytical framework combining inferential and predictive modelling. Using linear mixed-effects modelling, group-level expertise effects were assessed while accounting for repeated observations within participants and stimuli, thereby enabling inference about systematic expertise-related differences. In contrast, machine-learning classifiers provide a means of evaluating the discriminative value of fixation duration and 3DFD features by learning non-linear decision boundaries in multivariate feature spaces. Combining these approaches allows both inferential (linear mixed-effects) and predictive (machine-learning) aspects of neurosurgical visual expertise to be examined within an integrated analytical framework.
Three hypotheses were tested in this study using an integrated analytical framework.
  • H1: Fixation duration will be longer for neurosurgery experts than for naïve observers when viewing pathological brain MRIs, indicating selective, task-relevant temporal allocation of visual attention.
  • H2: The structural complexity of image regions fixated by observers, quantified using 3DFD, will be lower for neurosurgery experts than for naïve observers when viewing pathological brain MRIs, but in a pathology-dependent manner, reflecting prioritization of diagnostically informative structure over surface-level complexity.
  • H3: Joint modelling of fixation duration and 3DFD, reflecting complementary local temporal and structural sampling, will improve discrimination between naïve observers and neurosurgery experts compared with either feature alone, when evaluated using non-linear machine-learning classifiers.
By testing these hypotheses across multiple brain MRI stimulus groups, this study examines how temporal fixation behavior and spatial image complexity jointly encode neurosurgery-specific visual expertise. To our knowledge, no prior study has jointly modelled fixation behavior and local image-derived structural complexity at fixation locations in neurosurgeons. This is the first systematic study to apply 3DFD to neurosurgical eye-tracking data and to integrate 3DFD with fixation-duration metrics for characterizing visual expertise in neurosurgical image visualization. By explicitly contrasting linear statistical inference with non-linear classification performance, this work establishes a computational foundation for future neurosurgery-specific expert-like systems.

2. Materials and Methods

2.1. Participants

The participants for this study were recruited via email announcement and word-of-mouth. Each participant was contacted individually to explain the study, experiment setup, and inclusion/exclusion criteria.
Expert participants. The inclusion criteria for neurosurgery expert participants in our study were based on whether the participant held a board certification or other post-graduate neurosurgery qualification. Accordingly, the expert participants in our study were a combination of consultant neurosurgeons (n = 24) and trainee neurosurgery registrars (n = 16).
The expert participants in our study were recruited in two sets. The first set of participants (Np = 27) was recruited from the 77th Annual Scientific Meeting of the Neurosurgical Society of Australasia, which took place in Sydney (Australia) on 14–16 September 2022. The second set of participants (Np = 13) was recruited from Macquarie University, Sydney, Australia. Overall, there were 11 females, 28 males, and 1 preferred not to say. The expert participants were aged from 26 to 61 years. All participants reported normal or corrected-to-normal vision.
Naïve participants. The inclusion criteria for naïve participants in our study were a lack of any prior experience reading medical images. Accordingly, the naïve participants (Np = 29) in our study were recruited from a population of Macquarie University students and staff. There were 16 female and 13 male participants, aged from 23 to 52 years. All participants reported normal or corrected-to-normal vision.

Participant Grouping for Visual Expertise Analysis

Groups: The expert group in our study was split into two groups. The neurosurgery consultants (Np = 24) were grouped into the expert consultant group (EC), and the neurosurgery registrars (Np = 16) into the expert registrar group (ER). A naïve group (N) of participants (Np = 29) was formed as the third group in our study. The resulting participant groups and demographic characteristics are presented in Table 1.

2.2. Stimuli

Each participant in this study was presented with a set of 39 de-identified T1-weighted brain MRIs as stimuli. The stimuli were two-dimensional MRI images stored as PNG files. The stimulus images were standardized to a size of 1920 × 1080. The MRI stimuli set presented to participants was a mixture of images taken in different views (axial, sagittal, and coronal). Twenty brain MRIs did not contain any pathology, and the other 19 included some pathologies (e.g., brain or skull base tumors and aneurysms). The normal and pathological images were intermixed, and a fixed sequence of the intermixed images was created. This sequence of images was then presented to each participant during their eye-tracking data collection session.

Stimulus Grouping for Visual Expertise Analysis

Seven stimulus groups (A–E, StimP, StimN) were constructed from the 39 stimulus images for visual expertise analysis. Six groups (A–E, StimP) contained a unique collection of brain pathological MR images with various types of pathology. These groups represented a diverse set of MR images commonly seen by neurosurgeons in their neurosurgery practice. One group (StimN) contained normal brain MR images. The full taxonomy of stimulus groups, including image type, pathological status, and image counts, is provided in Table 2. The pathological images used in our study are provided in Appendix A Figure A1.

2.3. Apparatus

A portable eye tracker (EyeLink Portable Duo 1000) from SR Research Ltd. (Ottawa, ON, Canada) was used in this study. Stimuli were presented on a 15-inch laptop screen with a 1920 × 1080 display resolution operating at a 60 Hz refresh rate. The eye-tracking system operates based on infrared video-oculography, using high-speed optical sensors to continuously capture pupil and corneal reflection signals. Ocular movements were measured using the eye-tracker at a sampling rate of 1000 Hz. Figure 1 shows an illustration of our experimental setup with this eye-tracker.

2.4. Ethics Approval and Consent

All procedures were approved by the Macquarie University Medicine and Health Sciences Subcommittee (Ref: 52021919128177; approved on 18 May 2021). As per ethics approval, this study meets the requirements set out in the National Statement on Ethical Conduct in Human Research 2007 (updated July 2018), which embodies the ethical principles of the Declaration of Helsinki 1964. As per the ethics approval, this study falls within the category of “Clinical research that is not a clinical trial—observational study”.
Written consent was obtained from all participants prior to data collection.

2.5. Free-View Task Study Design

A free-view task study was designed where participants were only required to view the stimulus images and did not engage in any additional tasks. Since no specific diagnostic task was imposed, the fixation strategies were expected to reflect spontaneous expertise-driven search patterns rather than task-directed instructions.

2.6. Experimental Procedure

2.6.1. Instructions to Participants Before Data Collection Session

Each participant was asked to complete a consent form and demographic questionnaire. This questionnaire included their age, neurosurgery expertise in reading brain MRIs (if any), and the average number of radiological cases reviewed per week. Following the completion of forms, the stimulus presentation sequence was explained to the participant with the aid of an illustration printed on paper.
A stimulus presentation sequence is demonstrated in Figure 1c. Participants were informed that they can view anywhere on the noise mask but must fixate centrally at the cross on the noise mask with fixation. Participants were also informed that following the two noise mask images, they would view a series of stimulus images on the laptop screen but were free to view anywhere on those stimulus images.

2.6.2. Eye-Tracker Setup and Calibration Before Data Collection Session

The operational mode of the eye-tracker was chosen as head free-to-move Remote mode to allow data collection without any head support. Participants were asked to sit on a chair as illustrated in Figure 1b. A target sticker was placed on their forehead to aid the eye-tracker’s software in finding the target-to-camera distance between 480 and 520 mm. The eye-tracker was then calibrated using a 13-point calibration procedure. During calibration, participants were asked to keep their heads still. The laptop display was subtended to a maximum visual angle of 32° horizontally and 25° vertically. It was ensured that the calibration for both eyes was good in the eye-tracking software, with a calibration error of <1°.

2.6.3. Data Collection Session

Short rest breaks were provided between blocks for the participants experiencing fatigue during their recording sessions. Throughout the data collection session, participants were asked to keep their heads still. The sequence of images (trial) consisted of a noise mask presented for 500 ms, a noise mask with cross fixation presented for 1500 ms, and then a stimulus image presented for 5000 ms. Each trial started with a 1-point calibration to ensure that participants were properly calibrated throughout the experiment. The scanpath (eye gaze) data were collected for each of the 39 trials at a sampling rate of 1 KHz. Therefore, there were 5000 samples in the scanpath data per trial.

2.7. Eye-Tracking Raw Data Processing and Report Generation

The raw data acquired from the eye-tracker comes in a form called EyeLink Data Files (EDF). The EDF files contain time-stamped gaze signals captured by the optical sensor in the eye-tracker, including eye position samples recorded at the device sampling rate. During this preprocessing stage, the raw gaze signals were parsed into discrete fixation and saccade events using the eye-tracker’s “EyeLink host online parser” algorithm, which is a velocity- and acceleration-threshold-based algorithm executed during data recording. In our experiments, one EDF file was generated for each participant. These EDF files were subsequently processed offline using EyeLink Data Viewer software (version 4.2.1) to generate reports for each participant [32]. Only the right eye’s data were utilized in this study for analysis.

2.8. Feature Extraction

All features analyzed in this study were derived from fixation events detected from the eye-tracking sensor’s gaze signals, linking the sensor-level measurements to quantitative descriptors of visual behavior.

2.8.1. Fixation Duration

For each participant, a report was generated using EyeLink Data Viewer software (version 4.2.1) containing the coordinates of the fixation locations for each stimulus viewed and the duration of the fixation at that location. A template for this report is provided in Appendix A Table A1. Let n represent the index of the n -th fixation location, and x n and y n represent the horizontal and vertical coordinates of the fixation location at n , respectively. Then, the eye’s fixation position at n can be represented with a point in 2D space as:
P o s n = x n , y n
Let the eye’s fixation duration at P o s n be represented as F i x D u r n . Let us denote the participant by P and the stimulus by s . From each participant’s report, we extracted all fixation durations for each stimulus s and arranged them in an array as:
F i x D u r P s = F i x D u r P s 1 , F i x D u r P s 2 , , F i x D u r P s N P s
Here, N P s denotes the number of fixations for stimulus s and participant P . We then sorted the fixation durations for the stimulus s in descending order as:
F i x D u r d e c e n d P s = d e c e n d i n g F i x D u r P s 1 , F i x D u r P s 2 , , F i x D u r P s N P s
Let the first M P s maximum fixation durations of this sorted fixation duration array be represented as:
F i x D u r d e c e n d f i r s t M m a x P s = F i x D u r d e c e n d P s 1 , F i x D u r d e c e n d P s 2 , , F i x D u r d e c e n d P s M P s
In this study, we empirically choose the M P s = 10, which corresponds to the first 10 maximum fixations durations for each stimulus for each participant.
  • Heat map of fixation duration for visualization of fixations
The heat map of fixation durations was calculated for each stimulus and for each participant using the EyeLink Data Viewer software (version 4.2.1) provided by EyeLink [32]. In this software, the heat map is generated by applying a 2D Gaussian filter to each of the fixations. The center of the Gaussian function is chosen as the location of fixation. The width of the Gaussian function is determined by an adjustable standard deviation value, i.e., sigma, which represents the visual angle in degrees. A larger sigma value results in a greater area being affected by the fixation. We set sigma as 1.0 degrees, which is the default setting in the EyeLink Data Viewer software (version 4.2.1). In the software, the height of the Gaussian function is weighted by the duration of the fixation.

2.8.2. Three-Dimensional Fractal Dimension (3DFD)

Fractal dimension (FD) is a metric that quantifies local geometric complexity, which may arise from spatial irregularity or heterogeneity within an image. In this study, the computation of FD for each pixel of the stimulus image is referred to as the 3DFD map. The 3DFD map represents the complexity of the structures present in the stimulus image. For each fixation point P o s n , the corresponding fractal dimension value was extracted from the 3DFD map. Let the 3DFD value at P o s n be represented as F D 3 D n . Thus, for every fixation location, F D 3 D n denotes the local structural complexity of the image region that attracted the observer’s gaze. We collected the 3DFD values for each participant P and stimulus s as:
F D 3 D P s = F D 3 D P s 1 , F D 3 D P s 2 , , F D 3 D P s N P s
Let F D 3 D d e c e n d P s represent the sorted values of F D 3 D P s arranged in descending order of fixation durations. The F D 3 D d e c e n d P s values corresponding to the first M P s maximum fixation durations will then be represented as:
F D 3 D d e c e n d f i r s t M m a x P s = F D 3 D d e c e n d P s 1 , F D 3 D d e c e n d P s 2 , , F D 3 D d e c e n d P s M P s
1.
Generation of 3DFD map from stimulus image
The 3DFD maps were calculated for each of the 39 stimulus images by implementing the box counting method [33]. The MRI stimuli were analyzed as two-dimensional grayscale images rather than as reconstructed volumetric MRI datasets. For 3DFD computation, each grayscale MRI image was represented as a three-dimensional intensity surface, where the x- and y-axes corresponded to image coordinates, and the z-axis corresponded to grayscale intensity. Therefore, the term 3DFD refers to the fractal dimension of the grayscale intensity surface derived from each 2D MRI image, rather than to the fractal dimension of a volumetric brain reconstruction from multiple 2D MRI slices.
There were two settings in our implementation of the box counting method: shift and patchsize. For each point within the stimulus image, a moving window of patchsize × patchsize = 100 × 100 was created around the point, and the grayscale pixel values included in this patch were used to calculate the 3DFD value at that point. A shift × shift = 10 × 10 was used to move the patch through the image from left-to-right and from top-to-bottom. We implemented the box counting method for the calculation of 3DFD maps in the Python 3.10 programming language. Our implementation generated the 3DFD values of each stimulus in a 3-column csv format. A MATLAB code was written to convert the 3-column CSV file into an image of the same size as the stimuli image.
2.
Foveal vision-driven processing of 3DFD maps to calculate F D 3 D n
For each stimulus s, all the fixation points F n     n 1 , N P s were read from their eye-tracking data, and circular regions of interest (ROIs) were created by taking these points as centres. Foveal vision is defined as the central 1.5–2 degrees of the visual field. Vision within the fovea is generally called central vision. Considering the foveal vision, the radius of the ROI was taken as 1.5 degrees because this approximates the central foveal field used for detailed visual inspection at fixation. This corresponded to a radius of 30 pixels in our stimulus image. Mean values of the pixels within the ROIs were taken from the 3DFD maps and used to calculate F D 3 D n .
After calculating F D 3 D n for all fixation locations, the values corresponding to the M P s = 10 longest fixation-duration locations were selected for the main analysis. Thus, the M P s = 10 points used in the analysis refer to the ten longest fixation-duration locations for each participant–stimulus pair, not to MRI slices or anatomical sections. At each of these fixation locations, the local 3DFD value was extracted from the precomputed 3DFD map described in item 1 of Section 2.8.2.

2.9. Linear Mixed-Effects Modelling and Statistical Analysis

2.9.1. Model Equations

Fixation duration model:
F i x a t i o n D u r a t i o n ~ 1 + G r o u p × I m a g e T y p e + 1 | P a r t i c i p a n t + 1 | S t i m u l i
3DFD model:
F D 3 D   ~   1 + G r o u p × I m a g e T y p e   + 1   |   P a r t i c i p a n t +   1   |   S t i m u l i  
These base models were fitted separately for each of the stimulus groups (A–E, StimP) by combining the data from these pathology stimulus groups with the normal stimuli group (StimN).
In both models, G r o u p = N , E C , E R and I m a g e T y p e = N o r m a l , P a t h o l o g y were treated as fixed effects using naïve observers (N) and N o r m a l image-type as the reference categories, respectively. An interaction term G r o u p × I m a g e T y p e was included in the model to test if the type of pathology (tumor, cerebrovascular pathologies etc.) affects the differences across N , E C , E R groups. In both models, the random intercepts for Participant captured the differences across the participants within a group, and the random intercepts for Stimuli accounted for the stimuli-to-stimuli variability within the combined stimulus group (A–E, StimP) and normal group (StimN).
A total of 12 models were created, 6 for each of the fixation duration and 3DFD. All models were fitted using MATLAB’s fitlme function with maximum likelihood as the method for estimating the model parameters.
Appendix Mathematical Form of Linear Mixed-Effects Modelling Equations provides details of the mathematical equations corresponding to the R-notation Equations (7) and (8).

2.9.2. Effect Size Computation

For each fixed-effect contrast (N vs. EC, N vs. ER), Cohen’s d was computed as:
d = β σ e
where β is the model-estimated fixed-effect coefficient and σ e is the residual standard deviation obtained from the mixed-effects model. This definition reflects the magnitude of the effect relative to within-subject variability and is appropriate for linear mixed-effects modelling. Effect sizes were interpreted according to established benchmarks: small ( d ≈ 0.1–0.3), medium ( 0.3–0.5), and large (>0.5).

2.9.3. False Discovery Rate (FDR) Correction

False discovery rate (FDR) correction was applied to control multiple comparisons within each family of statistical tests. In this context, a family of statistical tests refers to all hypothesis tests conducted for the same outcome variable (e.g., fixation duration or 3DFD) across all stimulus groups (A–E, StimP) combined with group (StimN). No additional multiple-comparison correction beyond FDR correction was applied.

2.9.4. Sample Size of Data

Across all participants and stimuli, the dataset comprises 69 participants, 39 stimuli and 10 fixation locations per participant–stimulus pair. Therefore, the maximum number of fixation-level observations available for the primary analysis was: 69 × 39 × 10 = 26,910 observations. Because separate LMMs were fitted for each pathology stimulus group by comparing that group with StimN, the number of observations entering each model varied by stimulus group, with the StimP model using the maximum observation count and the pathology-specific models using fewer observations. Because these observations were repeated within participants and stimuli, they were not treated as statistically independent observations. Instead, participant and stimulus were included as random intercepts in the mixed-effects models. This modelling framework allowed the repeated-measures structure of the dataset to be used while accounting for within-participant and within-stimulus dependency. It should be noted that the fixation locations where 3DFD values were ≤1.7 were not included in our LMM analysis.

2.9.5. Software for Statistical Data Processing

Appendix A.4.1 provides details of the software used.

2.10. Non-Linear Modelling Using Supervised Machine Learning Methods

To test hypothesis H3, we examined whether joint non-linear modelling of geometric complexity (3DFD) and fixation duration improves discrimination between naïve and expertise groups compared with models based on either feature alone. Supervised machine-learning classifiers were therefore trained using three feature configurations: fixation duration only, 3DFD only, and a combined feature set incorporating both measures.

2.10.1. Feature Engineering for Machine Learning Classifiers

The features were created in three main steps. Firstly, all fixation durations for a given image–participant pair were sorted in descending order, and the 10 longest fixation durations were selected. The 10 longest durations generally represented most of the temporal engagement of participants with a stimulus, so it was empirically chosen and fixed across all the stimuli. For each of these 10 fixations, their fixation coordinates were taken, and the corresponding 3DFD values were retrieved from the precomputed 3DFD map at those coordinates. Three sets of features were then created as shown in Table 3.
The choice of 10 features is consistent with the number of data points taken from each of the participant–stimulus data points in mixed-effects modelling ( M P s = 10).

2.10.2. Machine Learning Classifiers

Two supervised machine-learning classifiers were trained to classify the visual expertise of participants into naïve (N) or expert consultant (EC), naïve (N) or expert registrar (ER), expert consultant (EC) or expert registrar (ER). For each of these binary classifiers, two classifiers were developed using Random Forest and Support Vector Machine.
  • Random Forest (RF)
A Random Forest classifier was implemented using the TreeBagger algorithm with 200 trees (TreeBagger in MATLAB 2025b). Each tree was grown on a bootstrap-resampled subset of the training data. At each split, a random subset of predictors was sampled automatically by MATLAB’s internal heuristic for classification forests. Trees were grown to full depth without pruning to maximize ensemble diversity, and class predictions were obtained through majority voting across trees. This configuration has been widely used by other researchers in biomedical pattern-recognition tasks due to its merits, such as its effectiveness in handling non-linear decision boundaries, robustness to noise and reduction in the risk of overfitting by averaging across diverse trees [34,35].
2.
Support Vector Machine (SVM)
A Support Vector Machine classifier with a radial basis function (RBF) kernel (fitcsvm in MATLAB 2025b) was trained to model non-linear decision boundaries in the feature space. The features were standardized before training to improve optimization and margin estimation. MATLAB’s internal automatic hyperparameter selection was used for the RBF kernel scale. This kernel-based SVM is well-suited for non-linear classification problems and is consistent with other biomedical modelling studies when dimensionality is moderate [36]. For completeness, we also evaluated a linear-kernel SVM, in addition to the RBF SVM reported in the main analysis.

2.10.3. Software for Machine Learning Classifiers

Appendix A.4.2 provides details of the software used.

2.10.4. 10-Fold Cross-Validation and Performance Evaluation

We used 10-fold cross-validation as the primary method for reporting all performance metrics, including accuracy, sensitivity, specificity, precision, F1 score, and AUC. For this 10-fold cross-validation, the classifier-level dataset was partitioned into 10 folds, with 9 folds used for training and the remaining fold for testing. This process was repeated until each fold had served once as the test set, and the performance metrics from these 10 runs were averaged to obtain a single performance estimate.
All cross-validation partitions were performed at the participant level, ensuring that all participant–stimulus samples from a given participant were assigned exclusively to either the training or testing set within a fold, thereby preventing subject-level data leakage. Because the classification analyses were performed as binary group comparisons, the actual classifier-level sample size varied according to the two groups being compared and the number of stimuli in each stimulus group. The details of the total size of the input dataset for 10-fold subject-wise cross-validation are provided in Appendix D, Table A11 and Table A12. The exact fold size could vary slightly because fold assignment was performed at the participant level. We additionally conducted a sensitivity analysis by repeating the same participant-level cross-validation with K ∈ {2, 3, …, 9}, examining the stability of the performance metrics across different choices of fold sizes.

3. Results

We report inferential (linear mixed-effects) and predictive (machine learning) findings addressing H1–H3 across all stimulus sets. Throughout this section, we adopt the notation A–E to denote pathology-based stimulus groups: A (tumors); B (cerebrovascular pathologies); C (other pathologies); D (inflammatory changes); E (malformations), as presented in Table 2.

3.1. H1: Expertise-Related Differences in Fixation-Duration Behavior

3.1.1. Qualitative Comparison of Fixation-Duration Distributions

A comparison of the averaged fixation duration based heatmaps (Figure 2 and Figure 3) for naïve, consultant and registrar groups indicates that there are differences in their gaze pattern. The experts (consultant and registrar groups) appear to focus on a pathology for a longer duration than the Naïve participants. The Naïve participants do not appear to fixate on a region for a longer duration; instead, their fixations are dispersed over a wider area on the stimulus (Figure 2b and Figure 3b). The focus of naïve participants appears to be dispersed not only for a smaller-sized pathology in Figure 2b, but also for a relatively larger-sized pathology in Figure 3b.
Figure 4 shows boxplots comparing naïve participants viewing normal images (NN), expert consultants viewing normal images (ECN), expert registrars viewing normal images (ERN), and the corresponding groups viewing pathological images (NP, ECP, ERP). Across all stimulus groups in Figure 4, the swarm overlays for naïve participants (NN, NP) show wider horizontal spread, forming broader “violin-like” point clouds. This pattern indicates a greater dispersion of fixation durations relative to the slightly tighter distributions observed in expert groups (ECN, ERN, ECP, ERP). Note that the fixation durations in these boxplots and violin plots were taken from the eye tracking reports (see Appendix A Table A1 for report template), where only data corresponding to M P s = 10 maximum fixation durations were used.
Importantly, within each subpanel (a–e) in Figure 4, the total number of plotted fixation-level observations differs between normal and pathology because the number of stimuli per condition is not the same: the normal (StimN) set contributes 20 images, whereas the paired pathology sets contribute A = 10, B = 4, C = 3, D = 1, and E = 1 images (Table 2). As each stimulus contributes its top 10 fixations per participant, aggregating across unequal numbers of stimuli yields different totals of plotted events within a subpanel. In contrast, panel (f) with “All pathological images combined” pools 20 normal and 19 pathological images, so the total number of plotted events, and thus the overall appearance of the violins, is very similar between conditions.

3.1.2. Quantitative Assessment of Fixation Duration Using Mixed-Effects Modelling

Using normal images and the naïve group (N) as reference categories, the mixed-effects models showed no significant ImageType main effects across stimulus groups A–E or in the combined analysis (all p_FDR > 0.05). In contrast, expertise-related effects emerged primarily through Group × ImageType interactions. For expert consultants (EC), EC × ImageType effects were FDR-significant in stimulus groups A (β = 19.091, p_FDR = 0.004), B (β = 25.271, p_FDR = 0.006), and StimP (β = 20.464, p_FDR <0.001), but not in C, D, or E. For expert registrars (ER), ER × ImageType effects were FDR-significant in stimulus groups A (β = 45.649, p_FDR < 0.001), B (β = 42.304, p_FDR < 0.001), C (β = 26.567, p_FDR = 0.033), E (β = 42.393, p_FDR = 0.042), and StimP (β = 39.733, p_FDR < 0.001), but not in D. Cohen’s d values indicated small effect sizes for EC interactions (d ≈ 0.11–0.15) and consistently larger, but still small effect sizes for ER interactions (d ≈ 0.14–0.26), supporting greater reliability of registrar-related effects.
The M P s = 5 maximum fixation-duration results reported in Appendix B Table A2 showed a highly convergent qualitative pattern with the M P s = 10 findings reported in Table 4. As with M P s = 10, ImageType main effects relative to the naïve–normal reference were absent, and fixation-duration differences were primarily expressed through Group × ImageType interactions. The same stimulus groups that showed reliable effects at M P s = 10 (notably A, B, and the aggregated StimP analysis) also exhibited interaction effects at M P s = 5, with additional ER-specific effects observed in C and E, while stimulus group D remained non-significant across both maximum thresholds. Effect size estimates at M P s = 5 mirrored those at M P s = 10, with EC interactions remaining small and ER interactions consistently in the small-to-moderate range, indicating that the fixation-duration findings are stable and reliable across fixation-count thresholds.

3.2. H2: Expertise-Related Differences in Structural Complexity (3DFD) at Fixation Locations

3.2.1. Qualitative Comparison of 3DFD Distributions at Fixation Locations

A qualitative comparison of 3DFD distributions is presented in Figure 5. Distributions are shown for naïve participants (NN), expert consultants (ECN) and expert registrars (ERP) when they are viewing normal images and pathological images (NP, ECP, ERP). For the interpretation of violin “fullness,” the aggregation logic and stimulus-count differences described in Section 3.1.1 (regarding Figure 4) also apply to Figure 5. Like Figure 4, the panel (f) in Figure 5 also has normal and pathology stimuli of 20 and 19, respectively.

3.2.2. Quantitative Assessment of 3DFD Using Mixed-Effects Modelling

Using normal images and the naïve group (N) as reference categories, the M P s = 10 mixed-effects models showed no FDR-significant ImageType main effects on 3DFD across stimulus groups A, B, C, D, or in the aggregated analysis (StimP) (all p_FDR ≥ 0.051). A clear exception was observed for stimulus group E (malformations), where a strong positive ImageType main effect was present (β = 0.103, p_FDR < 0.001, d = 0.857), indicating substantially higher 3DFD for pathological relative to normal images within the naïve group. In contrast to fixation-duration results, expertise-related effects on 3DFD were expressed selectively and in a stimulus-dependent manner. For stimulus group B, both EC × ImageType (β = −0.031, p_FDR < 0.001, d = −0.244) and ER × ImageType (β = −0.026, p_FDR < 0.001, d = −0.205) interaction effects were FDR-significant, indicating reduced 3DFD for experts relative to the naïve–normal reference. Conversely, for stimulus group D, strong positive interaction effects were observed for both EC (β = 0.078, p_FDR < 0.001, d = 0.582) and ER (β = 0.082, p_FDR < 0.001, d = 0.612), indicating markedly higher 3DFD values for experts on pathological images relative to naïve participants. No FDR-significant Group × ImageType interactions were observed in stimulus groups A, C, or E, nor in the StimP analysis. Cohen’s d values indicate large ImageType effects for E, moderate-to-large interaction effects for stimulus group D, small effects for stimulus group B, and negligible effects elsewhere, highlighting that 3DFD-based expertise effects at M P s = 10 are strongly stimulus-specific.
The M P s = 5 (Appendix B Table A3) 3DFD results reported in Appendix B showed a slightly convergent qualitative pattern with the M P s = 10 findings reported in Table 5. As with M P s = 10, the Group × ImageType interactions remained restricted to the same stimulus categories, with negative expert interactions in stimulus group B and positive expert interactions in stimulus group D, and no reliable interaction effects in A, C, E, or StimP. Cohen’s d values at M P s = 5 mirrored those at M P s = 10, with large ImageType effects in E, moderate-to-large interaction effects in D, smaller interaction effects in B, and negligible effects elsewhere, largely supporting the conclusion that the 3DFD findings are robust to the choice of maximum number of fixations and reflect stable, stimulus-dependent differences.

3.3. H3: Expertise-Related Differences in Joint Non-Linear Modelling of 3DFD and Fixation Duration

Table 6, Table 7 and Table 8 report 10-fold cross-validation results for the Random Forest classifier. The results are reported as mean ± standard deviation across cross-validation folds. Accuracy is expressed as a percentage, whereas sensitivity, specificity, precision, F1 score, and area under the receiver operating characteristic curve (AUC) are reported on the 0–1 scale. Accuracy and specificity values for the combined feature set are highlighted to summarize the improved overall discrimination and reduced false-positive rates achieved through joint modelling of 3DFD and fixation duration. Sensitivity, specificity, precision, F1, and AUC are defined with respect to the expert group as the positive class.
Across all three binary classification tasks (N vs. EC, N vs. ER, and EC vs. ER), the Random Forest classifiers trained on the combined feature set consistently demonstrated superior overall performance, particularly in terms of accuracy, and generally for specificity as well, across stimulus sets. The fixation-duration-only classifiers generally showed lower discriminative performance, while classifiers trained on 3DFD alone often performed strongly but with greater variability across stimuli. The overall pattern indicates that although geometric complexity provides a substantial discriminative capacity on its own, fixation duration contributes complementary temporal information that enhances classification when jointly modelled. The advantage of the combined feature representation was observed consistently across all stimulus groups (A–E, StimP, StimN). The only exception to this trend was the stimulus group D for N vs. ER and EC vs. ER classifiers, where the 3DFD-only classifier performed slightly better than the combined-feature-set-based classifier.
The SVM classifiers showed a consistent pattern across all three binary tasks. Fixation-duration-only SVM models generally yielded the weakest performance, whereas 3DFD-only models often performed better but with greater variability across stimulus groups. Unlike the Random Forest classifiers, SVMs did not show benefit from combining fixation duration and 3DFD, with combined-feature models often performing poorer than single-feature models. Full numerical results for all stimulus groups are provided in Appendix B Table A4, Table A5, Table A6, Table A7, Table A8 and Table A9 for M P s = 10 longest fixation duration locations and K = 10 folds.
Multi-K robustness analysis results for K ∈ {2, 3, …, 10}, as provided in Appendix B Figure A3 (for M P s = 10 longest fixation duration locations) and A4 (for M P s = 5 longest fixation duration locations) for Random Forest and SVM classifiers, show that accuracy remained broadly consistent across cross-validation settings for Random Forest models. However, the SVM accuracies vary with K across all pathology groups: A (tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), and StimP (combined stimuli images), for both M ρ s = 5 and M ρ s = 10 . Across K ∈ {2, 3, …, 10} and stimulus groups, linear-SVM (Appendix B Figure A5) and RBF-SVM (in Appendix B Figure A3) achieved similar accuracies, and both were consistently below Random Forest, particularly for the combined (3DFD + fixation duration) features; thus, including a linear kernel did not alter the conclusion that Fandom Forest provides the strongest and most stable performance.

3.4. Linking Mixed-Effects and Machine-Learning Analyses

Across all three binary classification tasks (N vs. EC, N vs. ER, and EC vs. ER), classifiers trained on fixation duration alone generally showed lower performance, while 3DFD-only classifiers often performed strongly but with greater variability across stimuli. This pattern mirrors the mixed-effects findings, where fixation duration and 3DFD exhibited independent but non-identical contributions to group differentiation. Moreover, while the mixed-effects models tested group-level associations separately for fixation duration and 3DFD, the classification analysis tested whether these features, alone or in combination, could discriminate expertise status in unseen participants. Thus, the machine-learning analysis addressed the predictive component of H3, which is not addressed by the LMMs.

4. Discussion

We demonstrated that the visual expertise of neurosurgeons could be characterized by jointly analyzing fixation durations derived from eye-tracking data and local image structural complexity at the locations of fixations, during free viewing of brain MRIs. Fixation duration and 3DFD captured complementary aspects of gaze behavior, and their combined consideration provided a more informative description of expertise-related variation than either feature in isolation.

4.1. Heat Maps of Fixation Duration

The heat maps of fixation durations presented in Figure 2 and Figure 3 and in Appendix B Figure A2 represent the group-level heat maps of fixation durations averaged across all the participants within the group. The qualitative analysis in Figure 2b and Figure 3b indicates that the naïve participants do not focus on a particular region within the stimulus image and instead allocate their attention throughout the image. This is attributed to their lack of knowledge of the domain of stimuli. Experts, on the other hand, are trained in image interpretation, which helps them focus on pathology for a longer duration (Figure 2b and Figure 3b), (Figure 2c and Figure 3c). We emphasize that the gaze pattern is only suggestive of attentional allocation, and we acknowledge that covert attention may not always align fully with gaze position. Notably, this is not general behavior and, in some prominent pathologies, such as the one shown in Appendix B Figure A2b, the allocation of visual attention is less dependent on expertise. In this example, naïve participants could allocate most of their visual attention to pathological regions for similar durations as the experts. Note, however, that the pattern of fixation distributions of naïve participants still shows that they get attracted to the other normal regions of the stimulus image.

4.2. Linear Mixed-Effects Modelling

4.2.1. Expertise Effects in Fixation Duration

Consistent with the mixed-effects modelling results (Table 4 and Appendix B Table A2), experts demonstrated longer fixation durations on pathological images relative to naïve participants, with these differences expressed through Group × ImageType interactions rather than uniform main effects, which supports hypothesis H1. When interpreted alongside the qualitative observations reported in Section 4.1, these findings suggest that experts allocate visual attention more selectively, prolonging fixations when clinically relevant abnormalities are present. Rather than reflecting a general tendency toward longer fixations, this pattern indicates a context-dependent modulation of fixation duration, whereby experts engage in sustained visual inspection, specifically when pathology is detected. This targeted prolongation of fixations on pathological images is consistent with established models of visual expertise, which propose that experts are able to suppress irrelevant visual information while dedicating increased attentional resources to diagnostically meaningful regions. In the free-viewing paradigm used in this study, experts appeared to rapidly identify salient pathological features and subsequently maintain fixation on these regions for longer durations, supporting the inference that fixation duration reflects strategic information extraction rather than inefficient search behavior.
Importantly, expertise-related differences in fixation duration were not uniform across all stimulus categories but varied as a function of both image type and observer group (Table 4, Appendix B Table A2). Across pathological stimulus categories, both consultants and registrars exhibited longer fixation durations than naïve participants, irrespective of the number of maximum fixation points considered ( M P s = 5 or M P s = 10). This pattern is consistent with the expectation that when trained observers identify pathology, they engage in more sustained scrutiny of abnormal regions even under unconstrained free-viewing conditions.
Notably, fixation-duration effects were often more pronounced in registrars than in consultants, suggesting that prolonged fixation may reflect an intermediate stage of expertise. One plausible explanation is that registrars, while capable of identifying pathological regions, may rely on extended visual confirmation as part of their interpretative process. With increasing experience, consultants may achieve similar diagnostic outcomes with less prolonged fixation, reflecting greater efficiency in visual processing and decision-making. Under this explanation, elevated fixation duration in registrars may represent a transitional marker of developing expertise, rather than a monotonic indicator of expert performance.

4.2.2. Image Complexity and 3DFD Effects

Higher fractal dimension (FD) values correspond to increased geometric complexity in visual stimuli [37]. The 3DFD analysis revealed stimulus-dependent effects of ImageType, with pathological images exhibiting higher 3DFD values than normal images for specific pathology categories, most prominently for malformations (stimulus group E) and, to a lesser extent, in the aggregated analysis. These findings indicate that certain pathological stimuli are characterized by greater structural irregularity and visual complexity, and that 3DFD is sensitive to these differences when they are present. Importantly, the absence of uniform ImageType effects across all stimulus groups suggests that image complexity varies substantially across pathological types, rather than reflecting a global distinction between pathological and normal images. This stimulus-specific pattern is consistent with our hypothesis H2, which predicts that expertise-related 3DFD differences should emerge only for pathology categories where local structural complexity is diagnostically relevant.
Because 3DFD values were computed at fixation locations, these effects reflect the complexity of image regions that observers actively sampled. Importantly, expertise-related differences in 3DFD were not uniform across stimulus groups. Significant Group × ImageType interaction effects were observed selectively, with experts exhibiting significantly lower 3DFD values than naïve participants in stimulus group B, suggesting reduced sampling of highly complex regions in this pathology category of cerebrovascular pathologies. In contrast, experts showed significantly higher 3DFD values in stimulus group D, indicating that expert observers selectively engaged with more structurally complex regions when such complexity was diagnostically relevant in this pathology category of inflammatory changes. No reliable expertise-related modulation of 3DFD was observed in stimulus groups A, C, E, or in the aggregated analysis.
These stimulus-dependent effects indicate that experts are not simply less influenced by visual complexity but rather modulate their sampling of complex image regions based on diagnostic relevance. In some cases, experts deprioritize visually irregular regions that lack diagnostic value, whereas in others they deliberately focus on structurally complex areas that are clinically informative. This pattern aligns with models of perceptual expertise, suggesting that experts rely less on low-level salience and more on semantic knowledge, structured search strategies, and internal templates when interpreting medical images [38,39].

4.2.3. Integration of Complexity and Fixation Behavior

Taken together, the fixation duration and 3DFD results reveal a complementary pattern of attentional control. Naïve participants tend to be attracted to visually complex regions and distribute fixations more broadly, whereas experts exhibit greater selectivity in where they fixate, coupled with longer fixation durations when pathology is identified. Thus, expertise is characterized not by increased sensitivity to visual complexity per se, but by strategic allocation of attention based on diagnostic relevance.
The relatively small effect sizes observed in the 3DFD models, contrasted with larger and more consistent effects in fixation duration, suggest that 3DFD primarily captures image-driven differences, whereas fixation duration more strongly reflects observer-driven behavioral modulation. This distinction is important, as it indicates that 3DFD and fixation duration quantify independent yet complementary dimensions of visual expertise: one reflecting the property of the visual stimulus at sampled locations, and the other reflecting how observers engage with diagnostically meaningful information over time.

4.2.4. Comparison with Prior Studies in Fixation Duration and Geometric Complexity

A prior study by Suman et al. [10] reported expertise-related differences in visual search behavior during interpretation of pathological medical images, broadly consistent with the fixation-duration effects observed in the present study. However, the expert cohort in their study comprised radiologists from a range of subspecialties (e.g., neuroradiology, head and neck, chest, musculoskeletal, and interventional radiology), with only 4 out of 31 experts being neurosurgeons. As a result, their findings reflect general radiological expertise rather than expertise specific to neurosurgical image interpretation. In contrast, the present study explicitly focuses on neurosurgeons, allowing the reported fixation-duration patterns to be interpreted within a domain-specific clinical context.
Importantly, Suman et al. [10] quantified expertise effects using fractal dimension analysis of scanpaths, reporting higher geometric complexity of scanpaths when experts viewed pathological stimuli. While this approach captures differences in eye-movement organization, it does not directly relate scanpath complexity to the geometric complexity of the underlying image content. In contrast, the current study applied a 3D fractal dimension analysis to the stimulus images themselves, evaluated at fixation locations. This approach enables a more direct examination of how observers sample image regions of varying structural complexity, thereby linking fixation behavior to the geometric properties of the visual stimulus. Together, these methodological differences highlight that scanpath-based and image-based fractal analyses capture complementary aspects of visual expertise, with the present findings extending prior work by explicitly relating gaze behavior to the structural complexity of pathological image regions.

4.3. Non-Linear Modelling with Supervised Machine Learning

In addition to inferential statistical analysis, machine learning classifiers were employed to evaluate whether the identified features can generalize to the prediction of unseen participants. While statistical models establish group-level differences, predictive modelling provides a complementary assessment of the discriminative utility and practical applicability of these features in real-world scenarios such as developing “expert-like” systems for assessing visual expertise.
It is important to clarify that the role of machine learning in this study is not to replace statistical inference, but to complement it. While linear mixed-effects models establish the presence and nature of expertise-related differences, machine learning evaluates whether these differences are sufficiently robust to enable accurate prediction at the individual participant level. This unified analytical framework provides both explanatory and predictive validation of the proposed features.

4.3.1. Joint Modelling of Complexity Using 3DFD and Fixation Duration

The advantage of the combined feature representation was observed consistently across stimulus sets, including both pathology-specific and aggregated pathological image groups (A–E, StimP) as well as in the normal stimuli group StimN. The higher accuracy and specificity in the combined-feature classifiers achieved more reliable overall correctness while reducing false-positive classifications. This indicates that our joint modelling of the two spatial event-level categories of features, the fixation duration and spatial image complexity, provides a more complete characterization of visual expertise than either feature alone. The machine-learning findings support our hypothesis H3 by demonstrating that joint non-linear modelling of the spatial image complexity and the temporal allocation of visual attention yields the most robust and consistent differentiation between expertise groups.
Although SVM performance gains were less pronounced for the combined feature set (Appendix B Figure A3, Figure A4 and Figure A5, Table A4, Table A5 and Table A6), this does not contradict the combined-feature advantage, as SVMs optimize global decision margins and may not fully exploit conditional feature interactions that are naturally captured by tree-based models such as Random Forest. This is true regardless of whether the SVM kernel is linear or RBF.
We assert that the aim of the machine learning analysis in our study was not to optimize or compare classifier architectures, but to assess whether joint modelling of 3DFD and fixation duration improves group discrimination relative to single-feature representations. Accordingly, conclusions regarding feature complementarity are based on consistent performance gains observed across stimulus sets and supported by robust classifiers (e.g., Random Forest), rather than on the absolute performance of any single algorithm.

4.3.2. 10-Fold Cross-Validation and Performance Evaluation of Classifiers

Our primary evaluation of classifier performance was based on 10-fold cross-validation. Despite this being the standard practice, to ensure that our findings were not inadvertently dependent on our choice of K = 10, we additionally conducted a robustness analysis by repeating the same participant-level cross-validation with K ∈ {2, 3, …, 9}. The results from this sensitivity check confirmed that performance remained broadly stable across different fold sizes (see Appendix B Figure A3, Figure A4 and Figure A5), supporting the generalizability of the reported results and providing a comprehensive assessment of the discriminative potential of the 3DFD and fixation-duration features.
We adopted a participant-level cross-validation strategy to strengthen the validity of our evaluation. By ensuring that all observations from a single participant were confined to either the training or testing set within each fold, the analysis avoided subject-level data leakage arising from repeated observations per participant. As a result, the reported classification metrics provide a more rigorous assessment of model generalizability, as performance estimates reflect the ability to classify unseen participants rather than repeated samples from the same individual.

4.3.3. Choosing 10 Longest Fixation-Duration Locations

The 10 longest fixation-duration locations were selected to capture fixation events most likely to reflect sustained attentional engagement with the stimulus. This threshold provided a consistent participant–stimulus feature representation for both mixed-effects modelling and machine-learning analysis. To assess whether the findings depended on this threshold, we repeated the main analyses using the top five fixation-duration locations. The broadly similar results obtained for M P s = 10 and M P s = 5 indicate that the principal findings were robust to the selected fixation-count threshold.

4.4. Heterogeneity in Brain MR Images and Formation of Stimulus Groups

The decision to divide the stimuli into multiple groups (A–E, StimP, StimN) was driven by the inherent heterogeneity of pathological brain MR images. Different pathology types, such as tumors, cerebrovascular pathologies, inflammatory changes, and malformations, often present with distinct structural characteristics, making it unlikely that a single aggregated pathology group would capture these meaningful differences. Without this grouping, clinically relevant variability could have been obscured, potentially masking how expertise interacts with specific pathology types and limiting our ability to examine how these differences shape expert visual strategies. It is also important to acknowledge that this study did not account for heterogeneity arising from different imaging planes (e.g., axial, sagittal, coronal), anatomical regions, or acquisition conditions within both normal and pathological images. These factors may further influence gaze behavior and complexity sampling but were beyond the scope of the current analysis. Future work that systematically examines these additional sources of heterogeneity may provide deeper insights into how experts adapt their strategies across diverse imaging contexts.

4.5. Treating Imbalance in the Three Participant Groups and Stimuli Groups

The participant groups in this study were not perfectly balanced, with fewer expert registrars ER (Np = 16) than naïve observers N (Np = 29) or consultant neurosurgeons EC (Np = 24). This imbalance is not uncommon in medical and surgical eye-tracking studies, where recruitment of specialist clinicians is often constrained. For example, Li et al. [19] studied simulated surgical training using 8 experts and 6 novices; Wilson et al. [21] studied laparoscopic expertise using 8 experienced surgeons and 6 novices; and Dalveren and Cagiltay [22] studied eye movements in surgical residents using 14 novices and 9 intermediates. These studies illustrate the practical recruitment constraints commonly encountered when studying expert medical populations.
Acknowledging the participant imbalance, we addressed it analytically by adopting linear mixed-effects models rather than simple group-wise comparisons for analysis. Adopting this modelling approach explicitly accounted for between-participant and between-stimulus variability. Specifically, our models included random intercepts for Subject and StimuliID, thereby preventing observations from the same participant or stimulus from being treated as statistically independent. This is preferable to conventional ANOVA-based approaches, which are less flexible for unbalanced hierarchical eye-tracking data. Our approach is consistent with a prior study conducted by Suman et al. [10], who used linear mixed-effects models in their eye-tracking study and explicitly noted that these models accommodate unbalanced group sizes. Similarly, other prior eye tracking studies by Chen et al. [40] and Bertram et al. [41] also had imbalanced data, and they all used LMMs.
We also acknowledge that participant-group imbalance can influence machine-learning classifiers because the participant–stimulus-level samples inherit the unequal group sizes. Unlike LMMs, Random Forest and SVM classifiers do not automatically remove class imbalance unless explicit reweighting or resampling is applied. Therefore, we do not claim that the classifiers corrected class imbalance. Instead, we reduced the risk of biased interpretation by using participant-level cross-validation and by reporting multiple class-sensitive performance metrics, including sensitivity, specificity, precision, F1 score, and AUC, in addition to accuracy. These metrics allow classifier performance to be evaluated beyond overall accuracy and help identify whether classification performance is driven disproportionately by the majority class. Noting this, a comparison of our overall results with the results from a subset of our participant data where balancing of groups was maintained (Appendix B Figure A6, Table A7, Table A8 and Table A9) confirmed that the conclusions are still the same, i.e., the performance of the combined feature set is higher or closer to the 3DFD feature set.
Notably, the number of stimuli also varied across pathology groups (A = 10, B = 4, C = 3, D = 1, E = 1), reflecting the availability of clinically representative pathological brain MRI cases during our study design. It should be noted that separate mixed-effects models were fitted for each pathology stimulus group by comparing that pathology group with the normal stimulus group. Therefore, although the pathology stimulus groups differed in size, each mixed-effects model evaluated one pathology category against the same normal stimulus group, with stimulus-level variability accounted for using random intercepts for StimuliID. Nevertheless, the findings from the pathology categories containing few stimuli, particularly the inflammatory changes (group D) and malformations (group E), should be interpreted with caution and therefore may be treated as stimulus-specific exploratory findings, rather than definitive pathology-category effects.

4.6. Selecting Fixation Duration as the Primary Behavioral Feature

Among a range of available eye-tracking metrics, fixation duration was selected as the primary behavioral feature in this study due to its direct interpretability as a proxy for attentional allocation and cognitive processing load [4,9,18,19,23]. Importantly, fixation duration can be co-localized with image-derived features (3DFD) at fixation locations, enabling a direct linkage between gaze behavior and underlying image structure. In contrast, temporal sequence-based metrics such as scanpath similarity or saccade dynamics are less suited for such spatially localized analysis. This careful selection of fixation duration allows for a more interpretable and methodologically coherent integration with local image complexity measures.

4.7. Other Methodologies Used for the Characterization of Visual Attention

Alternate methodologies such as fMRI, EEG, and MEG have been employed in prior work to identify the neural substrates of attention and test cognitive theories of attention [1]. However, these approaches require lengthy setup and computationally intensive processing, making them less feasible for studies involving time-constrained clinical experts. Eye-tracking, by contrast, offers a portable, cost-effective, and computationally efficient alternative for studying visual attention. Given the practical challenges of recruiting neurosurgeons with demanding schedules, using eye-tracking cameras was the most suitable choice for our focused study on the neurosurgery experts.

4.8. Data Collection and Sampling Rate

A key strength of this study is the use of a high eye-tracking sampling rate, which is critical for accurate fixation detection and feature computation. Sampling rate directly influences the precision of fixation-based measures, as lower rates can introduce temporal discretization errors and compromise the reliability of derived metrics such as the number of fixation points [42]. Different studies have used different sampling rates of eye-trackers for data collection. Suman et al. [10] used a sampling rate of 250 Hz for their eye-tracking data collection. In their eye-tracking study with neuroradiologists, Crowe et al. [16] used Eye-Link 1000 (SR Research, Mississauga, ON, Canada), which had a high-sampling rate of 1000 Hz. Like Crowe et al. [16], our study also adopted a high-sampling rate of 1000 Hz. This choice ensures superior temporal resolution and stability in our fixation-duration-based analyses.
Importantly, our eye-tracker combined this high sampling rate with portability, enabling data collection in diverse environments, including laboratory settings and external venues such as conference desk spaces. This methodological advantage allowed us to maintain data quality without sacrificing logistical flexibility, a critical consideration when working with time-constrained neurosurgical experts. By reducing temporal discretization error in fixation detection, our approach strengthens the robustness and generalizability of the reported findings [10,22,23,42].

4.9. Translational Implications for Developing Expert-like Systems Based on Neurosurgical Visual Expertise

A central motivation of this study was to identify neurosurgery-specific gaze behaviors that could inform the development of computational systems capable of exhibiting “expert-like” visual strategies. To make this link explicit, we synthesize here the most salient empirical findings and outline how each can be operationalized in the design of intelligent models for neurosurgical image interpretation.
First, use expert fixation duration as a behavior-grounded supervisory signal to train attention: models should up-weight features from regions analogous to those attracting experts’ longest fixations when pathology is present, operationalizing the observed Group × ImageType effect where experts dwell longer on abnormal images. Practically, precompute a 3DFD map for each image, extract features at the top- M P s fixation locations (5–10) to capture the most informative scrutiny, and associate these with patch embeddings during training; this retains fidelity to our data pipeline while remaining computationally tractable.
Second, design models to fuse temporal allocation and structural complexity, because these signals are complementary and improve discrimination beyond either alone. Concretely, adopt dual-branch architectures (image branch + fixation/3DFD branch) with a fusion or gating layer that learns non-linear interactions; our Random Forest classifiers showed consistently higher accuracy and specificity with the combined feature set across tasks and stimulus groups, justifying this integration.
Finally, make attention pathology-aware: learn category-specific complexity priors so that 3DFD influences attention only when diagnostically relevant (e.g., higher weighting for inflammatory changes, lower for certain cerebrovascular cases), reflecting the stimulus-dependent expertise effects in our 3DFD analyses. Together, these steps translate our empirical markers of selective fixation prolongation, complexity-dependent sampling, and their synergy into concrete modelling choices for systems that attend and prioritize like neurosurgical experts.
As an alternative expert-like multi-agent (“mixture-of-expertise”) architecture, researchers can train stimulus-specific classifiers aligned with our pathology groupings (A–E, e.g., tumors, cerebrovascular, others, inflammatory, malformations) and then weight or aggregate their outputs into a final expertise score for the observer. This design is directly supported by our results: per-group Random Forest classifiers routinely achieve higher performance than models trained on the combined “StimP” set (e.g., N vs. EC: group A/B/C accuracies ≈ 94–99% vs. 80.9% on StimP; N vs. ER: group A/B/C ≈ 93–96% vs. 87.1% on StimP; EC vs. ER: group A/B/C ≈ 95–97% vs. 87.4% on StimP), indicating that combining heterogeneous stimuli dampens discrimination, whereas category-specific specialization preserves expert-like signals. A multi-agent system can therefore mirror a neurosurgeon’s pathology-dependent strategies by letting each agent specialize on its image type and fusing their calibrated outputs (e.g., via learned weights or Bayesian ensembling) into a single, robust expertise estimate.
As a general note, our framework is designed to classify the observer’s level of visual expertise (e.g., N vs. EC/ER), and not to classify images as normal versus pathological. Any model derived from this work should therefore report expertise level or expert-likeness of viewing behavior, rather than diagnostic labels for the input image. These translational directions could be incorporated into future neurosurgical training simulators by providing objective feedback on whether a trainee’s fixation-duration patterns and fixation-linked complexity sampling resemble those of more experienced neurosurgeons. In this context, the proposed framework may support observer-level assessment of visual expertise rather than image-level diagnostic classification.

4.10. Limitations and Future Work

Although the study demonstrates consistent methodological implementation and clear behavioral patterns, several limitations warrant consideration. First, the 3DFD measure captures local image complexity but does not encode higher-level structural or anatomical relevance. Incorporating multi-scale fractal features or complementary descriptors, such as entropy-based measures, may further strengthen the characterization of visual behavior. Moreover, because 3DFD values were computed only at fixation locations rather than across the full image, the reported ImageType and Group × ImageType effects should be interpreted as reflecting differences in the structural complexity of sampled regions, rather than as global measures of image complexity. Second, the expert consultant (EC) and expert registrar (ER) groups differ in experience level, which may give rise to distinct trajectories of search efficiency and attentional control that are not fully captured in the present cross-sectional design. Longitudinal tracking of surgical trainees would enable a more direct examination of how fixation behavior and complexity sampling evolve with increasing expertise.
Future work, including integrating temporal scanpath measures, such as fixation transitions, sequence structure, and gaze entropy, could provide a more comprehensive account of expertise by jointly modelling where observers look, for how long, and in what order. Furthermore, the integrated spatial–temporal framework developed in our study could benefit and extend emerging machine-learning approaches that leverage expert gaze for adaptive or explainable decision-making, such as eye-guided multimodal fusion frameworks [43] that fuse imaging with expert fixation maps to improve interpretability and reliability, and machine-learning systems [44] designed to discriminate radiologists’ experience levels using spatiotemporal eye-tracking encodings.
Future work may also extend this framework to volumetric MRI datasets to examine whether fixation-linked complexity measures derived from full 3D anatomical reconstructions provide additional information beyond the 2D slice-based analysis used here. Such work should also evaluate how the number of MRI slices/segments, slice thickness, inter-slice spacing, and anatomical coverage influence volumetric 3DFD estimation and its usefulness for neurosurgical visualization and training.
Future studies could further extend this framework to immersive virtual-reality environments by visualizing healthy and pathological volumetric MRI data in 3D using head-mounted displays with integrated eye tracking. Such a framework could enable fixation-linked analysis of how neurosurgeons inspect anatomical and pathological structures from multiple viewpoints and may provide additional insight into visual expertise during 3D neurosurgical image interpretation and training.
Also, while the mixed-effects framework accounts for the imbalance in stimuli, findings for smaller stimulus groups, such as groups D (Inflammatory changes) and E (Malformations), should be validated in future studies with larger stimulus datasets.
In addition, future work could incorporate model-explainability analyses such as Gini index, Boruta, or SHAP to quantify the contribution of individual fixation- and complexity-based features to classifier decisions, thereby providing deeper insight into feature-level relevance within expert-like models. Finally, future work should also evaluate the applicability of the translational implications of the integrated spatial–temporal framework developed in this study to other medical imaging domains such as CT, angiography, ultrasound, and other clinical specialties such as ophthalmology, ENT and radiology, where visual expertise may depend on different image structures and task demands.

5. Conclusions

The convergence of mixed-effects modelling and machine-learning analyses indicates that fixation duration and geometric complexity (3DFD) capture complementary aspects of gaze behavior associated with visual expertise. While 3DFD reflects the structural properties of image regions sampled during viewing, fixation duration indexes the temporal allocation of attention to diagnostically relevant areas. Importantly, expertise-related differences were most consistently characterized when these spatial and temporal features were considered jointly rather than in isolation. Quantitatively, joint modelling of fixation duration and 3DFD yielded high discriminative performance, with supervised Random Forest classifiers achieving 93–100% accuracy across pathology-specific stimulus groups and approximately 80–87% accuracy for the aggregated stimuli dataset. The agreement between inferential mixed-effects results and predictive classification performance supports an integrated spatial–temporal feature framework as a principled approach for characterizing visual expertise in medical image interpretation.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/jemr19030062/s1, Supplementary File contains MATLAB and Python code, data files, and stimuli images used in this study. It includes a Word document describing folder contents and usage instructions.

Author Contributions

Conceptualization, A.D.I. and P.K.; methodology, P.K.; software, P.K. and C.R.; formal analysis, P.K.; investigation, P.K.; resources, A.D.I.; data curation, P.K.; writing—original draft preparation, P.K.; writing—review and editing, G.A., C.R. and A.D.I.; visualization, P.K.; supervision, A.D.I.; project administration, A.D.I. and P.K.; funding acquisition, A.D.I. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by Macquarie University and an Australian Research Council (ARC) Future Fellowship granted to Prof Antonio Di Ieva in 2019 (FT190100623).

Institutional Review Board Statement

Ethical approval for this study was granted by the Macquarie University Medicine and Health Sciences Subcommittee (reference number 52021919128177, approved on 18 May 2021). As per ethics approval, this study meets the requirements set out in the National Statement on Ethical Conduct in Human Research 2007 (updated July 2018), which embodies the ethical principles of the Declaration of Helsinki 1964. As per the ethics approval, this study falls within the category of “Clinical research that is not a clinical trial—observational study”.

Informed Consent Statement

Written consent was obtained from all subjects involved in the study.

Data Availability Statement

Data, along with the data processing pipeline codes, are contained within the Supplementary Materials. The original contributions presented in this study are included in the article/Appendix material. Further inquiries can be directed to the corresponding authors.

Acknowledgments

We would like to thank all the participants for their valuable time spent participating in this study. The data collection work was carried out at Macquarie University at the Computational NeuroSurgery (CNS) Lab and the Neurosurgery Unit at Macquarie University Hospital. Additional data were collected at the Neurosurgical Society of Australasia (NSA) 77th Annual Scientific Meeting, Sydney, Australia, 14–16 September 2022.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
FDFractal dimension
3DFDThree-dimensional fractal dimension
MRIMagnetic resonance imaging
NNaïve
ERExpert registrar
ECExpert consultant
NNNaïve normal
NPNaïve pathology
ECNExpert consultant normal
ECPExpert consultant pathology
ERNExpert registrar normal
ERPExpert registrar pathology
AUCArea under curve

Appendix A

Appendix A.1. Eye-Tracking Report Template

Table A1. A template of the report containing the fixation duration values at different fixation locations for participant ID N02. The total number of fixation points in this template report is N P s = N 2 1 by considering the subject number P = 2 and the stimulus number for this participant being s = 1.
Table A1. A template of the report containing the fixation duration values at different fixation locations for participant ID N02. The total number of fixation points in this template report is N P s = N 2 1 by considering the subject number P = 2 and the stimulus number for this participant being s = 1.
Participant IDStimulus NameX Coordinate of
Fixation
Y Coordinate of
Fixation
Fixation Duration at the (X, Y) Coordinates
N02BRAIN_1_PATH_Meningioma_ax.png x 1 y 1 F i x D u r 1
N02BRAIN_1_PATH_Meningioma_ax.png x 2 y 2 F i x D u r 2
N02BRAIN_1_PATH_Meningioma_ax.png x N 2 1 y N 2 1 F i x D u r N 2 1

Appendix A.2. Pathological Stimuli

Figure A1. MR images used as pathological stimuli in this study. Stimulus groups A–E correspond to the stimulus groups with pathology types of tumors, cerebrovascular pathologies, others, inflammatory changes, and malformations. The name of the pathology is shown for each of the stimulus images.
Figure A1. MR images used as pathological stimuli in this study. Stimulus groups A–E correspond to the stimulus groups with pathology types of tumors, cerebrovascular pathologies, others, inflammatory changes, and malformations. The name of the pathology is shown for each of the stimulus images.
Jemr 19 00062 g0a1aJemr 19 00062 g0a1b

Appendix A.3. Eye-Tracking Image Data Processing

Masks for Removing Background Pixels in Stimuli Images

A method was implemented in Matlab to create masks by manually clicking on the stimulus image. Using this method, masks were manually created from each of the 39 stimulus images and applied to the heat maps of fixation duration and the 3DFD maps. The masks ensured that the background pixels in the stimulus image were removed from analysis.

Appendix A.4. Software Programming Languages Used for Data Processing

Appendix A.4.1. Linear Mixed-Effects Modelling and Statistical Analysis

All the data processing for statistical analysis, such as data preparation, processing, mixed-effects modeling, and effect size computation, was performed in MATLAB R2025b using the Statistics and Machine Learning Toolbox. The numerical values in the statistical analysis results tables were generated programmatically to ensure reproducibility across all the stimulus groups.

Appendix A.4.2. Machine Learning Models

All the data preparation for feeding into the three machine learning models and the development of the three models were performed in MATLAB R2025b. The numerical values in the tables summarizing the machine learning results were generated programmatically to ensure reproducibility across all the stimulus groups and the three binary classifiers.

Appendix B

Appendix B.1. Heatmap of Fixation Duration

Figure A2. Qualitative comparison of the heatmaps of fixation duration for naïve, consultant and registrar groups. (a) One of the MRI structural images taken from stimulus group A – Tumors, and used as stimuli with pathology (Glioblastoma) highlighted in the red circle. (bd) Heat maps of fixation duration averaged over all the participants in a group and superimposed on the structural MRI for (b) naïve, (c) consultant, and (d) registrar groups. The fixation duration values below 50 ms are not displayed in (bd). Note that the red circle shown in (a) is not present in the stimulus when it is viewed by the participants during their eye-tracking data collection session.
Figure A2. Qualitative comparison of the heatmaps of fixation duration for naïve, consultant and registrar groups. (a) One of the MRI structural images taken from stimulus group A – Tumors, and used as stimuli with pathology (Glioblastoma) highlighted in the red circle. (bd) Heat maps of fixation duration averaged over all the participants in a group and superimposed on the structural MRI for (b) naïve, (c) consultant, and (d) registrar groups. The fixation duration values below 50 ms are not displayed in (bd). Note that the red circle shown in (a) is not present in the stimulus when it is viewed by the participants during their eye-tracking data collection session.
Jemr 19 00062 g0a2

Appendix B.2. Linear Mixed-Effects Modelling and Statistical Analysis

Appendix B.2.1. H1: Expertise-Related Differences in Fixation-Duration Behavior

Table A2. Fixed-effects results for Fixation Duration × ImageType across stimulus groups (A–E, and StimP). For these results, we chose the M P s = 5, which corresponds to the first 5 maximum fixation durations for each stimulus for each participant.
Table A2. Fixed-effects results for Fixation Duration × ImageType across stimulus groups (A–E, and StimP). For these results, we chose the M P s = 5, which corresponds to the first 5 maximum fixation durations for each stimulus for each participant.
Stimulus GroupPredictorβSEtdfpp_FDRCohen’s d
AImageType
(tumors)
−8.0588.898−0.910,0440.370.365−0.038
AEC × ImageType (tumors)29.72110.492.8910,0440.010.01 *0.139
AER × ImageType (tumors)73.45611.566.3610,0440.0000.000 ***0.343
BImageType
(cerebrovascular pathologies)
−7.96512.039−0.6680340.5080.508−0.039
BEC × ImageType (cerebrovascular pathologies)52.07714.0473.7180340.0000.000 ***0.257
BER × ImageType (cerebrovascular pathologies)76.53015.4724.9580340.0000.000 ***0.377
CImageType (others)1.43814.3370.1076990.9200.9200.007
CEC × ImageType (others)37.06215.9242.3376990.0200.040 *0.182
CER × ImageType (others)46.65317.5402.6676990.0080.024 *0.229
DImageType
(inflammatory)
25.42721.4101.1970290.2350.4370.119
DEC × ImageType (inflammatory)17.79227.7230.6470290.5210.5210.083
DER × ImageType (inflammatory)26.91430.5360.8870290.3780.4540.126
EImageType (malfomations)−20.53920.958−0.9870290.3270.327−0.101
EEC × ImageType (malfomations)39.43026.2911.5070290.1340.2010.194
EER × ImageType (malfomations)82.70528.9592.8670290.0040.012 *0.408
StimPImageType
(all pathologies)
−5.4367.913−0.713,0590.4920.492−0.025
StimPEC × ImageType (all pathologies)35.4728.8724.0013,0590.0000.000 ***0.169
StimPER × ImageType (all pathologies)67.9119.7726.9513,0590.0000.000 ***0.318
Note. Fixed-effects estimates are from linear mixed-effects models of the formula Fixation Duration ~ Group × ImageType + (1|Subject) + (1|StimuliID). The naïve group (N) served as the reference category, and normal images served as the reference level for ImageType. β = fixation duration estimate; d = Cohen’s d. Significant p and p_FDR values (p/p_FDR < 0.05) are shown in bold. Significance levels on FDR-corrected p-values (p_FDR) are marked as: p_FDR < 0.05 (*), p_FDR < 0.001 (***).

Appendix B.2.2. H2: Expertise-Related Differences in Structural Complexity (3DFD) at Fixation Locations For M P s = 5 Maximum Fxation Durations

Table A3. Fixed-Effects Results for 3DFD × ImageType across stimulus groups (A–E, and StimP). For these results, we chose the M P s = 5, which corresponds to the first 5 maximum fixation durations for each stimulus for each participant.
Table A3. Fixed-Effects Results for 3DFD × ImageType across stimulus groups (A–E, and StimP). For these results, we chose the M P s = 5, which corresponds to the first 5 maximum fixation durations for each stimulus for each participant.
Stimulus GroupPredictorβSEtdfpp_FDRCohen’s d
AImageType
(tumors)
0.0310.0102.9610,0440.0090.009 **0.386
AEC × ImageType (tumors)−0.0020.004−0.6310,0440.530.5320.025
AER × ImageType (tumors)−0.0030.004−0.7810,0440.430.52−0.037
BImageType
(cerebrovascular pathologies)
0.0250.0151.6880340.0920.1100.355
BEC × ImageType (cerebrovascular pathologies)−0.0120.005−2.4280340.0160.026 *−0.171
BER × ImageType (cerebrovascular pathologies)−0.0180.005−3.2880340.0010.003 **−0.256
CImageType
(others)
0.0040.0160.2476990.8080.8080.058
CEC × ImageType (others)−0.0050.005−0.8476990.3980.478−0.072
CER × ImageType (others)−0.0060.006−0.9676990.3380.478−0.087
DImageType
(inflammatory)
−0.0030.028−0.1170290.9090.909−0.034
DEC × ImageType (inflammatory)0.0410.0113.6170290.0000.000 ***0.467
DER × ImageType (inflammatory)0.0540.0134.2970290.0000.000 ***0.615
EImageType (malfomations)0.0890.0283.2470290.0010.003 **1.239
EEC × ImageType (malfomations)0.0050.0090.5270290.6060.6490.070
EER × ImageType (malfomations)0.0050.0100.4570290.6490.6490.070
StimPImageType
(all pathologies)
0.0270.0092.9313,0590.0030.009 **0.327
StimPEC × ImageType (all pathologies)−0.0020.003−0.6213,0590.5330.533−0.024
StimPER × ImageType (all pathologies)−0.0030.004−0.913,0590.3670.44−0.036
Note. Fixed-effects estimates are from linear mixed-effects models of the formula Mean3DFD ~ Group × ImageType + (1|Subject) + (1|StimuliID). The naïve group (N) served as the reference category, and normal images served as the reference level for ImageType. β = unstandardized 3DFD estimate; d = Cohen’s d. Significant p and p_FDR values (p/p_FDR < 0.05) are shown in bold. Significance levels on FDR-corrected p-values (p_FDR) are marked as: p_FDR < 0.05 (*), p_FDR < 0.01 (**), p_FDR < 0.001 (***).

Appendix B.3. Non-Linear Modelling Using Supervised Machine Learning Methods

Appendix B.3.1. H3: Unbalanced Participant Groups of N (Np = 29), EC (Np = 24), ER (Np = 16)

Figure A3. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets. Panels correspond to stimulus sets A (tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant. RBF kernel is used in SVM.
Figure A3. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets. Panels correspond to stimulus sets A (tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant. RBF kernel is used in SVM.
Jemr 19 00062 g0a3
Classifiers trained on the combined feature representation generally achieve higher accuracy and specificity across stimulus sets. Performance differences between Random Forest and SVM classifiers are consistent with the N vs. EC and N vs. ER classifiers for the combined feature space. The performance differences between Random Forest and SVM classifiers are consistent with the interaction-driven structure of the combined feature space, with tree-based models more effectively capturing non-linear interactions between geometric complexity and fixation behavior.
Figure A4. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets, when the data corresponding to the first 5 longest fixation durations are used. Panels correspond to stimulus sets A(tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 5, which corresponds to the first 5 maximum fixation durations for each stimulus for each participant. RBF kernel is used in SVM.
Figure A4. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets, when the data corresponding to the first 5 longest fixation durations are used. Panels correspond to stimulus sets A(tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 5, which corresponds to the first 5 maximum fixation durations for each stimulus for each participant. RBF kernel is used in SVM.
Jemr 19 00062 g0a4aJemr 19 00062 g0a4b
Classifiers trained on the combined feature representation generally achieve higher accuracy and specificity across stimulus sets. Performance differences between Random Forest and SVM classifiers are consistent with the N vs. EC and N vs. ER classifiers for the combined feature space. The performance differences between Random Forest and SVM classifiers are consistent with the interaction-driven structure of the combined feature space, with tree-based models more effectively capturing non-linear interactions between geometric complexity and fixation behavior.
Figure A5. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets, when the first 10 longest fixation durations are used. Panels correspond to stimulus sets A (tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant. A linear kernel is used in SVM.
Figure A5. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets, when the first 10 longest fixation durations are used. Panels correspond to stimulus sets A (tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant. A linear kernel is used in SVM.
Jemr 19 00062 g0a5aJemr 19 00062 g0a5b
Classifiers trained on the combined feature representation generally achieve higher accuracy and specificity across stimulus sets. Performance differences between Random Forest and SVM classifiers are consistent with the N vs. EC and N vs. ER classifiers for the combined feature space. The performance differences between Random Forest and SVM classifiers are consistent with the interaction-driven structure of the combined feature space, with tree-based models more effectively capturing non-linear interactions between geometric complexity and fixation behavior.
Table A4. Classification performance of SVM classifier for N vs. EC using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Table A4. Classification performance of SVM classifier for N vs. EC using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Stimulus GroupFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD75.53 ± 13.640.97 ± 0.060.61 ± 0.260.67 ± 0.180.73 ± 0.140.86 ± 0.10
AFixation Duration50.67 ± 14.820.21 ± 0.160.81 ± 0.170.47 ± 0.230.43 ± 0.120.59 ± 0.13
ACombined56.67 ± 18.660.00 ± 0.001.00 ± 0.000.00 ± 0.000.35 ± 0.080.82 ± 0.13
B3DFD84.58 ± 13.711.00 ± 0.000.69 ± 0.240.76 ± 0.200.82 ± 0.140.97 ± 0.04
BFixation Duration70.83 ± 20.260.85 ± 0.280.66 ± 0.230.68 ± 0.180.69 ± 0.200.84 ± 0.14
BCombined54.67 ± 16.870.10 ± 0.320.95 ± 0.160.03 ± 0.110.37 ± 0.110.96 ± 0.06
C3DFD78.78 ± 17.260.99 ± 0.030.63 ± 0.310.72 ± 0.240.77 ± 0.180.94 ± 0.08
CFixation Duration58.78 ± 25.470.67 ± 0.460.68 ± 0.290.40 ± 0.320.55 ± 0.280.77 ± 0.20
CCombined54.00 ± 16.760.10 ± 0.320.94 ± 0.180.03 ± 0.090.36 ± 0.100.88 ± 0.15
D3DFD86.00 ± 18.971.00 ± 0.000.76 ± 0.330.80 ± 0.230.84 ± 0.220.98 ± 0.05
DFixation Duration82.33 ± 14.741.00 ± 0.000.67 ± 0.320.73 ± 0.200.79 ± 0.200.97 ± 0.11
DCombined70.33 ± 23.750.90 ± 0.320.61 ± 0.340.60 ± 0.300.67 ± 0.260.97 ± 0.11
E3DFD88.00 ± 13.980.93 ± 0.170.85 ± 0.200.86 ± 0.190.87 ± 0.150.97 ± 0.07
EFixation Duration70.33 ± 23.750.90 ± 0.320.62 ± 0.290.60 ± 0.300.69 ± 0.250.93 ± 0.12
ECombined60.00 ± 24.940.60 ± 0.520.75 ± 0.330.38 ± 0.390.54 ± 0.281.00 ± 0.00
StimP3DFD58.21 ± 18.810.55 ± 0.300.71 ± 0.190.58 ± 0.140.56 ± 0.180.69 ± 0.10
StimPFixation Duration53.04 ± 13.040.26 ± 0.150.80 ± 0.110.49 ± 0.170.47 ± 0.100.58 ± 0.07
StimPCombined56.67 ± 18.660.00 ± 0.001.00 ± 0.000.00 ± 0.000.35 ± 0.080.66 ± 0.10
StimN3DFD58.40 ± 20.010.48 ± 0.310.78 ± 0.210.65 ± 0.220.55 ± 0.200.71 ± 0.12
StimNFixation Duration56.67 ± 15.930.23 ± 0.150.85 ± 0.060.52 ± 0.170.48 ± 0.130.54 ± 0.10
StimNCombined56.67 ± 18.660.00 ± 0.001.00 ± 0.000.00 ± 0.000.35 ± 0.080.68 ± 0.10
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.
Table A5. Classification performance of SVM classifier for N vs. ER using M P s = 10 longest fixation-duration locations (10-fold cross-validation).
Table A5. Classification performance of SVM classifier for N vs. ER using M P s = 10 longest fixation-duration locations (10-fold cross-validation).
Stimulus SetFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD64.60 ± 22.300.06 ± 0.110.98 ± 0.040.24 ± 0.420.43 ± 0.150.87 ± 0.10
AFixDur66.70 ± 20.280.08 ± 0.080.98 ± 0.040.53 ± 0.500.45 ± 0.100.80 ± 0.12
ACombined65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.85 ± 0.11
B3DFD67.62 ± 25.080.53 ± 0.450.86 ± 0.210.53 ± 0.430.60 ± 0.270.96 ± 0.03
BFixDur66.12 ± 23.040.03 ± 0.081.00 ± 0.000.20 ± 0.420.41 ± 0.120.91 ± 0.11
BCombined65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.92 ± 0.11
C3DFD60.83 ± 19.710.00 ± 0.000.95 ± 0.100.00 ± 0.000.37 ± 0.080.90 ± 0.14
CFixDur65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.81 ± 0.17
CCombined65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.89 ± 0.11
D3DFD78.00 ± 22.390.60 ± 0.470.92 ± 0.140.65 ± 0.470.71 ± 0.290.94 ± 0.17
DFixDur60.00 ± 20.000.30 ± 0.480.85 ± 0.220.13 ± 0.220.44 ± 0.181.00 ± 0.00
DCombined65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.091.00 ± 0.00
E3DFD81.50 ± 27.990.70 ± 0.480.94 ± 0.120.65 ± 0.470.76 ± 0.331.00 ± 0.00
EFixDur65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.94 ± 0.13
ECombined65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.94 ± 0.12
StimP3DFD64.50 ± 22.880.01 ± 0.030.99 ± 0.020.10 ± 0.320.39 ± 0.100.71 ± 0.19
StimPFixDur65.82 ± 20.090.10 ± 0.050.97 ± 0.030.62 ± 0.330.46 ± 0.100.67 ± 0.09
StimPCombined65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.75 ± 0.16
StimN3DFD64.75 ± 23.100.01 ± 0.021.00 ± 0.010.10 ± 0.320.39 ± 0.100.75 ± 0.14
StimNFixDur62.42 ± 20.960.02 ± 0.020.96 ± 0.030.25 ± 0.340.39 ± 0.090.55 ± 0.09
StimNCombined65.00 ± 23.210.00 ± 0.001.00 ± 0.000.00 ± 0.000.38 ± 0.090.72 ± 0.16
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.
Table A6. Classification performance of SVM classifier for EC vs. ER using M P s = 10 longest fixation-duration locations (10-fold cross-validation).
Table A6. Classification performance of SVM classifier for EC vs. ER using M P s = 10 longest fixation-duration locations (10-fold cross-validation).
Stimulus SetFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD64.67 ± 27.630.57 ± 0.480.61 ± 0.320.58 ± 0.440.59 ± 0.310.96 ± 0.06
AFixDur51.50 ± 22.350.34 ± 0.370.61 ± 0.300.52 ± 0.420.45 ± 0.210.82 ± 0.11
ACombined43.92 ± 31.070.00 ± 0.000.75 ± 0.420.00 ± 0.000.28 ± 0.160.93 ± 0.09
B3DFD82.71 ± 20.300.70 ± 0.480.70 ± 0.310.62 ± 0.450.71 ± 0.290.97 ± 0.03
BFixDur47.29 ± 23.220.20 ± 0.350.69 ± 0.340.33 ± 0.450.38 ± 0.220.91 ± 0.12
BCombined38.12 ± 22.350.00 ± 0.000.69 ± 0.410.00 ± 0.000.26 ± 0.120.94 ± 0.07
C3DFD70.83 ± 28.930.61 ± 0.510.65 ± 0.310.59 ± 0.440.64 ± 0.330.98 ± 0.04
CFixDur50.28 ± 30.830.37 ± 0.480.62 ± 0.340.30 ± 0.410.45 ± 0.320.98 ± 0.02
CCombined39.17 ± 22.240.00 ± 0.000.70 ± 0.400.00 ± 0.000.26 ± 0.120.99 ± 0.02
D3DFD85.00 ± 17.480.67 ± 0.470.74 ± 0.330.62 ± 0.460.70 ± 0.251.00 ± 0.00
DFixDur54.17 ± 30.240.35 ± 0.470.68 ± 0.330.32 ± 0.430.47 ± 0.320.88 ± 0.21
DCombined40.00 ± 22.500.10 ± 0.320.68 ± 0.380.05 ± 0.160.30 ± 0.191.00 ± 0.00
E3DFD91.67 ± 13.610.70 ± 0.480.81 ± 0.320.65 ± 0.470.76 ± 0.270.94 ± 0.14
EFixDur56.67 ± 30.880.40 ± 0.520.66 ± 0.310.28 ± 0.390.48 ± 0.310.83 ± 0.26
ECombined42.50 ± 22.030.10 ± 0.320.70 ± 0.360.05 ± 0.160.31 ± 0.190.94 ± 0.14
StimP3DFD47.24 ± 27.860.32 ± 0.420.62 ± 0.330.38 ± 0.420.43 ± 0.290.88 ± 0.11
StimPFixDur49.96 ± 21.580.14 ± 0.130.72 ± 0.320.54 ± 0.460.40 ± 0.170.66 ± 0.09
StimPCombined59.17 ± 35.450.00 ± 0.000.90 ± 0.320.00 ± 0.000.34 ± 0.160.86 ± 0.12
StimN3DFD44.67 ± 24.840.26 ± 0.370.63 ± 0.340.37 ± 0.410.39 ± 0.260.87 ± 0.11
StimNFixDur47.37 ± 22.350.05 ± 0.050.74 ± 0.310.47 ± 0.430.34 ± 0.140.64 ± 0.12
StimNCombined59.17 ± 35.450.00 ± 0.000.90 ± 0.320.00 ± 0.000.34 ± 0.160.84 ± 0.11
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.

Appendix B.3.2. H3: Balanced Participant Groups of N (Np = 16), EC (Np = 16), ER (Np = 16)

Figure A6. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets, with an equal number of participants (Np = 16). Panels correspond to stimulus sets A (tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant. RBF kernel is used in SVM.
Figure A6. Classification performance for (a) N vs. EC, (b) N vs. ER, and (c) EC vs. ER, using Random Forest (RF) and SVM classifiers across stimulus sets, with an equal number of participants (Np = 16). Panels correspond to stimulus sets A (tumors), B (cerebrovascular pathologies), C (others), D (inflammatory), E (malformations), StimP (all pathologies taken together), and StimN (normal stimuli). For each stimulus set, classifier performance is shown for fixation duration only, 3DFD only, and the combined feature representation. Each box plot in the panels corresponds to the distributions of classifier accuracies for K = [2,3,4,5,6,7,8,9,10] in the K-fold cross-validation. For these results, we chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant. RBF kernel is used in SVM.
Jemr 19 00062 g0a6
In agreement with the results in Figure A3 for the unbalanced participant groups, for the balanced group (Np = 16), the classifiers trained on the combined feature representation achieved either higher or very close accuracy and other performance metrics across the stimulus sets for all the Random Forest binary classifiers.
Table A7. Classification performance of RBF classifier for N vs. EC using M P s = 10 longest fixation-duration locations (10-fold cross-validation), with balanced data for each class (n = 16).
Table A7. Classification performance of RBF classifier for N vs. EC using M P s = 10 longest fixation-duration locations (10-fold cross-validation), with balanced data for each class (n = 16).
Stimulus SetFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD86.17 ± 18.270.85 ± 0.300.70 ± 0.410.75 ± 0.360.75 ± 0.240.96 ± 0.05
AFixDur59.67 ± 11.280.59 ± 0.260.52 ± 0.310.60 ± 0.330.53 ± 0.130.70 ± 0.14
ACombined86.00 ± 20.570.87 ± 0.310.67 ± 0.450.74 ± 0.370.73 ± 0.270.96 ± 0.05
B3DFD95.83 ± 10.580.89 ± 0.310.85 ± 0.340.85 ± 0.340.86 ± 0.220.99 ± 0.04
BFixDur84.58 ± 18.400.76 ± 0.300.74 ± 0.420.81 ± 0.340.75 ± 0.260.94 ± 0.13
BCombined96.67 ± 10.540.90 ± 0.320.85 ± 0.340.85 ± 0.340.87 ± 0.221.00 ± 0.01
C3DFD95.56 ± 9.370.88 ± 0.310.83 ± 0.320.84 ± 0.320.85 ± 0.210.99 ± 0.04
CFixDur85.56 ± 10.940.71 ± 0.310.81 ± 0.330.78 ± 0.340.75 ± 0.190.93 ± 0.13
CCombined98.89 ± 3.510.90 ± 0.320.88 ± 0.310.88 ± 0.320.89 ± 0.211.00 ± 0.00
D3DFD100.00 ± 0.000.90 ± 0.320.90 ± 0.320.90 ± 0.320.90 ± 0.211.00 ± 0.00
DFixDur96.67 ± 10.540.85 ± 0.340.90 ± 0.320.90 ± 0.320.87 ± 0.221.00 ± 0.00
DCombined96.67 ± 10.540.85 ± 0.340.90 ± 0.320.90 ± 0.320.87 ± 0.221.00 ± 0.00
E3DFD100.00 ± 0.000.90 ± 0.320.90 ± 0.320.90 ± 0.320.90 ± 0.211.00 ± 0.00
EFixDur96.67 ± 10.540.85 ± 0.340.90 ± 0.320.90 ± 0.320.87 ± 0.221.00 ± 0.00
ECombined100.00 ± 0.000.90 ± 0.320.90 ± 0.320.90 ± 0.320.90 ± 0.211.00 ± 0.00
StimP3DFD69.65 ± 15.130.77 ± 0.280.51 ± 0.350.63 ± 0.330.60 ± 0.180.81 ± 0.13
StimPFixDur58.60 ± 5.540.55 ± 0.220.53 ± 0.250.59 ± 0.290.53 ± 0.110.67 ± 0.10
StimPCombined72.94 ± 16.480.79 ± 0.290.55 ± 0.370.67 ± 0.340.64 ± 0.210.84 ± 0.16
StimN3DFD72.46 ± 18.300.75 ± 0.280.59 ± 0.360.66 ± 0.350.64 ± 0.220.82 ± 0.21
StimNFixDur50.54 ± 14.480.49 ± 0.230.47 ± 0.260.55 ± 0.310.47 ± 0.170.56 ± 0.19
StimNCombined74.04 ± 18.850.78 ± 0.280.58 ± 0.390.67 ± 0.350.65 ± 0.240.79 ± 0.26
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.
Table A8. Classification performance of RBF classifier for N vs. ER using M P s = 10 longest fixation-duration locations (10-fold cross-validation), with balanced data for each class (n = 16).
Table A8. Classification performance of RBF classifier for N vs. ER using M P s = 10 longest fixation-duration locations (10-fold cross-validation), with balanced data for each class (n = 16).
Stimulus SetFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD86.75 ± 17.070.83 ± 0.300.72 ± 0.380.75 ± 0.360.75 ± 0.210.93 ± 0.11
AFixDur74.67 ± 15.630.70 ± 0.290.68 ± 0.380.72 ± 0.330.67 ± 0.210.89 ± 0.12
ACombined86.92 ± 17.990.86 ± 0.310.70 ± 0.420.75 ± 0.360.75 ± 0.250.98 ± 0.03
B3DFD97.50 ± 5.620.89 ± 0.310.88 ± 0.320.87 ± 0.320.88 ± 0.211.00 ± 0.00
BFixDur80.21 ± 16.170.71 ± 0.340.70 ± 0.420.73 ± 0.380.68 ± 0.240.89 ± 0.12
BCombined97.50 ± 7.910.90 ± 0.320.86 ± 0.330.86 ± 0.330.87 ± 0.211.00 ± 0.00
C3DFD88.61 ± 14.010.82 ± 0.310.76 ± 0.380.78 ± 0.340.77 ± 0.200.98 ± 0.04
CFixDur84.44 ± 12.780.77 ± 0.310.78 ± 0.350.79 ± 0.330.76 ± 0.200.97 ± 0.06
CCombined91.67 ± 11.420.87 ± 0.310.77 ± 0.360.80 ± 0.330.80 ± 0.201.00 ± 0.01
D3DFD96.67 ± 10.540.87 ± 0.320.90 ± 0.320.90 ± 0.320.89 ± 0.231.00 ± 0.00
DFixDur96.67 ± 10.540.85 ± 0.340.90 ± 0.320.90 ± 0.320.87 ± 0.221.00 ± 0.00
DCombined100.00 ± 0.000.90 ± 0.320.90 ± 0.320.90 ± 0.320.90 ± 0.211.00 ± 0.00
E3DFD96.67 ± 10.540.87 ± 0.320.90 ± 0.320.90 ± 0.320.89 ± 0.231.00 ± 0.00
EFixDur100.00 ± 0.000.90 ± 0.320.90 ± 0.320.90 ± 0.320.90 ± 0.211.00 ± 0.00
ECombined100.00 ± 0.000.90 ± 0.320.90 ± 0.320.90 ± 0.320.90 ± 0.211.00 ± 0.00
StimP3DFD71.01 ± 16.950.77 ± 0.280.55 ± 0.410.66 ± 0.350.61 ± 0.220.80 ± 0.20
StimPFixDur63.20 ± 12.800.54 ± 0.250.63 ± 0.300.63 ± 0.320.57 ± 0.160.73 ± 0.16
StimPCombined74.21 ± 18.180.79 ± 0.290.59 ± 0.420.70 ± 0.360.65 ± 0.240.86 ± 0.19
StimN3DFD72.54 ± 14.770.76 ± 0.280.58 ± 0.370.67 ± 0.340.63 ± 0.180.81 ± 0.18
StimNFixDur53.71 ± 11.580.50 ± 0.230.52 ± 0.280.58 ± 0.340.49 ± 0.140.62 ± 0.16
StimNCombined72.29 ± 15.990.76 ± 0.280.58 ± 0.400.67 ± 0.350.63 ± 0.200.80 ± 0.19
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.
Table A9. Classification performance of RBF classifier for EC vs. ER using M P s   = 10 longest fixation-duration locations (10-fold cross-validation), with balanced data for each class (n = 16).
Table A9. Classification performance of RBF classifier for EC vs. ER using M P s   = 10 longest fixation-duration locations (10-fold cross-validation), with balanced data for each class (n = 16).
Stimulus SetFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD92.33 ± 12.670.86 ± 0.310.79 ± 0.360.82 ± 0.340.82 ± 0.220.94 ± 0.12
AFixDur77.25 ± 14.180.66 ± 0.280.79 ± 0.300.75 ± 0.320.70 ± 0.180.87 ± 0.09
ACombined93.08 ± 10.740.86 ± 0.310.82 ± 0.340.84 ± 0.330.84 ± 0.220.99 ± 0.04
B3DFD95.83 ± 10.580.89 ± 0.310.85 ± 0.340.85 ± 0.340.86 ± 0.220.98 ± 0.06
BFixDur88.96 ± 13.750.78 ± 0.310.85 ± 0.320.85 ± 0.340.82 ± 0.230.91 ± 0.16
BCombined96.67 ± 10.540.90 ± 0.320.85 ± 0.340.85 ± 0.340.87 ± 0.220.99 ± 0.03
C3DFD93.89 ± 11.250.86 ± 0.320.85 ± 0.340.85 ± 0.340.85 ± 0.220.96 ± 0.12
CFixDur91.11 ± 12.610.83 ± 0.310.85 ± 0.340.85 ± 0.340.83 ± 0.220.96 ± 0.08
CCombined94.44 ± 12.000.88 ± 0.320.85 ± 0.340.85 ± 0.340.86 ± 0.230.97 ± 0.08
D3DFD96.67 ± 10.540.90 ± 0.320.85 ± 0.340.85 ± 0.340.87 ± 0.220.94 ± 0.18
DFixDur90.83 ± 14.930.77 ± 0.420.85 ± 0.340.75 ± 0.420.80 ± 0.270.94 ± 0.18
DCombined96.67 ± 10.540.90 ± 0.320.85 ± 0.340.85 ± 0.340.87 ± 0.220.94 ± 0.18
E3DFD96.67 ± 10.540.90 ± 0.320.85 ± 0.340.85 ± 0.340.87 ± 0.221.00 ± 0.00
EFixDur96.67 ± 10.540.90 ± 0.320.85 ± 0.340.85 ± 0.340.87 ± 0.221.00 ± 0.00
ECombined96.67 ± 10.540.90 ± 0.320.85 ± 0.340.85 ± 0.340.87 ± 0.221.00 ± 0.00
StimP3DFD79.12 ± 9.380.67 ± 0.240.74 ± 0.320.74 ± 0.320.70 ± 0.150.83 ± 0.08
StimPFixDur61.97 ± 10.650.51 ± 0.240.66 ± 0.280.64 ± 0.320.57 ± 0.140.70 ± 0.13
StimPCombined85.88 ± 9.600.76 ± 0.280.78 ± 0.340.79 ± 0.320.76 ± 0.180.92 ± 0.05
StimN3DFD77.08 ± 16.150.69 ± 0.260.72 ± 0.360.72 ± 0.340.68 ± 0.190.84 ± 0.14
StimNFixDur58.54 ± 8.560.46 ± 0.190.64 ± 0.250.62 ± 0.300.54 ± 0.110.65 ± 0.07
StimNCombined81.67 ± 14.860.72 ± 0.260.76 ± 0.360.76 ± 0.340.72 ± 0.190.90 ± 0.10
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.

Appendix C

Mathematical Form of Linear Mixed-Effects Modelling Equations

Fixation-duration model:
F i x a t i o n D u r a t i o n i j                                                     =   β 0   +   β 1 E C i   +   β 2 E R i   +   β 3 P a t h o l o g y j   +   β 4 ( E C i   ×   P a t h o l o g y j )                                                     +   β 5 ( E R i   ×   P a t h o l o g y j )   +   u i   +   v j   +   ε i j
Here, i     1 , 2 , , N p a r t i c i p a n t s and j     1 , 2 , , N s t i m u l i denote indices for participant (Subject) and stimulus (StimuliID), respectively. N p a r t i c i p a n t s is the total number of participants, and N s t i m u l i is the total number of stimuli. The random intercepts are defined as:
u i ~ N 0 , σ S u b j e c t 2 v j ~ N 0 , σ S t i m u l i I D 2 ε i j ~ N 0 , σ 2
where:
u i represents the participant-specific deviation from the grand intercept.
v j represents the stimulus-specific deviation from the grand intercept.
ε i j represents the residual error for the observation from participant i viewing stimulus j .
The variance components have the following interpretations:
  • σ S u b j e c t 2 is the variance of the participant-level random intercepts. It quantifies how much baseline fixation duration varies between participants after accounting for the fixed effects.
  • σ S t i m u l i I D 2 is the variance of the stimulus-level random intercepts. It quantifies how much baseline fixation duration varies between stimuli after accounting for the fixed effects.
  • σ 2 is the residual variance, representing within-participant × stimulus variability not explained by either the fixed effects or the random intercepts.
    All random components are assumed to be mutually independent.
3DFD model:
M e a n 3 D F D i j = γ 0 + γ 1 E C i + γ 2 E R i + γ 3 P a t h o l o g y j + γ 4 ( E C i × P a t h o l o g y j ) + γ 5 ( E R i × P a t h o l o g y j ) + u i + v j + ε i j
The notation, indexing, and random-effects structure for the 3DFD model are identical to those described for the fixation-duration model (Equations (A1) and (A2)). The coefficients γ k correspond to fixed effects for Mean3DFD rather than fixation duration.

Appendix D

Table A10. Dataset size for N vs. EC classification. Participant count: N (Np = 29), EC (Np = 24), total = 53 participants.
Table A10. Dataset size for N vs. EC classification. Participant count: N (Np = 29), EC (Np = 24), total = 53 participants.
Stimulus GroupNo. of StimuliTotal Classifier SamplesApprox. Test Fold SizeApprox. Training Size
A1053053477
B421221191
C315916143
D153548
E153548
StimP191007101906
StimN201060106954
Full dataset3920672071860
Table A11. Dataset size for N vs. ER classification. Participant count: N (Np = 29), ER (Np = 16), total = 45 participants.
Table A11. Dataset size for N vs. ER classification. Participant count: N (Np = 29), ER (Np = 16), total = 45 participants.
Stimulus GroupNo. of StimuliTotal Classifier SamplesApprox. Test Fold SizeApprox. Training Size
A1045045405
B418018162
C313514121
D145540
E145540
StimP1985586769
StimN2090090810
Full dataset3917551761579
Table A12. Dataset size for EC vs. ER classification. Participant count: EC (Np = 24), ER (Np = 16), total = 40 participants.
Table A12. Dataset size for EC vs. ER classification. Participant count: EC (Np = 24), ER (Np = 16), total = 40 participants.
Stimulus GroupNo. of StimuliTotal Classifier SamplesApprox. Test Fold SizeApprox. Training Size
A1040040360
B416016144
C312012108
D140436
E140436
StimP1976076684
StimN2080080720
Full dataset3915601561404

References

  1. Downing, P.; Liu, J.; Kanwisher, N. Testing Cognitive Models of Visual Attention with FMRI and MEG. Neuropsychologia 2001, 39, 1329–1342. [Google Scholar] [CrossRef]
  2. Brefczynski, J.A.; DeYoe, E.A. A Physiological Correlate of the “spotlight” of Visual Attention. Nat. Neurosci. 1999, 2, 370–374. [Google Scholar] [CrossRef] [PubMed]
  3. Cristino, F.; Mathôt, S.; Theeuwes, J.; Gilchrist, I.D. ScanMatch: A Novel Method for Comparing Fixation Sequences. Behav. Res. Methods 2010, 42, 692–700. [Google Scholar] [CrossRef] [PubMed]
  4. Alamudun, F.; Yoon, H.-J.; Hudson, K.B.; Morin-Ducote, G.; Hammond, T.; Tourassi, G.D. Fractal Analysis of Visual Search Activity for Mass Detection during Mammographic Screening. Med. Phys. 2017, 44, 832–846. [Google Scholar] [CrossRef]
  5. Gandomkar, Z.; Tay, K.; Brennan, P.C.; Kozuch, E.; Mello-Thoms, C. Can Eye-Tracking Metrics Be Used to Better Pair Radiologists in a Mammogram Reading Task? Med. Phys. 2018, 45, 4844–4856. [Google Scholar] [CrossRef] [PubMed]
  6. Kelly, B.S.; Rainford, L.A.; Darcy, S.P.; Kavanagh, E.C.; Toomey, R.J. The Development of Expertise in Radiology: In Chest Radiograph Interpretation, “Expert” Search Pattern May Predate “Expert” Levels of Diagnostic Accuracy for Pneumothorax Identification. Radiology 2016, 280, 252–260. [Google Scholar] [CrossRef]
  7. Kok, E.M.; Jarodzka, H.; de Bruin, A.B.H.; BinAmir, H.A.N.; Robben, S.G.F.; van Merriënboer, J.J.G. Systematic Viewing in Radiology: Seeing More, Missing Less? Adv. Health Sci. Educ. Theory Pract. 2015, 21, 189–205. [Google Scholar] [CrossRef]
  8. Kundel, H.L.; Nodine, C.F.; Conant, E.F.; Weinstein, S.P. Holistic Component of Image Perception in Mammogram Interpretation: Gaze-Tracking Study. Radiology 2007, 242, 396–402. [Google Scholar] [CrossRef]
  9. Voisin, S.; Pinto, F.; Morin-Ducote, G.; Hudson, K.B.; Tourassi, G.D. Predicting Diagnostic Error in Radiology via Eye-Tracking and Image Analytics: Preliminary Investigation in Mammography. Med. Phys. 2013, 40, 101906. [Google Scholar] [CrossRef]
  10. Suman, A.A.; Russo, C.; Carrigan, A.; Nalepka, P.; Liquet-Weiland, B.; Newport, R.A.; Kumari, P.; Di Ieva, A. Spatial and Time Domain Analysis of Eye-Tracking Data during Screening of Brain Magnetic Resonance Images. PLoS ONE 2021, 16, e0260717. [Google Scholar] [CrossRef]
  11. Miller, T.; Durlik, I.; Łobodzińska, A.; Dorobczyński, L.; Jasionowski, R.; Miller, T.; Durlik, I.; Łobodzińska, A.; Dorobczyński, L.; Jasionowski, R. AI in Context: Harnessing Domain Knowledge for Smarter Machine Learning. Appl. Sci. 2024, 14, 11612. [Google Scholar] [CrossRef]
  12. Zhu, F.; Zhang, J.; Dang, R.; Hu, B.; Wang, Q. MTNet: Multimodal Transformer Network for Mild Depression Detection through Fusion of EEG and Eye Tracking. Biomed. Signal Process. Control 2025, 100, 106996. [Google Scholar] [CrossRef]
  13. Singh, L.K.; Khanna, M.; Pooja, A. Novel Multimodality Based Dual Fusion Integrated Approach for Efficient and Early Prediction of Glaucoma. Biomed. Signal Process. Control 2022, 73, 103468. [Google Scholar] [CrossRef]
  14. Öksüz, C.; Urhan, O.; Güllü, M.K. Brain Tumor Classification Using the Fused Features Extracted from Expanded Tumor Region. Biomed. Signal Process. Control 2022, 72, 103356. [Google Scholar] [CrossRef]
  15. Chen, Z.; Liu, Z.; Song, Y. Gaze-Guided Vision Transformer for Chest X-Ray Image Classification. Biomed. Signal Process. Control 2026, 111, 108298. [Google Scholar] [CrossRef]
  16. Crowe, E.M.; Gilchrist, I.D.; Kent, C. New Approaches to the Analysis of Eye Movement Behaviour across Expertise While Viewing Brain MRIs. Cogn. Res. Princ. Implic. 2018, 3, 12. [Google Scholar] [CrossRef]
  17. Matsumoto, D.; Kikuchi, T.; Takagi, Y.; Kojima, S.; Kobayashi, R.; Ueda, D.; Yamamoto, K.; Kawabe, S.; Mori, H. Eye Tracking as a Tool to Quantify the Effects of CAD Display on Radiologists’ Interpretation of Chest Radiographs. Eur. J. Radiol. Open 2026, 16, 100731. [Google Scholar] [CrossRef]
  18. Cavaro-Ménard, C.; Tanguy, J.-Y.; Le Callet, P. Eye-Position Recording during Brain MRI Examination to Identify and Characterize Steps of Glioma Diagnosis. In Medical Imaging 2010: Image Perception, Observer Performance, and Technology Assessment; SPIE: Bellingham, WA, USA, 2010; Volume 7627, pp. 120–127. [Google Scholar]
  19. Li, S.; Duffy, M.C.; Lajoie, S.P.; Zheng, J.; Lachapelle, K. Using Eye Tracking to Examine Expert-Novice Differences during Simulated Surgical Training: A Case Study. Comput. Hum. Behav. 2023, 144, 107720. [Google Scholar] [CrossRef]
  20. Khan, R.S.A.; Tien, G.; Atkins, M.S.; Zheng, B.; Panton, O.N.M.; Meneghetti, A.T. Analysis of Eye Gaze: Do Novice Surgeons Look at the Same Location as Expert Surgeons during a Laparoscopic Operation? Surg. Endosc. 2012, 26, 3536–3540. [Google Scholar] [CrossRef] [PubMed]
  21. Wilson, M.; McGrath, J.; Vine, S.; Brewer, J.; Defriend, D.; Masters, R. Psychomotor Control in a Virtual Laparoscopic Surgery Training Environment: Gaze Control Parameters Differentiate Novices from Experts. Surg. Endosc. 2010, 24, 2458–2464. [Google Scholar] [CrossRef]
  22. Dalveren, G.G.M.; Cagiltay, N.E. Distinguishing Intermediate and Novice Surgeons by Eye Movements. Front. Psychol. 2020, 11, 542752. [Google Scholar] [CrossRef] [PubMed]
  23. Eivazi, S.; Hafez, A.; Fuhl, W.; Afkari, H.; Kasneci, E.; Lehecka, M.; Bednarik, R. Optimal Eye Movement Strategies: A Comparison of Neurosurgeons Gaze Patterns When Using a Surgical Microscope. Acta Neurochir. 2017, 159, 959–966. [Google Scholar] [CrossRef] [PubMed]
  24. Upadhyaya, S. Inspecting the Retina: Oculomotor Patterns and Accuracy in Fundus Image Interpretation by Novice Versus Experienced Eye Care Practitioners. J. Eye Mov. Res. 2026, 19, 11. [Google Scholar] [CrossRef]
  25. Dewhurst, R.; Nyström, M.; Jarodzka, H.; Foulsham, T.; Johansson, R.; Holmqvist, K. It Depends on How You Look at It: Scanpath Comparison in Multiple Dimensions with MultiMatch, a Vector-Based Approach. Behav. Res. Methods 2012, 44, 1079–1100. [Google Scholar] [CrossRef]
  26. Newport, R.A.; Russo, C.; Liu, S.; Suman, A.A.; Di Ieva, A. SoftMatch: Comparing Scanpaths Using Combinatorial Spatio-Temporal Sequences with Fractal Curves. Sensors 2022, 22, 7438. [Google Scholar] [CrossRef] [PubMed]
  27. Wolfe, J.M.; Horowitz, T.S. Five Factors That Guide Attention in Visual Search. Nat. Hum. Behav. 2017, 1, 58. [Google Scholar] [CrossRef]
  28. Wolfe, J.M. Guided Search 4.0: Current Progress with a Model of Visual Search. In Integrated Models of Cognitive Systems; Oxford University Press: New York, NY, USA, 2012. [Google Scholar] [CrossRef]
  29. Markia, B.; Mezei, T.; Báskay, J.; Pollner, P.; Mátyás, A.; Simon, Á.; Várallyay, P.; Banczerowski, P.; Erőss, L. Consistency and Grade Prediction of Intracranial Meningiomas Based on Fractal Geometry Analysis. Neurosurg. Rev. 2025, 48, 598. [Google Scholar] [CrossRef]
  30. Castiglione, J.A.; Drake, A.W.; Hussein, A.E.; Johnson, M.D.; Palmisciano, P.; Smith, M.S.; Robinson, M.W.; Stahl, T.L.; Jandarov, R.A.; Grossman, A.W.; et al. Complex Morphologic Analysis of Cerebral Aneurysms Through the Novel Use of Fractal Dimension as a Predictor of Rupture Status: A Proof of Concept Study. World Neurosurg. 2023, 175, e64–e72. [Google Scholar] [CrossRef]
  31. Di Ieva, A.; Niamah, M.; Menezes, R.J.; Tsao, M.; Krings, T.; Cho, Y.-B.; Schwartz, M.L.; Cusimano, M.D. Computational Fractal-Based Analysis of Brain Arteriovenous Malformation Angioarchitecture. Neurosurgery 2014, 75, 72–79. [Google Scholar] [CrossRef][Green Version]
  32. EyeLink Data Viewer, Version 4.2.1; SR Research Lab: Oakville, ON, Canada, 2021. Available online: https://www.Sr-Research.Com/Data-Viewer/ (accessed on 25 March 2026).
  33. Mandelbrot, B.B.; Wheeler, J.A. The Fractal Geometry of Nature. Am. J. Phys. 1983, 51, 286–287. [Google Scholar] [CrossRef]
  34. Speiser, J.L. A Random Forest Method with Feature Selection for Developing Medical Prediction Models with Clustered and Longitudinal Data. J. Biomed. Inform. 2021, 117, 103763. [Google Scholar] [CrossRef]
  35. Gleißner, C.; Kaczmarz, S.; Kufer, J.; Schmitzer, L.; Kallmayer, M.; Zimmer, C.; Wiestler, B.; Preibisch, C.; Göttler, J. Hemodynamic MRI Parameters to Predict Asymptomatic Unilateral Carotid Artery Stenosis with Random Forest Machine Learning. Front. Neuroimaging 2022, 1, 1056503. [Google Scholar] [CrossRef] [PubMed]
  36. Cristianini, N.; Shawe-Taylor, J. An Introduction to Support Vector Machines and Other Kernel-Based Learning Methods. In An Introduction to Support Vector Machines and Other Kernel-based Learning Methods; Cambridge University Press: Cambridge, UK, 2000. [Google Scholar] [CrossRef]
  37. Di Ieva, A.; Grizzi, F.; Jelinek, H.; Pellionisz, A.J.; Losa, G.A. Fractals in the Neurosciences, Part I: General Principles and Basic Neurosciences. Neuroscientist 2014, 20, 403–417. [Google Scholar] [CrossRef] [PubMed]
  38. Alexander, R.G.; Waite, S.; Macknik, S.L.; Martinez-Conde, S. What Do Radiologists Look for? Advances and Limitations of Perceptual Learning in Radiologic Search. J. Vis. 2020, 20, 17. [Google Scholar] [CrossRef]
  39. Sheridan, H.; Reingold, E.M. The Holistic Processing Account of Visual Expertise in Medical Image Perception: A Review. Front. Psychol. 2017, 8, 252697. [Google Scholar] [CrossRef]
  40. Chen, Y.; Sudin, E.S.; Partridge, G.J.; Taib, A.G.; Darker, I.T.; Phillips, P.; James, J.J.; Satchithananda, K.; Sharma, N.; Michell, M.J. Measuring reader fatigue in the interpretation of screening digital breast tomosynthesis (DBT). Br. J. Radiol. 2023, 96, 20220629. [Google Scholar] [CrossRef]
  41. Bertram, R.; Helle, L.; Kaakinen, J.K.; Svedström, E. The effect of expertise on eye movement behaviour in medical image perception. PLoS ONE 2013, 8, e66169. [Google Scholar] [CrossRef]
  42. Andersson, R.; Nyström, M.; Holmqvist, K.; Andersson, R.; Nyström, M.; Holmqvist, K. Sampling Frequency and Eye-Tracking Measures: How Speed Affects Durations, Latencies, and More. J. Eye Mov. Res. 2010, 3, 1–12. [Google Scholar] [CrossRef]
  43. Moradizeyveh, S.; Hanif, A.; Liu, S.; Qi, Y.; Beheshti, A.; Di Ieva, A. Eye-Guided Multimodal Fusion: Toward an Adaptive Learning Framework Using Explainable Artificial Intelligence. Sensors 2025, 25, 4575. [Google Scholar] [CrossRef] [PubMed]
  44. Martinez, S.; Ramirez-Tamayo, C.; Faruqui, S.H.A.; Clark, K.; Alaeddini, A.; Czarnek, N.; Aggarwal, A.; Emamzadeh, S.; Mock, J.R.; Golob, E.J. Discrimination of Radiologists’ Experience Level Using Eye-Tracking Technology and Machine Learning: Case Study. JMIR Form. Res. 2025, 9, e53928. [Google Scholar] [CrossRef]
Figure 1. Eye-tracking experiment. (a) Eye tracker; (b) data collection setup showing participant’s position with respect to Eye Tracker; (c) stimulus presentation sequence; (d) eye-tracking data (scan path) overlayed on a stimulus.
Figure 1. Eye-tracking experiment. (a) Eye tracker; (b) data collection setup showing participant’s position with respect to Eye Tracker; (c) stimulus presentation sequence; (d) eye-tracking data (scan path) overlayed on a stimulus.
Jemr 19 00062 g001
Figure 2. Qualitative comparison of the heatmaps of fixation duration for naïve, consultant and registrar groups. (a) One of the MRI structural images taken from stimulus group D—Inflammatory changes, and used as stimuli with pathology (Multiple sclerosis) highlighted in the red circle. (bd) Heat maps of fixation duration averaged over all the participants in a group and superimposed on the structural MRI for (b) naïve, (c) consultant, and (d) registrar groups. The fixation duration values below 50 ms are not displayed in (bd). Note that the red circle shown in (a) is not present in the stimulus when it is viewed by the participants during their eye-tracking data collection session.
Figure 2. Qualitative comparison of the heatmaps of fixation duration for naïve, consultant and registrar groups. (a) One of the MRI structural images taken from stimulus group D—Inflammatory changes, and used as stimuli with pathology (Multiple sclerosis) highlighted in the red circle. (bd) Heat maps of fixation duration averaged over all the participants in a group and superimposed on the structural MRI for (b) naïve, (c) consultant, and (d) registrar groups. The fixation duration values below 50 ms are not displayed in (bd). Note that the red circle shown in (a) is not present in the stimulus when it is viewed by the participants during their eye-tracking data collection session.
Jemr 19 00062 g002
Figure 3. Qualitative comparison of the heatmaps of fixation duration for naïve, consultant and registrar groups. (a) One of the MRI structural images taken from stimulus group A—Tumors, and used as stimuli with pathology (Convexity meningioma) highlighted in the red circle. (bd) Heat maps of fixation duration averaged over all the participants in a group and superimposed on the structural MRI for (b) naïve, (c) consultant, and (d) registrar groups. The fixation duration values below 50 ms are not displayed in (bd). Note that the red circle shown in (a) is not present in the stimulus when it is viewed by the participants during their eye-tracking data collection session.
Figure 3. Qualitative comparison of the heatmaps of fixation duration for naïve, consultant and registrar groups. (a) One of the MRI structural images taken from stimulus group A—Tumors, and used as stimuli with pathology (Convexity meningioma) highlighted in the red circle. (bd) Heat maps of fixation duration averaged over all the participants in a group and superimposed on the structural MRI for (b) naïve, (c) consultant, and (d) registrar groups. The fixation duration values below 50 ms are not displayed in (bd). Note that the red circle shown in (a) is not present in the stimulus when it is viewed by the participants during their eye-tracking data collection session.
Jemr 19 00062 g003
Figure 4. Fixation duration distributions across participant groups for different pathological stimulus categories. Only M P s = 10 maximum fixation durations are taken. Subpanels (ae) show tumors, cerebrovascular pathologies, others, inflammatory changes, and malformations, respectively; (f) shows all pathological stimuli combined. Boxplots (shown in red color) compare naïve participants viewing normal images (NN), expert consultants viewing normal images (ECN), expert registrars viewing normal images (ERN), and the corresponding groups viewing pathological images (NP, ECP, ERP). Boxes indicate the interquartile range with median values; whiskers denote the non-outlier range. Individual fixation events are overlaid as swarm plots (shown in blue color), where horizontal dispersion reflects the local density of observations along the fixation-duration axis. The y-axis is constrained to a range of [0–900] ms, which covers more than the 99th percentile of the y-axis values. Note that in each subpanel (ae), normal images (StimN) contribute 20 stimuli, whereas the paired pathology category contributes A = 10, B = 4, C = 3, D = 1, E = 1 stimuli (Table 2). As each stimulus contributes its top 10 fixations per participant, the total number of plotted fixation events differs between normal and pathology within a subpanel and influences the density and visual appearance of the plotted observations.
Figure 4. Fixation duration distributions across participant groups for different pathological stimulus categories. Only M P s = 10 maximum fixation durations are taken. Subpanels (ae) show tumors, cerebrovascular pathologies, others, inflammatory changes, and malformations, respectively; (f) shows all pathological stimuli combined. Boxplots (shown in red color) compare naïve participants viewing normal images (NN), expert consultants viewing normal images (ECN), expert registrars viewing normal images (ERN), and the corresponding groups viewing pathological images (NP, ECP, ERP). Boxes indicate the interquartile range with median values; whiskers denote the non-outlier range. Individual fixation events are overlaid as swarm plots (shown in blue color), where horizontal dispersion reflects the local density of observations along the fixation-duration axis. The y-axis is constrained to a range of [0–900] ms, which covers more than the 99th percentile of the y-axis values. Note that in each subpanel (ae), normal images (StimN) contribute 20 stimuli, whereas the paired pathology category contributes A = 10, B = 4, C = 3, D = 1, E = 1 stimuli (Table 2). As each stimulus contributes its top 10 fixations per participant, the total number of plotted fixation events differs between normal and pathology within a subpanel and influences the density and visual appearance of the plotted observations.
Jemr 19 00062 g004
Figure 5. Distributions of local geometric complexity, quantified by the 3DFD, across participant groups for different pathological stimulus categories. The 3DFD values are taken only for the top 10 maximum fixation duration locations. Subpanels (ae) correspond to tumors, cerebrovascular pathologies, others, inflammatory changes, and malformations, respectively; (f) shows all pathological stimuli combined. Boxplots (shown in red color) compare naïve participants viewing normal images (NN), expert consultants viewing normal images (ECN), expert registrars viewing normal images (ERN), and the corresponding groups viewing pathological images (NP, ECP, ERP). Boxes represent the interquartile range with median values, and whiskers denote the non-outlier range. Individual 3DFD values at fixation locations are overlaid as swarm plots (shown in blue color), where horizontal dispersion reflects the local density of observations along the 3DFD axis. The y-axis is constrained to a range of [1.7–2.3], which covers more than 99 percentiles of the y-axis values. Note that in each subpanel (ae), normal images (StimN) contribute 20 stimuli, whereas the paired pathology category contributes A = 10, B = 4, C = 3, D = 1, E = 1 stimuli (Table 2). As each stimulus contributes its top 10 fixations per participant, the total number of plotted fixation events differs between normal and pathology within a subpanel and influences the density and visual appearance of the plotted observations.
Figure 5. Distributions of local geometric complexity, quantified by the 3DFD, across participant groups for different pathological stimulus categories. The 3DFD values are taken only for the top 10 maximum fixation duration locations. Subpanels (ae) correspond to tumors, cerebrovascular pathologies, others, inflammatory changes, and malformations, respectively; (f) shows all pathological stimuli combined. Boxplots (shown in red color) compare naïve participants viewing normal images (NN), expert consultants viewing normal images (ECN), expert registrars viewing normal images (ERN), and the corresponding groups viewing pathological images (NP, ECP, ERP). Boxes represent the interquartile range with median values, and whiskers denote the non-outlier range. Individual 3DFD values at fixation locations are overlaid as swarm plots (shown in blue color), where horizontal dispersion reflects the local density of observations along the 3DFD axis. The y-axis is constrained to a range of [1.7–2.3], which covers more than 99 percentiles of the y-axis values. Note that in each subpanel (ae), normal images (StimN) contribute 20 stimuli, whereas the paired pathology category contributes A = 10, B = 4, C = 3, D = 1, E = 1 stimuli (Table 2). As each stimulus contributes its top 10 fixations per participant, the total number of plotted fixation events differs between normal and pathology within a subpanel and influences the density and visual appearance of the plotted observations.
Jemr 19 00062 g005
Table 1. Summary of the participant groups, their age, gender and number in each group.
Table 1. Summary of the participant groups, their age, gender and number in each group.
Participant Group LabelGroup NameNumber of ParticipantsAge Range
(Years)
Gender
(Males)
Gender
(Females)
Gender
(Prefer not to Say)
NNaïve2923–521316-
ECExpert consultants2429–611851
ERExpert registrars1626–42106-
Table 2. Summary of the stimulus groups, types of MR brain images within each group and their number.
Table 2. Summary of the stimulus groups, types of MR brain images within each group and their number.
Stimuli Group LabelImage TypeStimuli Group (Based on Type of Images)Number of Stimuli
APathologicalTumors 10
BPathologicalCerebrovascular pathologies4
CPathologicalOthers 3
DPathologicalInflammatory changes1
EPathologicalMalformations 1
StimPPathologicalAll pathological MR images combined19
StimNNormalNormal MR images20
Table 3. Feature sets used in classifiers.
Table 3. Feature sets used in classifiers.
Feature SetNumber of Features
3DFD only10
Fixation duration only10
Combined 3DFD and fixation duration20
Table 4. Fixed-Effects Results for Fixation Duration × ImageType across stimulus groups (A–E, and StimP). We chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant.
Table 4. Fixed-Effects Results for Fixation Duration × ImageType across stimulus groups (A–E, and StimP). We chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant.
Stimulus GroupPredictorβSEtdfpp_FDRCohen’s d
AImageType
(tumors)
−8.8335.675−1.5620,0940.120.144−0.049
AEC × ImageType (tumors)19.0916.2133.0720,0940.0000.004 **0.106
AER × ImageType (tumors)45.6496.5876.9320,0940.0000.000 ***0.244
BImageType
(cerebrovascular pathologies)
−5.7008.004−0.7116,0740.4760.476−0.033
BEC × ImageType (cerebrovascular pathologies)25.2718.4033.0116,0740.0030.006 **0.147
BER × ImageType (cerebrovascular pathologies)42.3049.2564.5716,0740.0000.000 ***0.247
CImageType
(others)
0.9829.5420.1015,4040.9180.9180.006
CEC × ImageType (others)18.8649.4921.9915,4040.0470.0940.110
CER × ImageType (others)26.56710.4552.5415,4040.0110.033 *0.155
DImageType
(inflammatory)
−6.22814.265−0.4414,0640.6620.662−0.035
DEC × ImageType (inflammatory)22.52616.3341.3814,0640.1680.2520.126
DER × ImageType (inflammatory)25.54017.9921.4214,0640.1560.2520.143
EImageType (malfomations)−3.29414.055−0.2314,0640.8150.815−0.019
EEC × ImageType (malfomations)17.69615.5841.1414,0640.2560.3840.104
EER × ImageType (malfomations)42.39317.1662.4714,0640.0140.042 *0.249
StimPImageType
(all pathologies)
−6.1955.021−1.226,1240.220.26−0.034
StimPEC × ImageType (all pathologies)20.4645.1603.8926,1240.0000.000 ***0.112
StimPER × ImageType (all pathologies)39.7335.6836.8626,1240.0000.000 ***0.216
Note. Fixed-effects estimates are from linear mixed-effects models of the form Fixation Duration ~ Group × ImageType + (1|Subject) + (1|StimuliID). The naïve group (N) served as the reference category, and normal images served as the reference level for ImageType. β = Fixation duration estimate; d = Cohen’s d. Significant p and p_FDR values (p/p_FDR < 0.05) are shown in bold. Significance levels on FDR-corrected p-values (p_FDR) are marked as: p_FDR < 0.05 (*), p_FDR < 0.01 (**), p_FDR < 0.001 (***).
Table 5. Fixed-Effects Results for 3DFD × ImageType across stimulus groups (A–E, and StimP). We chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant.
Table 5. Fixed-Effects Results for 3DFD × ImageType across stimulus groups (A–E, and StimP). We chose the M P s = 10, which corresponds to the first 10 maximum fixation durations for each stimulus for each participant.
Stimulus GroupPredictorβSEtdfpp_FDRCohen’s d
AImageType
(tumors)
0.0220.0112.1220,0940.0340.1020.161
AEC × ImageType (tumors)0.0090.0051.8620,0940.0620.1170.066
AER × ImageType (tumors)0.0090.0051.7620,0940.0780.1170.066
BImageType
(cerebrovascular pathologies)
0.0320.0152.1216,0740.0340.0510.252
BEC × ImageType (cerebrovascular pathologies)−0.0310.006−5.0116,0740.0000.000 ***−0.244
BER × ImageType (cerebrovascular pathologies)−0.0260.007−3.8116,0740.0000.000 ***−0.205
CImageType
(others)
−0.0070.017−0.4115,4040.6830.683−0.056
CEC × ImageType (others)0.0070.0071.0215,4040.3080.4620.056
CER × ImageType (others)0.0060.0080.8215,4040.4120.4940.048
DImageType
(inflammatory)
−0.0340.028−1.2314,0640.2190.263−0.254
DEC × ImageType (inflammatory)0.0780.0126.3714,0640.0000.000 ***0.582
DER × ImageType (inflammatory)0.0820.0146.0714,0640.0000.000 ***0.612
EImageType
(malfomations)
0.1030.0283.6914,0640.0000.000 ***0.857
EEC × ImageType (malfomations)−0.0110.011−1.0214,0640.3090.371−0.092
EER × ImageType (malfomations)−0.0100.012−0.7914,0640.4290.429−0.083
StimPImageType
(all pathologies)
0.0190.0092.0226,1240.0270.0810.147
StimPEC × ImageType (all pathologies)0.0040.0040.6626,1240.5070.5070.021
StimPER × ImageType (all pathologies)0.0050.0040.9226,1240.3590.4310.028
Note. Fixed-effects estimates are from linear mixed-effects models of the form Mean3DFD ~ Group × ImageType + (1|Subject) + (1|StimuliID). The naïve group (N) served as the reference category, and normal images served as the reference level for ImageType. β = 3DFD estimate; d = Cohen’s d. Significant p and p_FDR values (p/p_FDR < 0.05) are shown in bold. Significance levels on FDR-corrected p-values (p_FDR) are marked as: p_FDR < 0.001 (***).
Table 6. Classification performance for N vs. EC Random Forest classifier using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Table 6. Classification performance for N vs. EC Random Forest classifier using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Stimulus GroupFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD91.27 ± 8.380.90 ± 0.080.93 ± 0.120.91 ± 0.150.90 ± 0.080.97 ± 0.06
AFixation Duration64.43 ± 11.670.41 ± 0.160.82 ± 0.110.60 ± 0.230.59 ± 0.100.62 ± 0.12
ACombined94.03 ± 8.560.94 ± 0.080.94 ± 0.110.93 ± 0.130.94 ± 0.090.97 ± 0.06
B3DFD96.50 ± 6.260.98 ± 0.040.97 ± 0.110.97 ± 0.110.96 ± 0.061.00 ± 0.01
BFixation Duration83.75 ± 5.430.74 ± 0.150.90 ± 0.110.88 ± 0.170.81 ± 0.050.91 ± 0.06
BCombined98.50 ± 4.741.00 ± 0.000.97 ± 0.080.97 ± 0.090.98 ± 0.051.00 ± 0.00
C3DFD95.33 ± 8.340.99 ± 0.040.93 ± 0.140.93 ± 0.140.95 ± 0.080.99 ± 0.02
CFixation Duration87.33 ± 9.140.83 ± 0.160.94 ± 0.090.89 ± 0.150.86 ± 0.090.91 ± 0.11
CCombined97.33 ± 5.621.00 ± 0.000.96 ± 0.090.95 ± 0.110.97 ± 0.061.00 ± 0.01
D3DFD98.00 ± 6.320.97 ± 0.081.00 ± 0.001.00 ± 0.000.98 ± 0.081.00 ± 0.00
DFixation Duration94.00 ± 13.500.95 ± 0.160.97 ± 0.110.97 ± 0.110.94 ± 0.141.00 ± 0.00
DCombined100.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.00
E3DFD100.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.00
EFixation Duration94.00 ± 13.500.93 ± 0.170.97 ± 0.110.95 ± 0.160.93 ± 0.140.98 ± 0.05
ECombined100.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.001.00 ± 0.00
StimP3DFD75.14 ± 7.280.57 ± 0.090.89 ± 0.100.79 ± 0.210.71 ± 0.060.77 ± 0.08
StimPFixation Duration59.84 ± 9.450.33 ± 0.060.81 ± 0.080.55 ± 0.230.54 ± 0.060.61 ± 0.06
StimPCombined80.86 ± 7.310.66 ± 0.050.92 ± 0.100.85 ± 0.170.78 ± 0.060.83 ± 0.06
StimN3DFD73.38 ± 9.990.60 ± 0.110.83 ± 0.130.74 ± 0.220.70 ± 0.100.80 ± 0.09
StimNFixation Duration57.75 ± 9.940.39 ± 0.100.73 ± 0.050.52 ± 0.160.54 ± 0.080.60 ± 0.08
StimNCombined77.67 ± 9.370.64 ± 0.090.88 ± 0.130.80 ± 0.210.75 ± 0.090.82 ± 0.07
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.
Table 7. Classification performance for N vs. ER Random Forest classifier using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Table 7. Classification performance for N vs. ER Random Forest classifier using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Stimulus SetFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD92.15 ± 9.480.80 ± 0.290.95 ± 0.110.84 ± 0.350.88 ± 0.180.94 ± 0.08
AFixDur79.85 ± 12.080.52 ± 0.300.92 ± 0.150.74 ± 0.370.70 ± 0.160.84 ± 0.09
ACombined93.70 ± 7.000.84 ± 0.300.96 ± 0.090.86 ± 0.330.89 ± 0.170.98 ± 0.03
B3DFD95.12 ± 8.220.86 ± 0.310.96 ± 0.090.87 ± 0.320.91 ± 0.181.00 ± 0.00
BFixDur83.87 ± 12.960.66 ± 0.370.90 ± 0.160.78 ± 0.340.75 ± 0.180.91 ± 0.13
BCombined95.62 ± 9.340.90 ± 0.320.95 ± 0.110.86 ± 0.330.92 ± 0.191.00 ± 0.00
C3DFD92.83 ± 9.690.87 ± 0.310.94 ± 0.100.85 ± 0.320.89 ± 0.180.99 ± 0.02
CFixDur86.50 ± 9.980.67 ± 0.380.94 ± 0.140.72 ± 0.420.78 ± 0.190.90 ± 0.17
CCombined93.67 ± 8.560.84 ± 0.320.95 ± 0.100.86 ± 0.330.90 ± 0.181.00 ± 0.00
D3DFD97.50 ± 7.910.90 ± 0.320.97 ± 0.080.90 ± 0.320.94 ± 0.181.00 ± 0.00
DFixDur95.50 ± 9.560.85 ± 0.340.97 ± 0.080.90 ± 0.320.92 ± 0.191.00 ± 0.00
DCombined95.50 ± 9.560.80 ± 0.420.97 ± 0.080.80 ± 0.420.89 ± 0.241.00 ± 0.00
E3DFD98.00 ± 6.320.85 ± 0.341.00 ± 0.000.90 ± 0.320.93 ± 0.171.00 ± 0.00
EFixDur97.50 ± 7.910.90 ± 0.320.97 ± 0.080.90 ± 0.320.94 ± 0.181.00 ± 0.00
ECombined100.00 ± 0.000.90 ± 0.321.00 ± 0.000.90 ± 0.320.95 ± 0.161.00 ± 0.00
StimP3DFD80.08 ± 9.900.54 ± 0.220.93 ± 0.080.77 ± 0.320.72 ± 0.110.81 ± 0.09
StimPFixDur74.32 ± 13.950.38 ± 0.250.94 ± 0.060.68 ± 0.310.63 ± 0.120.73 ± 0.09
StimPCombined87.05 ± 6.960.67 ± 0.260.96 ± 0.080.85 ± 0.320.81 ± 0.140.89 ± 0.07
StimN3DFD78.33 ± 9.970.53 ± 0.200.91 ± 0.100.74 ± 0.320.71 ± 0.120.79 ± 0.10
StimNFixDur64.58 ± 16.280.20 ± 0.160.90 ± 0.080.51 ± 0.360.50 ± 0.110.60 ± 0.10
StimNCombined81.97 ± 7.230.57 ± 0.210.93 ± 0.080.78 ± 0.320.75 ± 0.110.84 ± 0.07
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.
Table 8. Classification performance of Random Forest classifier for EC vs. ER using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Table 8. Classification performance of Random Forest classifier for EC vs. ER using M P s = 10 longest fixation duration locations (10-fold cross-validation).
Stimulus SetFeature SetAccuracySensitivitySpecificityPrecisionF1AUC
A3DFD94.00 ± 8.680.68 ± 0.470.85 ± 0.320.66 ± 0.470.76 ± 0.250.98 ± 0.02
AFixDur78.92 ± 10.930.47 ± 0.360.81 ± 0.300.63 ± 0.450.63 ± 0.190.85 ± 0.07
ACombined94.92 ± 8.140.68 ± 0.470.86 ± 0.310.67 ± 0.470.77 ± 0.261.00 ± 0.01
B3DFD96.25 ± 6.720.68 ± 0.470.88 ± 0.320.70 ± 0.480.78 ± 0.261.00 ± 0.00
BFixDur90.62 ± 9.430.56 ± 0.420.88 ± 0.320.70 ± 0.480.72 ± 0.230.95 ± 0.06
BCombined96.88 ± 7.930.69 ± 0.480.88 ± 0.320.70 ± 0.480.79 ± 0.261.00 ± 0.00
C3DFD95.00 ± 8.050.66 ± 0.460.88 ± 0.320.70 ± 0.480.77 ± 0.261.00 ± 0.00
CFixDur93.06 ± 10.330.62 ± 0.450.88 ± 0.320.70 ± 0.480.74 ± 0.240.98 ± 0.05
CCombined96.67 ± 8.050.69 ± 0.480.88 ± 0.320.70 ± 0.480.78 ± 0.261.00 ± 0.00
D3DFD97.50 ± 7.910.70 ± 0.480.88 ± 0.320.70 ± 0.480.79 ± 0.271.00 ± 0.00
DFixDur92.50 ± 12.080.63 ± 0.460.88 ± 0.320.70 ± 0.480.74 ± 0.251.00 ± 0.00
DCombined95.00 ± 10.540.67 ± 0.470.88 ± 0.320.70 ± 0.480.77 ± 0.261.00 ± 0.00
E3DFD100.00 ± 0.000.70 ± 0.480.90 ± 0.320.70 ± 0.480.80 ± 0.261.00 ± 0.00
EFixDur97.50 ± 7.910.67 ± 0.470.90 ± 0.320.70 ± 0.480.77 ± 0.251.00 ± 0.00
ECombined100.00 ± 0.000.70 ± 0.480.90 ± 0.320.70 ± 0.480.80 ± 0.261.00 ± 0.00
StimP3DFD79.21 ± 9.410.51 ± 0.360.80 ± 0.300.62 ± 0.450.65 ± 0.190.90 ± 0.04
StimPFixDur70.31 ± 12.980.33 ± 0.250.80 ± 0.300.58 ± 0.430.55 ± 0.130.65 ± 0.10
StimPCombined87.41 ± 7.270.59 ± 0.410.83 ± 0.300.64 ± 0.460.71 ± 0.220.96 ± 0.03
StimN3DFD76.79 ± 9.550.50 ± 0.360.77 ± 0.300.62 ± 0.450.64 ± 0.210.92 ± 0.05
StimNFixDur64.62 ± 16.290.29 ± 0.210.76 ± 0.280.54 ± 0.420.51 ± 0.140.65 ± 0.07
StimNCombined82.42 ± 8.130.55 ± 0.390.80 ± 0.300.64 ± 0.460.67 ± 0.210.95 ± 0.06
Note. Within each stimulus group, boldface indicates the feature set with the highest values for both Accuracy and Specificity across the three feature sets. In cases of tied highest values, boldface was assigned to the feature set that also showed the highest value for the other metric.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Kumari, P.; Azemi, G.; Russo, C.; Di Ieva, A. Characterizing Visual Neurosurgical Expertise in Brain MRI Visualization Using Eye-Tracking and 3D Fractal Dimension Analysis. J. Eye Mov. Res. 2026, 19, 62. https://doi.org/10.3390/jemr19030062

AMA Style

Kumari P, Azemi G, Russo C, Di Ieva A. Characterizing Visual Neurosurgical Expertise in Brain MRI Visualization Using Eye-Tracking and 3D Fractal Dimension Analysis. Journal of Eye Movement Research. 2026; 19(3):62. https://doi.org/10.3390/jemr19030062

Chicago/Turabian Style

Kumari, Poonam, Ghasem Azemi, Carlo Russo, and Antonio Di Ieva. 2026. "Characterizing Visual Neurosurgical Expertise in Brain MRI Visualization Using Eye-Tracking and 3D Fractal Dimension Analysis" Journal of Eye Movement Research 19, no. 3: 62. https://doi.org/10.3390/jemr19030062

APA Style

Kumari, P., Azemi, G., Russo, C., & Di Ieva, A. (2026). Characterizing Visual Neurosurgical Expertise in Brain MRI Visualization Using Eye-Tracking and 3D Fractal Dimension Analysis. Journal of Eye Movement Research, 19(3), 62. https://doi.org/10.3390/jemr19030062

Article Metrics

Back to TopTop