1. Introduction
Visual attention is guided by an interplay of bottom-up and top-down processes that determine how information is selected and processed within complex visual environments. Classic models of attentional selection have emphasized the role of low-level salience—such as contrast, color, and orientation—in capturing attention [
1], as well as the influence of task demands and prior knowledge in guiding gaze behavior [
2]. However, beyond these well-established factors, there is increasing interest in understanding how higher-order visual structure modulates attentional allocation in the absence of semantic content.
From a perceptual perspective, Gestalt principles such as symmetry, proximity, and continuity provide a framework for understanding how the visual system organizes information into coherent structures [
3]. Symmetry in particular occupies a privileged position in this literature: symmetrical configurations are judged as more aesthetically pleasing than their asymmetrical counterparts across a wide range of stimuli [
4], and classic eye-tracking work has shown that symmetry systematically reorganizes visual scanning, with fixations for symmetrical shapes clustering non-uniformly across the figure while both fixation count and fixation time scale directly with structural complexity [
5]. This symmetry preference generalizes across a heterogeneous range of stimulus categories, including abstract shapes, faces, and natural scenes, although its strength varies considerably by category [
6], and more recent eye-tracking work indicates that observers spontaneously scan along a pattern’s axis of reflection in a manner that is similar whether they are explicitly classifying or evaluating the stimulus, suggesting that this oculomotor signature of symmetry detection is not simply a byproduct of an evaluative task [
7]. At the same time, whether symmetry detection is itself an attentionally demanding process, or occurs preattentively, remains debated [
8], suggesting that the relationship between structural regularity and the deployment of visual attention is more nuanced than a simple facilitation account would predict. Recent studies using eye tracking have demonstrated that compositional structure more broadly influences gaze distribution and fixation patterns, highlighting the role of visual organization in guiding attention [
9,
10,
11].
Visual complexity has also been shown to systematically modulate attentional behavior. Increasing levels of complexity are associated with higher cognitive load and changes in fixation patterns, often resulting in more focal and localized attentional allocation [
12]. Modeling what makes an image look complex in the first place remains an active and unresolved problem in its own right: recent computational work shows that structural regularity, color variety, and semantic surprise each make separable, non-redundant contributions to perceived complexity, and that models relying on a single simple metric systematically fail to capture this multidimensionality [
13]—a caution directly relevant to any study, including the present one, that operationalizes complexity through a single generative dimension. At the same time, structured regularity may support perceptual fluency, enabling more distributed exploration of visual scenes. The relationship between complexity and subjective evaluation has a long history in experimental aesthetics: Berlyne’s collative-variables framework proposed that stimulus properties such as complexity, novelty, and ambiguity determine an observer’s arousal potential, with hedonic value following an inverted-U function of arousal—stimuli of intermediate complexity being preferred over both very simple and very complex configurations [
14]. This inverted-U prediction has, however, received markedly mixed empirical support in more recent work: a large-sample re-examination of complexity and liking in product design found scant evidence for the inverted-U pattern once cognitive effort, arousal, and perceived producer skill were accounted for, instead reporting a predominantly negative relationship between complexity and liking [
15]—a direction more consistent with the pattern we report below than with Berlyne’s original formulation. More recent work within this tradition has emphasized processing fluency—the subjective ease with which a stimulus is perceptually or conceptually processed—as a proximal determinant of aesthetic liking, independent of the observer’s explicit awareness of the stimulus properties driving that ease [
16]. A recent integrative review of the neuroaesthetics literature similarly frames aesthetic experience as emerging from the interaction of sensory-motor, emotion-valuation, and meaning-knowledge systems, each engaging distinct and only partially overlapping neural and behavioral signatures [
17]—consistent with the possibility that evaluative judgment and overt attentional/oculomotor engagement, though often studied together, need not be driven by a single unified mechanism. These accounts converge on the prediction that structural manipulations of complexity should shape subjective aesthetic response, but they are comparatively agnostic about whether such manipulations also shape the overt allocation of gaze.
This last point is particularly relevant given a growing body of evidence that overt attention (where the eyes are directed) and covert evaluative or preferential processes can dissociate. In face perception research, for instance, covert and overt measures of attention to attractive faces have been shown to diverge depending on task demands [
18], and developmental work on symmetry preference has found that although young children fixate symmetrical patterns longer than asymmetrical ones, they show no corresponding explicit preference for those patterns—a clear dissociation between spontaneous attentional capture and evaluative judgment [
19]. These findings caution against the assumption that a structural property which reliably shapes subjective judgment will necessarily leave an equally reliable signature in fixation-based attentional metrics, or vice versa.
A complementary way of characterizing attentional deployment during free viewing is the distinction between ambient and focal modes of visual processing. Ambient viewing is characterized by short fixations followed by large-amplitude saccades and is thought to support the rapid extraction of a scene’s global spatial layout, whereas focal viewing involves longer fixations following short saccades and supports detailed, localized processing [
20]. This distinction has been formalized in the coefficient
K, a parametric index derived from the standardized relationship between fixation duration and the amplitude of the subsequent saccade, with positive values indicating focal and negative values indicating ambient processing [
21]. Because harmony and complexity manipulations are natural candidates for shifting the balance between ambient and focal processing, the present study incorporates this framework alongside conventional eye-tracking metrics.
Despite these theoretical advances, relatively few studies have systematically isolated the effects of visual harmony and complexity using tightly controlled stimuli that eliminate semantic and contextual confounds. Many existing studies rely on real-world or artistic images, where multiple visual variables co-vary, making it difficult to attribute attentional effects to specific structural properties.
The present study addresses this gap by examining how controlled manipulations of visual structure modulate attentional allocation during free viewing. Using a set of stimuli derived from a common geometric base, we compare three conditions varying in harmony and complexity while maintaining constant low-level visual features. Eye-tracking measures—including time to first fixation (TTFF), fixation duration, number of fixations, scanpath entropy, and spatial dispersion—are used to characterize differences in visual exploration. To complement these overt gaze-based indices with an independent, non-oculomotor window onto attentional engagement, we additionally recorded webcam-based facial coding throughout stimulus presentation, yielding moment-by-moment estimates of neutral, happy, and surprise expressions together with the coefficient K described above.
We hypothesize that harmonious stimuli will elicit more distributed attentional patterns, whereas disorganized stimuli will induce focal attentional capture. A third condition of structured complexity is expected to produce intermediate patterns, reflecting a balance between global exploration and local processing. Given the dissociation reported between covert/overt attention and explicit preference in related domains [
18,
19], we also consider it plausible that subjective ratings of pleasantness and complexity may respond to the structural manipulation independently of, or even in the absence of, corresponding shifts in fixation-based attentional metrics. By isolating the role of visual structure and combining oculomotor, facial-coding, and subjective measures, this study aims to contribute to a more precise understanding of how formal properties guide attentional dynamics.
This study is intended to make two distinct kinds of contribution, and we flag both explicitly here so that the reader can evaluate each on its own terms. The first is substantive: as reported below, our attempt to manipulate complexity through structural disorder alone produced a pattern of subjective complexity judgments opposite to what was intended, suggesting that structural regularity/disorder is not interchangeable with perceived complexity as a construct—a dissociation with direct implications for how complexity is operationalized in eye-tracking and experimental-aesthetics research using generated or parametric stimuli. The second is methodological: analyzing this dataset surfaced two analytic issues—treating repeated within-participant observations as independent in correlational analyses, and the absence of an explicit multiplicity strategy across a large family of outcomes—that are common in webcam-based eye-tracking and facial-coding studies of this kind but are not always addressed. We report, in detail, how correcting for both changes the statistical picture (
Section 3 and
Section 4), with the aim of providing a worked example that other researchers analyzing similarly structured repeated-measures eye-tracking and facial-coding data can draw on directly.
2. Materials and Methods
2.1. Participants
A total of 50 participants were recruited through the RealEye online platform. All participants reported normal or corrected-to-normal vision and provided informed consent prior to participation. The study was conducted in accordance with the Declaration of Helsinki and approved by the corresponding institutional ethics committee.
Given the use of webcam-based eye tracking, participants were required to complete a calibration procedure prior to the experiment. Data quality was evaluated using RealEye’s built-in quality grade (a 1–6 composite index of calibration accuracy and data completeness). Only participants with a quality grade were retained for analysis, resulting in a final sample of (age range: 19–32 years, , ; 25 female, 14 male, 1 other), with the remaining 10 participants excluded.
2.2. Stimuli
The stimulus set consisted of nine images (three exemplars per condition) derived from a common geometric base and generated through a parametric tessellation procedure, systematically manipulated to vary in visual structure while controlling for low-level features. All stimuli were rendered at the same resolution and aspect ratio and were matched in overall luminance, color palette (a shared orange–teal–cream scheme), and approximate number of visual elements, so that the three conditions differed primarily in their spatial and structural organization rather than in low-level photometric properties.
The three experimental conditions were:
High harmony (low complexity, HH): A symmetrical and highly organized composition built around a central radial motif, characterized by regular spatial distribution, mirror symmetry along both the vertical and horizontal axes, and smooth visual transitions between color regions (see
Figure 1).
High complexity (low harmony, HC): A disorganized tessellation of small, irregularly sized triangular facets covering the same canvas, with disrupted symmetry, non-uniform spatial arrangement, and increased local contrast between adjacent facets, yielding a texture-like appearance without a clear global focal point (see
Figure 2).
Structured complexity (SC): A condition combining a high density of small geometric facets, similar in scale to the HC condition, with an underlying global symmetry and a salient central motif analogous to that of the HH condition—that is, local complexity embedded within global order (see
Figure 3).
The visual similarity between HH and SC relative to HC is a deliberate feature of this design rather than an oversight. SC was constructed specifically to test whether local complexity could be embedded within global order without disrupting harmony-related structure; to isolate that manipulation, SC retains the same global mirror symmetry and central radial motif as HH, differing from it primarily in the density and irregularity of the small facets that compose that structure. HC, by contrast, removes the global symmetry and central motif entirely. Harmony and complexity were therefore defined a priori in purely generative, geometric terms—degree of mirror symmetry, presence/absence of a central motif, and the size and irregularity of the constituent facets—rather than through a preliminary psychophysical scaling of participants’ own perception of these dimensions. This is an important distinction: as reported in
Section 3.2 and discussed in
Section 4, this a priori structural classification did not fully align with participants’ subjective experience of complexity, indicating that the correspondence between generative parameters and perceived harmony/complexity cannot simply be assumed. We return to this point, and to how it should be addressed in future stimulus development (e.g., via an independent pilot rating study prior to the main experiment), in
Section 4.
All stimuli were presented in full-screen format and standardized to identical viewing conditions. No areas of interest (AOIs) were predefined on the stimuli; all eye-tracking metrics were therefore computed over the entire stimulus area rather than for specific sub-regions (see
Section 2.5).
2.3. Apparatus
Eye movements were recorded using the RealEye platform, a webcam-based eye-tracking system with an approximate sampling rate of 30 Hz. Although less precise than laboratory-grade systems, RealEye has been shown to provide reliable measures of fixation patterns and attentional distribution in remote experimental settings. Fixations were identified with RealEye’s default filter settings (gaze-point interpolation at 30 Hz with a maximum gap of 50 ms, a velocity threshold of 150°/s, and minimum and maximum fixation durations of 80 and 400 ms, respectively). We retained these platform defaults, rather than a custom filter, to keep fixation identification consistent with the calibration and validation benchmarks reported for RealEye and to avoid introducing additional researcher degrees of freedom into a webcam-based signal that is already noisier than laboratory-grade recordings. The upper bound of 400 ms is within the range commonly used for dispersion-based fixation classifiers, but it will truncate or split any genuine fixations that exceed it; because sustained fixations beyond 400 ms are, if anything, more likely during effortful or difficult processing, this truncation could attenuate rather than inflate condition differences in mean fixation duration, and we return to this as a limitation on our fixation-duration measure in
Section 4.
The hardware and software constraints of this webcam-based approach do affect spatial and temporal precision, and this is worth making explicit. RealEye’s own validation reports a full-screen spatial accuracy on the order of 100–125 px (corresponding to roughly 2–5 degrees of visual angle depending on viewing distance and screen size, both uncontrolled in this remote setting), compared with sub-degree accuracy typical of laboratory-grade infrared trackers, and a sampling rate (here, ∼30 Hz) an order of magnitude lower than the 250–1000 Hz common in laboratory systems. Consistent with this, a direct comparison of webcam-based (including RealEye) and remote infrared eye tracking reported higher measurement error and reduced precision for the webcam modality, although effects large enough to be behaviorally meaningful (e.g., differences between highly discriminable stimulus categories) were still reliably detected under both modalities [
22]. Similarly, a direct comparison against a laboratory EyeLink 1000 system found that a modern webcam-based tracker approaches, without fully matching, laboratory-grade spatial accuracy [
23], and independent behavioral replications of classic in-lab eye-tracking paradigms using webcam-based, browser-delivered tracking have reported little degradation in the size of well-established effects despite reduced spatial and temporal resolution [
24]. This lower spatial/temporal precision and correspondingly reduced signal-to-noise ratio is a plausible contributor to the null condition effects on eye-tracking metrics reported in
Section 3.1, and we return to it explicitly in
Section 4.
In parallel with eye tracking, RealEye’s facial-coding module recorded participants’ facial expressions from the webcam feed throughout each trial, sampled at the same 30 Hz rate. Facial coding was based on the Facial Action Coding System (FACS) [
25] and yielded, for every 30-ms sample, the estimated intensity (0–1) of three expression categories, neutral, happy, and surprise, together with an instantaneous, standardized estimate of the coefficient
K [
21], an index of ambient (negative values) versus focal (positive values) visual-attentional engagement derived jointly from oculomotor dynamics.
2.4. Procedure
Participants completed the experiment remotely using a desktop or laptop computer equipped with a webcam. After providing consent, they performed a standard nine-point calibration procedure to ensure eye-tracking accuracy; webcam access was also required and calibrated for facial coding.
Each trial consisted of the presentation of a single stimulus for a duration of 6–7 s under free-viewing conditions. This duration was chosen to be long enough to allow multiple fixations and at least one full exploratory pass over each stimulus (yielding, in practice, a mean of approximately 24 fixations per trial; see
Section 3.1) while remaining short enough, across nine trials plus ratings, to limit fatigue in a webcam-based remote setting where sustained attention and calibration quality tend to degrade over longer sessions. Whether 6–7 s is sufficient for global structural properties (as opposed to local, early-fixation salience) to shape the overall distribution of fixations is, however, an open question that this design cannot fully resolve; we consider this explicitly as a candidate explanation for the null eye-tracking results in
Section 4. The nine stimuli (three exemplars × three conditions) were presented once each, in an order randomized independently for every participant to control for order and fatigue effects.
Following each stimulus, participants completed a brief set of subjective ratings:
Perceived pleasantness (1–5 Likert scale)
Perceived complexity (1–5 Likert scale)
(Optional) First perceived element (categorical response)
No specific instructions were given regarding where to look, in order to capture spontaneous attentional behavior. Participants were informed, as part of the consent process and prior to the experiment, that the platform would analyze their facial expressions from the webcam feed in addition to tracking their gaze; consistent with RealEye’s privacy-preserving design, no video was recorded or stored at any point—the platform processes the live webcam feed on the client side and returns only derived gaze coordinates and facial-expression estimates. The ethics approval covering this study explicitly included this webcam-based facial-expression analysis alongside eye tracking.
2.5. Measures
The primary dependent variables were derived from eye-tracking data. Because no areas of interest were predefined on the stimuli (
Section 2.2), all metrics below were computed over fixations landing anywhere within the presentation area:
Time to First Fixation (TTFF): latency, from stimulus onset, to the first recorded fixation.
Fixation Duration: mean duration of fixations per stimulus.
Number of Fixations: total number of fixations during stimulus presentation.
Scanpath Entropy: Shannon entropy of the spatial distribution of fixations across a grid superimposed on the stimulus, normalized by its theoretical maximum (); higher values indicate more spatially distributed viewing.
Heatmap Dispersion: spatial dispersion of fixation coordinates, computed as , where x and y are fixation coordinates expressed as a percentage of stimulus width and height.
From the facial-coding data, the following measures were derived for each participant and stimulus: the mean intensity of neutral, happy, and surprise expressions across the trial, and the trial-level mean of the coefficient K (RealEye’s aggregated per-trial estimate).
Additionally, subjective ratings of perceived pleasantness and complexity were recorded for each stimulus. For all measures, condition-level scores were computed by averaging across the three exemplars of each condition, separately for every participant.
2.6. Data Analysis
Statistical analyses were conducted using a repeated-measures design. Differences across the three stimulus conditions were evaluated using one-way repeated-measures ANOVA for each dependent variable (eye-tracking metrics, facial-coding metrics, and subjective ratings), with condition-level scores obtained by averaging across the three exemplars per condition for each participant.
Post-hoc comparisons were performed using paired-samples t-tests with Bonferroni correction for the three pairwise comparisons per variable. Effect sizes were reported using partial eta squared () for omnibus tests and Cohen’s d for pairwise comparisons.
Because averaging over exemplars conflates condition with the specific stimulus items used to instantiate it, and because only three exemplars were sampled per condition, we additionally fit linear mixed-effects models with crossed random intercepts for participant and item (i.e., not averaging across exemplars) for each primary dependent variable, with condition modeled as a fixed effect (treatment-coded contrasts against the HH baseline). Omnibus condition effects were evaluated via likelihood-ratio tests comparing models with and without the condition fixed effect, both fit by maximum likelihood; parameter estimates are reported from the corresponding restricted maximum likelihood (REML) fits. This analysis serves as a robustness check on the repeated-measures ANOVA results by explicitly partitioning item-level variance rather than treating it as part of the residual.
To contextualize the null eye-tracking results, we conducted a sensitivity power analysis based on the present design (
,
,
,
), determining the minimum partial eta squared for which the study had 80% power, using the noncentral
F distribution. We emphasize that this analysis was necessarily conducted post hoc, after the null pattern was observed, rather than as a prospective sample-size justification prior to data collection; no such prospective power calculation was performed when the study was designed. Consequently, the sensitivity analysis should be read only as evidence of what the present design could and could not detect, not as a substitute for a pre-registered power analysis, and the absence of significant eye-tracking or facial-coding effects is interpreted throughout as an absence of detectable differences under the present design and sample size, rather than as evidence that visual structure has no effect on attention or facial engagement [
26].
Correlational analyses examined relationships between subjective ratings (pleasantness and complexity), eye-tracking metrics, and facial-coding metrics. Because each of the 40 retained participants contributed one observation per condition (i.e., three non-independent observations per participant,
condition-level observations in total), treating these as 120 independent data points—as an ordinary Pearson correlation implicitly does—violates the independence assumption underlying that test and can bias both the point estimate and its associated
p-value. We therefore used repeated-measures correlation (rmcorr) as our primary correlational technique [
27], which estimates the common within-participant association between two variables while explicitly partitioning out between-participant variance, analogous to an ANCOVA with participant modeled as a categorical factor. Because we lacked access to the original
rmcorr software (
https://lmarusich.github.io/rmcorr/ (accessed on 8 September 2026)) in the offline analysis environment used here, we implemented the equivalent ANCOVA-based procedure directly (participant fixed intercepts plus a common slope, fit by ordinary least squares; the partial correlation
and its associated
F-test were derived from the reduction in residual sum of squares attributable to the slope term, with
for
participants). For comparison, we also report the naive Pearson correlation obtained by treating the 120 observations as independent, to make transparent how much the correction changes each estimate.
Given the number of statistical tests performed across outcome families, we distinguish primary from exploratory analyses and adjust for multiplicity accordingly. This designation of primary versus exploratory outcomes reflects which dependent variables the Introduction’s hypotheses target, rather than a formally pre-registered analysis plan. The seven primary tests correspond to the central hypotheses stated in the Introduction and are the repeated-measures ANOVAs on the five eye-tracking metrics and the two subjective ratings; for this family we report both the nominal p-value and significance under a Bonferroni-corrected threshold (). The four facial-coding ANOVAs are treated as a secondary, exploratory family, since facial coding was introduced to complement the primary oculomotor and subjective measures rather than to test an independently specified hypothesis. All correlational analyses (23 rmcorr tests across eye-tracking, facial-coding, and subjective variables) are likewise treated as exploratory and are corrected using the Benjamini–Hochberg false discovery rate (FDR) procedure across that full set of 23 tests; we report FDR-adjusted q-values alongside nominal p-values and flag which associations survive correction. Exploratory findings, whether or not nominally significant, are interpreted as hypothesis-generating rather than confirmatory.
Significance level was set at for all analyses. All analyses were performed in Python v3.14.7 (NumPy, SciPy, and pandas); linear mixed-effects models and the repeated-measures correlation procedure were implemented directly via profile REML/ML and ordinary-least-squares optimization, respectively, since packages providing this functionality (e.g., statsmodels, lme4, rmcorr) were not available in the analysis environment.
4. Discussion
This study examined whether controlled variations in visual harmony and complexity, isolated from semantic content, would relate to overt attentional allocation during free viewing, and complemented conventional eye-tracking metrics with webcam-based facial coding and the ambient/focal coefficient
K. Of the seven primary, family-wise-corrected omnibus tests (
Section 2), only one—perceived complexity—survived correction, in a direction opposite to the intended manipulation; the pleasantness effect and the correlational associations discussed below are, at most, nominally significant and did not survive appropriate correction for multiple comparisons or for the non-independence of repeated observations (
Section 2). None of the eye-tracking or facial-coding metrics differentiated the three structural conditions at even the nominal level. This pattern is best summarized not as a confirmed dissociation between overt attentional/affective signals and evaluative judgment—the framing our uncorrected initial analyses suggested—but as a single robust effect on subjective complexity judgment, several weaker and largely non-robust signals elsewhere, and no detectable condition effect on overt attention or facial engagement under the present design. This calls for a substantial revision of the framework proposed in the Introduction, and, as elaborated below, is at least broadly consistent with prior work showing that attention and preference need not move together [
18,
19], though our own findings support that claim far more tentatively than we originally reported.
4.1. The Complexity Manipulation Did Not Translate into Perceived Complexity
The most striking result is that the high-complexity (HC) condition was rated as less complex than both the high-harmony (HH) and structured-complexity (SC) conditions—the opposite of the intended manipulation. Several non-mutually-exclusive explanations are possible. First, the HC stimuli, generated as disordered tessellations of small triangular elements, may have produced a visually homogeneous texture in which the repetition of similarly sized units was perceived as more uniform (and therefore less complex) than the layered, multi-scale symmetry of the HH and SC stimuli, which combine large and small elements and salient central motifs. This would be consistent with classic findings that fixation counts and durations scale with structural complexity as defined by the number of distinguishable sides or elements in a shape, rather than with disorder per se [
5], and with evidence that the perceived “goodness” of a pattern is closely tied to its degree of symmetry independent of its element count [
4]. This interpretation is also consistent with recent computational work showing that structural regularity, color variety, and semantic surprise are separable, non-redundant contributors to perceived complexity [
13]: our generative manipulation targeted structural disorder specifically, but disorder alone is evidently not synonymous with the complexity that participants reported perceiving. Irregular tessellation may increase disorder while paradoxically reducing the number of subjectively distinguishable units, thereby lowering rather than raising perceived complexity. Second, if symmetry detection itself depends on attentional resources rather than occurring preattentively [
8], the salient bilateral symmetry of the HH and SC stimuli may have drawn attention to their internal structure in a way that made their local detail more, not less, apparent to observers—again working against the intended complexity ordering. Third, without a pilot manipulation check on this specific stimulus set, the a priori classification of conditions (based on the generative geometric procedure) may not correspond to participants’ phenomenological experience of complexity. This underscores the importance of validating structural manipulations against subjective complexity ratings before—rather than only after—the main data collection.
It is important to be explicit about what this failed manipulation does, and does not, allow us to conclude. It does not allow us to conclude that “visual complexity has no effect on attention,” nor does it validate any general claim about complexity as a perceptual or aesthetic construct: what we manipulated and validly confirmed was a specific generative dimension (disordered, small-scale tessellation vs. large-scale bilateral symmetry), and what participants’ complexity ratings tracked evidently diverged from that dimension in the HC condition specifically. Every null or positive result involving the HC condition in this study should therefore be read as a finding about this particular stimulus manipulation and how it was subjectively construed, not as a finding about complexity in general. Conversely, the null eye-tracking and facial-coding results (
Section 3.1 and
Section 3.5) cannot be cleanly attributed to the complexity manipulation having failed, because the HH-vs-SC contrast—which does not depend on the disputed HC condition and did track the intended symmetry/motif manipulation—was equally null on every oculomotor and facial-coding measure. What can be concluded with reasonable confidence is narrower than our original framing suggested: this specific generative manipulation of structural regularity, whatever exact perceptual dimension it engaged, shifted subjective complexity judgments but did not produce a statistically detectable change in overt gaze or facial engagement under the present design.
4.2. Overt Attention and Facial Engagement Were Insensitive to Structural Condition
None of the fixation-based or facial-coding metrics distinguished the three conditions, and, as detailed in
Section 2, this design was powered post hoc to detect only medium-to-large effects (
); the observed effect sizes for every eye-tracking and facial-coding metric fell well below this threshold (
–
and
–
, respectively). Following standard guidance on interpreting non-significant findings [
26], we treat this pattern as an absence of detectable differences under the present sample size and measurement precision, not as evidence that visual structure has no effect on overt attention or facial engagement; a true effect in the small-to-medium range cannot be ruled out. Several factors specific to the present design may additionally account for this null pattern. The 6–7 s exposure window, while sufficient to capture initial orienting, may be too brief for global structural properties to shape the overall distribution of fixations or facial engagement, particularly if such effects unfold gradually over longer viewing periods. The use of webcam-based eye tracking and facial coding (approximately 30 Hz sampling) also entails greater spatial, temporal, and signal-to-noise limitations than laboratory-grade systems, which may have obscured comparatively subtle condition effects—reflected in the modest effect sizes observed even for the non-significant contrasts (
–
for eye-tracking metrics;
–
for facial-coding metrics). It is also possible that, for abstract geometric patterns lacking semantic anchors, low-level salience (contrast, color, local symmetry) exerts a stronger pull on early fixations than global structural regularity, effectively equalizing scanning and facial-engagement behavior across conditions despite their differing formal properties.
In earlier analyses of this dataset, we additionally argued that the coefficient
K’s correlation with independent oculomotor indices of focal processing supported the construct validity of the combined measurement approach, and thus indicated that the null condition effects were unlikely to reflect a wholesale failure of measurement. Having corrected those correlations for the non-independence of repeated observations (
Section 3.6), we no longer consider this argument well supported: none of the
K–oculomotor associations survives appropriate correction, and the association with fixation duration reverses sign once repeated-measures structure is accounted for. We therefore do not rely on
K’s convergent validity to argue that the measurement pipeline was functioning correctly; the case that the null condition effects reflect limited power and precision rather than a broken instrument rests instead on the sensitivity analysis above, together with the platform validation evidence discussed in
Section 2.3 [
22]. We regard the status of
K as a valid index of ambient/focal processing in this particular abstract, brief-exposure paradigm as an open question rather than a settled one.
4.3. A Fragile, Nominal Association Between Scanpath Entropy and Pleasantness
In our initial, uncorrected analyses, both spatial dispersion and scanpath entropy appeared negatively correlated with pleasantness ratings, which we interpreted as running counter to the exploratory-processing account proposed in the Introduction. Once repeated-measures correlation and correction for multiple comparisons are applied (
Section 3.3), the dispersion–pleasantness association is no longer distinguishable from zero, and only the entropy–pleasantness association remains nominally significant (
,
)—itself not surviving FDR correction across the full exploratory correlation family (
). We therefore no longer treat this pattern as an established empirical result, but discuss below how it should be interpreted if it reflects a genuine, if weak, effect, while emphasizing that this discussion is now explicitly speculative rather than an interpretation of a confirmed finding.
The negative correlation between scanpath entropy and pleasantness ratings, to the extent it reflects a real effect, runs counter to the exploratory-processing account proposed in the Introduction, in which distributed viewing was expected to accompany more positive aesthetic responses. This pattern also runs counter to a substantial body of literature reporting the opposite association—that more, and more prolonged, looking tends to accompany greater liking rather than less. The gaze cascade effect, for instance, shows that gaze is progressively biased toward the item that will eventually be chosen or rated as more attractive, implicating a positive feedback loop between looking and preference [
30]—a pattern independently found to replicate under webcam-based, browser-delivered eye tracking of the same general kind used here [
31], suggesting the discrepancy with our own results is unlikely to be a simple artifact of webcam-based measurement per se; in an applied advertising context, increased number and duration of fixations on an advertisement has similarly been found to correlate with more positive evaluations of it [
32]. Both of these literatures, however, involve either an explicit choice between competing alternatives (gaze cascade) or naturalistic, semantically rich stimuli (advertisements)—conditions that differ substantially from the single-stimulus, free-viewing paradigm with abstract geometric patterns used here, and in which the mechanism linking gaze to preference (value accumulation toward a decision, or engagement with meaningful content) may not readily apply.
Closer to the present paradigm, single-stimulus free-viewing studies of pleasantness are more mixed. In an eye-tracking study of nature and urban scenes, images rated as more pleasant elicited fewer fixations but longer individual fixations [
33], a pattern that, in terms of fixation count at least, aligns with the direction we observed rather than the increased-exploration account. More broadly, reviews of fixation-duration operationalizations in aesthetic and appreciation research caution that longer looking does not unambiguously index greater liking, because processing difficulty—arising from lower visibility, greater clutter, or a harder-to-resolve percept—can independently prolong fixation and viewing time on stimuli that are not necessarily preferred [
34]. This is one alternative interpretation worth entertaining, if the entropy–pleasantness association reflects a genuine effect rather than a false positive: broader, less concentrated scanning (higher scanpath entropy) may reflect difficulty resolving a stimulus into a coherent percept—an uncertainty-driven search pattern—whereas preferred stimuli may instead sustain more focal engagement on a small number of privileged regions. This would be consistent with processing-fluency accounts of empirical aesthetics, in which stimuli that are easier to process (and often preferred) are associated with more efficient, less dispersed visual exploration [
16], and with Berlyne’s classic proposal that moderate, easily resolved arousal potential—rather than the sheer amount of exploratory activity a stimulus elicits—underlies hedonic value [
14].
Reconciling these conflicting literatures is beyond the scope of a single study, but the contrast is informative: gaze-cascade and advertising findings involve either explicit choice or content-rich stimuli where looking plausibly serves information gathering toward a decision or engagement with meaningful features, whereas the present free-viewing task with abstract, non-representational patterns and no choice requirement may instead be dominated by processing-fluency dynamics, in which dispersed, effortful scanning signals a percept that resists easy resolution. Given that our own evidence for this pattern is now nominal at best, we present this as one plausible account to be tested, not as an established finding. Future work directly manipulating task demands (choice vs. passive viewing) and stimulus meaningfulness (abstract vs. representational) within the same paradigm, with a sample size and correlational approach that accounts for repeated measures from the outset, would help establish whether an association of this kind exists and, if so, which boundary conditions govern its sign.
4.4. A Tentative Parallel with Dissociations Reported Elsewhere
The pattern found here—a robust effect on subjective complexity judgment alongside no detectable condition effect on fixation-based or facial-coding measures—is at least broadly compatible with a dissociation documented elsewhere between covert or overt attentional signals and explicit evaluative preference, though we emphasize that our own null oculomotor and facial-coding results should be read as an absence of detectable effects under this design (
Section 4), not as a demonstrated dissociation in the strong sense. In face perception, covert and overt measures of attention to facial attractiveness have been shown to diverge according to task demands, with reliable oculomotor preferences emerging only under some measurement conditions [
18]. In developmental work on symmetry, four-year-old children reliably fixate symmetrical patterns longer than asymmetrical ones yet show no corresponding explicit preference for them, directly dissociating spontaneous visual attention from aesthetic judgment [
19]. More broadly, the symmetry-preference literature suggests that the link between fixation behavior and aesthetic evaluation is itself categorically variable, holding for some stimulus domains but not others [
6], and that spontaneous oculomotor signatures of structural processing (e.g., axis-aligned scanning of symmetrical patterns) can be similar across classification and evaluation tasks even when explicit preference diverges from perceptual sensitivity [
7]. The present results are consistent with an extension of this general pattern to adult observers viewing abstract geometric stimuli—structural properties shifted what people said they preferred (for complexity, robustly so) without producing a statistically detectable signature in either oculomotor or facial-expressive channels under the present design—but given the corrections applied in this revision, we hold this extension more tentatively than in our initial analysis, and present it as a hypothesis the present data are consistent with rather than one they confirm. This raises the possibility that evaluative judgments of harmony and complexity are computed, at least in part, through processes that are not fully indexed by overt gaze allocation or momentary facial affect—a possibility that merits direct investigation with adequately powered samples, repeated-measures-aware analyses from the outset, and converging measures (e.g., pupillometry, EEG) in future work.
4.5. Two Contributions of This Study, Held to Different Standards of Evidence
We close the Discussion by returning explicitly to the dual contribution flagged in the Introduction, because the two halves of this study now rest on different standards of evidence and should be weighted accordingly by the reader. The substantive contribution—that structural disorder, as we operationalized it, does not translate straightforwardly into perceived complexity, and instead produced a reversed ordering—is a single-experiment finding from one specific stimulus set and should be treated as a hypothesis for the field to test further (
Section 5, item 1), not as a general claim about visual complexity. It is, however, the one primary result in this study that survives our strictest correction for multiple comparisons, and it converges with an independent, purely computational literature indicating that structural regularity is only one of several separable contributors to perceived complexity [
13].
The methodological contribution is, in our view, on firmer ground precisely because it does not depend on any single substantive effect being real. Analyzing this dataset surfaced two analytic problems that are neither unique to our data nor, we suspect, unique to this literature: correlating repeated within-participant observations without accounting for their non-independence, and reporting a large family of eye-tracking, facial-coding, and subjective outcomes without a stated multiplicity strategy. Both problems are straightforward to introduce inadvertently in exactly the kind of multi-condition, multi-exemplar, multi-measure webcam-based eye-tracking design that is becoming increasingly common as platforms like RealEye lower the barrier to running such studies remotely and at scale [
23,
24]. We report, in
Section 3 and
Section 4, a concrete demonstration of how much a study’s conclusions can change once these two corrections are applied—including the reversal of one association and the loss of significance for several others we had initially interpreted as supporting our measurement approach—together with the specific procedures used (repeated-measures correlation via an ANCOVA-equivalent formulation, and Benjamini–Hochberg FDR correction across the exploratory family designated in
Section 2). We offer this worked example, independent of whether the substantive complexity finding itself replicates, as a practical resource for researchers designing or reanalyzing similarly structured studies.
5. Conclusions
Controlled variations in visual harmony and complexity, isolated from semantic content, shaped participants’ subjective complexity judgments—the only one of seven primary, family-wise-corrected omnibus tests to survive correction—but did not produce statistically detectable differences in overt attentional allocation, as measured by fixation counts, durations, latencies, and spatial/entropy-based dispersion indices, nor in facial-coding measures of expression and ambient/focal engagement. The complexity manipulation itself did not translate into the intended perceptual ordering of conditions: what we can conclude is narrower than we originally framed it, and concerns this specific generative manipulation rather than visual complexity as a general construct (
Section 4). A nominally significant effect on pleasantness ratings, and nominally significant correlations between eye-tracking/facial-coding metrics and subjective ratings—including the coefficient
K’s associations with independent oculomotor indices, which we had initially read as supporting the validity of our combined measurement approach—did not survive correction for the non-independence of repeated observations and for the number of statistical tests performed, and are reported here as exploratory rather than confirmatory findings. Together, these findings suggest, more tentatively than our initial analysis proposed, that formal visual structure may influence evaluative, higher-order aesthetic responses more readily than early overt attentional or facial-affective deployment, at least under brief, free-viewing conditions with webcam-based measurement—a pattern broadly consistent with, though not decisive evidence for, the dissociation between attention and preference documented in other domains [
18,
19]. The absence of detectable oculomotor and facial-coding effects should not be read as evidence that visual structure has no effect on attention [
26]; rather, it reflects what a post-hoc-determined, webcam-based, brief-exposure design with
was, and was not, equipped to detect.
Based on the specific limitations identified throughout this study, we consolidate the following directions for future research:
Independently validate stimulus manipulations before data collection. A dedicated pilot study should collect subjective complexity and harmony ratings for candidate stimuli and use them to select or adjust the final exemplars, rather than relying solely on the a priori generative parameters used here (
Section 2 and
Section 4).
Conduct a prospective power analysis. Future replications should determine the target sample size before data collection, using the effect sizes and repeated-measures correlation structure reported here (
Section 3.4 and
Section 3.3) rather than relying on a post-hoc sensitivity analysis as this study necessarily did.
Increase spatial and temporal measurement precision. Replicating the design with laboratory-grade infrared eye tracking (sub-degree accuracy, ≥250 Hz) would clarify whether the null oculomotor effects reflect a genuine absence of structural modulation or the coarser precision of webcam-based recording (
Section 2.3).
Extend and vary exposure duration. Systematically comparing brief (<10 s) and extended (e.g., 20–30 s) free-viewing windows within the same design would establish whether global structural effects on gaze require more time to emerge than local salience-driven orienting.
Increase the stimulus set and statistical power. A larger item pool (beyond three exemplars per condition) together with a sample size informed by the sensitivity analysis reported here (
Section 3.4) would allow detection of the small (
) effects that the present design could not rule out.
Define areas of interest a priori. Region-specific measures (e.g., time to first fixation on the central motif versus peripheral facets) may be more sensitive to the harmony/complexity manipulation than whole-stimulus metrics.
Directly test the boundary conditions of the entropy–pleasantness relationship. Manipulating task demands (explicit choice vs. passive free viewing) and stimulus meaningfulness (abstract vs. representational) within a single paradigm would help establish whether the nominal association observed here is genuine and, if so, how it relates to the positive associations reported in gaze-cascade and advertising research (
Section 4).
Incorporate converging psychophysiological measures. Pupillometry and EEG could help determine whether evaluative judgments of harmony and complexity are computed through processes that are not fully indexed by overt gaze allocation or momentary facial affect.
Pre-register the analysis plan, including the correlational approach and multiplicity strategy. Specifying primary versus exploratory outcomes, the repeated-measures correlation method, and the multiple-comparison correction procedure in advance would remove the need for the kind of post-hoc reanalysis undertaken in this revision and would strengthen confidence in whichever pattern of results is obtained.