1. Introduction
Eye tracking is one of the most direct windows into human visual cognition. The gaze information channel paradigm [
1] models eye-tracking fixation sequences between Areas of Interest (AOIs) as a first-order Markov chain and derives a discrete information channel from the transition matrix. The stationary entropy
quantifies how uniformly the observer distributes gaze across regions; the conditional entropy
quantifies the randomness of gaze transitions; and the mutual information
measures sequential dependence.
A fixation-to-fixation trajectory carries information along several complementary dimensions simultaneously: which spatial region the gaze lands in (AOI), how long the fixation lasts (duration), how large the subsequent saccade is (amplitude), and in which direction the saccade travels (direction). Hao et al. introduced both the fixation duration channel [
2] and the saccade direction channel [
3].
This paper makes two methodological contributions: it introduces the saccade amplitude channel and the pupil diameter channel. These two channels were selected on complementary grounds. Saccade amplitude complements the previously introduced direction channel and represents the magnitude component of the saccadic displacement vector, constituting a natural next step in the same geometric decomposition. Pupil diameter, by contrast, is the only channel that is not purely spatial or temporal in nature: it is often associated with arousal and cognitive load [
4] and therefore introduces a qualitatively new physiological dimension into the multi-channel framework, complementing the four purely oculomotor channels. Prior work on saccade amplitude has focused on its marginal distribution [
5,
6] or on high-order Markovian patterns of short and long saccade lengths (pixel displacement) without computing
,
, or
[
7], and prior work on pupil diameter has treated it as a scalar physiological signal [
4] or used a temporally quantised marginal distribution [
8], rather than modelling the sequence of fixation-period pupil states as a communication channel. With all five channels computed on the same dataset, we perform a simultaneous assessment of cross-channel association among their
summary statistics, using Spearman rank correlations computed both across participants and across paintings, with Bonferroni correction for multiple comparisons.
We also derive two theoretical observations (Proposition 1 and Remark 2): a simple upper bound on the conditional entropy in terms of the expected self-transition probability, and a refinement monotonicity result showing that state-space refinement cannot decrease channel mutual information.
5. Discussion
The robustness of the amplitude channel across two categorisation schemes confirms that the sequential structure it captures is not an artefact of the particular bin boundaries chosen, a property guaranteed theoretically by Remark 2 and verified empirically by the near-identical normalised MI values (
Section 4.6). The dominance of medium–medium self-transitions is consistent with the ambient/focal framework of visual attention [
10], suggesting that observers tend to maintain their saccade amplitude regime across consecutive fixations rather than alternating freely between scales.
The link between transition persistence and mutual information, formalised in Proposition 1, is illustrated across all five channels in
Figure 13, which plots
(vertical axis) vs. the expected self-transition probability
(horizontal axis).
We verified the Proposition 1 bound
numerically for all five channels using the grand-pooled transition matrices;
Table 8 shows the actual
, the bound, and the slack (bound minus actual) for each channel.
The slack quantifies how far the actual channel is from the worst-case (maximally diffuse off-diagonal structure) assumed by the bound. The bound is tightest for the pupil channel (slack
bits), where high persistence leaves little room for off-diagonal entropy, and for the direction channel (slack
bits), which has eight states and a relatively diffuse transition structure. The bound is loosest for AOI (slack
bits): despite moderate
, its rich spatial composition produces a highly structured off-diagonal pattern that keeps
well below the worst case, illustrating that
alone does not determine
—off-diagonal structure matters too (Remark 1). The five channels occupy well-separated regions along the
axis (
Figure 13): direction (
) and amplitude (
) have the lowest persistence; AOI (
) and duration (
) have intermediate persistence; and the pupil channel (
) has the highest persistence and the highest MI among the four non-AOI channels. This between-channel pattern is broadly consistent with the expectation that higher persistence is associated with greater sequential dependence, with AOI as the notable exception owing to its large state space. A noteworthy detail is that amplitude (
,
) and duration (
,
) are the two lowest-MI channels, both well below the direction, pupil, and AOI channels, despite representing different physical quantities—consistent with both being components of the ambient/focal scanning mode captured by Krejtz et al.’s
coefficient [
10]. Their cross-channel Spearman correlation is nonetheless not significant (participants:
,
; paintings:
,
), confirming that they carry similar
amounts of sequential structure but not the
same structure. Within each channel,
varies only modestly across paintings and participants, so the dominant source of MI variation within a channel is the shape of the off-diagonal transition structure rather than persistence alone.
The pupil diameter channel shows the second highest
of all five channels (after the AOI channel), which is consistent with the slow physiological dynamics of pupil dilation [
14]. Pupil size reflects both psychosensory and luminance-driven processes, which unfold over seconds and are thus inherently persistent across consecutive fixations; this physiological timescale explains why self-transition probabilities are very high across both aggregations (around 0.78 for Small, 0.88 for Medium, and 0.90 for Large, nearly identical across participants and paintings) and why
substantially exceeds the amplitude, duration, and direction channels. The near-zero Pupil–Duration painting-level correlation (
) stands in contrast to the nominally significant participant-level correlation (
). The sensitivity analysis shows that this participant-level result is not robust, losing significance upon removal of P01 (
,
).
The overall pattern of weak and statistically non-significant cross-channel associations suggests that the channels may capture partially distinct aspects of gaze behaviour, although the present sample size does not allow strong conclusions regarding channel independence.
A connection worth noting is that the amplitude and duration channels together correspond to the components of Krejtz et al.’s
coefficient [
10], which quantifies the temporal correlation between fixation duration and subsequent saccade amplitude as a marker of the ambient–focal distinction. Despite this conceptual link, the Spearman correlation between duration and amplitude
values is not significant at either level (participants:
,
; paintings:
,
), suggesting that the channel measures capture different aspects of these dimensions than the scalar
statistic does.
The relationship between channel measures and computational aesthetics quantities [
1] is examined in
Section 4.11. Notably, the pupil channel shows a nominally significant negative correlation with Bense’s palette redundancy
(
,
, uncorrected), the AOI channel shows a marginal positive correlation with compositional complexity
(
,
), and the direction channel shows a marginal positive correlation with
(
,
). By contrast, neither permutation entropy
nor statistical complexity
shows a significant association with any channel, likely because all 12 Van Gogh paintings occupy a compressed high-entropy range that limits between-painting discrimination.
6. Limitations
Several limitations of this study should be noted.
Sample size. The dataset comprises 10 observers and 12 paintings. The small N limits statistical power for cross-channel association tests ( or per Spearman correlation) and means that individual outlier participants can substantially influence results, as shown by the sensitivity analysis for the Pupil–Duration pair. Replication with larger and more diverse samples is necessary before drawing strong conclusions.
Single dataset and stimulus type. All 12 stimuli are paintings by a single artist. The channel properties—particularly the finding that amplitude discriminates observers better than paintings while pupil discriminates observers better than paintings—may not generalise to other stimulus types (natural scenes, faces, diagrams) or tasks (visual search, reading).
Luminance confound in the pupil channel. As noted in
Section 3.3, fixation-period pupil size reflects a mixture of psychosensory arousal and the pupil light response to local luminance [
14]. Without per-fixation luminance estimates, these contributions cannot be separated. The pupil–
correlation (
) should therefore be interpreted with caution: it may partly reflect structured luminance transitions in paintings with a more ordered, limited palette (high
) rather than purely cognitive differences in arousal.
Choice of channel input order. The channel
constructed in this paper is a valid information channel regardless of whether the underlying gaze process is first-order Markov or not: the transition matrix
and the mutual information
are well-defined descriptors of the sequential structure between consecutive fixation states. The question of Markov order is therefore not one of validity but of completeness: a second-order channel with compound input
captures strictly more sequential dependence than the first-order channel, since by the data-processing inequality
. Bonev et al. [
7] reported empirical evidence of higher-order Markovian patterns in saccade length sequences (pixel displacement), suggesting that a second-order channel may capture additional sequential dependence beyond what is reported here. The standard tool for testing whether this additional structure is statistically detectable is the asymptotic likelihood-ratio G-test, but for the second-order amplitude channel (
transition matrix, 27 cells) the present dataset is too sparse to satisfy the Cochran condition for the
approximation in most units, so this question cannot be answered reliably here. A larger dataset would allow both reliable G-test evaluation and more stable estimation of the second-order channel MI.
Discretisation choices and pupil threshold bias. All channel measures depend on the choice of category boundaries. Although the amplitude channel showed robustness across two categorisation schemes (Remark 2 guarantees that finer discretisation cannot decrease MI), alternative binning strategies could alter absolute MI values and transition structures. The choice of boundaries for the pupil channel was motivated by the empirical distribution of the data; different datasets or viewing conditions may require different thresholds. More importantly, fixed absolute thresholds are a limitation for the pupil channel because individual baseline pupil diameter varies with age, iris pigmentation, and ambient illumination [
14]. Although the present experiment used constant room illumination and a single recording session per observer, the absolute categories (Small/Medium/Large) may not represent equivalent physiological states for all participants. A more principled alternative would be to normalise each observer’s pupil diameters (e.g., by per-observer z-score or tertile split) before assigning categories, producing equal-frequency states per observer and eliminating inter-individual baseline differences. Adopting this participant-normalised scheme would not be expected to alter the qualitative pattern of results (pupil showing the highest MI among the two new channels) but would change the exact numerical values and potentially the cross-channel correlation results; this is noted as a direction for future work.
Multiple comparisons. The exploratory correlation analyses in
Section 4.11 involve 25 pairwise tests across five channels and five aesthetic measures. With a nominal significance level of 0.05 (uncorrected), up to one or two false positives would be expected by chance alone. The single nominally significant result (Pupil–
,
) does not survive Bonferroni correction (threshold
) and should therefore be treated as hypothesis-generating rather than confirmatory.
7. Conclusions
This paper introduced two new gaze information channels: the saccade amplitude channel and the pupil diameter channel. Applied to 10 observers viewing 12 Van Gogh paintings:
The amplitude channel shows that observer-driven variation exceeds stimulus-driven variation in this dataset, in contrast to the AOI channel.
The pupil channel shows the highest among the two new channels (participant mean 0.489 bits, painting mean 0.746 bits), which is consistent with the slow dynamics of pupil responses.
Two theoretical observations (Proposition 1 and Remark 2) provide a simple upper bound on the conditional entropy in terms of transition persistence, and show that state-space refinement cannot decrease channel mutual information, directly explaining the 3-category to 4-category increase in amplitude MI.
Goodness-of-fit tests confirm significantly non-random sequential structure in both channels at for all pooled matrices.
Of the 20 pairwise Spearman correlations across five channels, 19 did not reach statistical significance, suggesting that the channels may capture partially distinct aspects of gaze behaviour, although the present sample size does not allow strong conclusions regarding channel independence; the single nominally significant result (Pupil–Duration, ) does not survive Bonferroni correction and is not robust to removal of one outlier participant.
In an exploratory comparison with five computational aesthetics measures, the pupil channel shows a nominally significant negative correlation with Bense’s palette redundancy (, , uncorrected, does not survive Bonferroni correction): paintings with more diverse colour palettes are associated with stronger sequential pupil dynamics, though this result is hypothesis-generating only, and the contribution of luminance-driven pupil responses cannot be separated from cognitive effects without per-fixation luminance data. The AOI channel shows a marginal positive association with compositional complexity (, ), and the direction channel with Kolmogorov compressibility (, ). Permutation entropy and statistical complexity show no significant association with any channel, consistent with the narrow high-entropy range occupied by all 12 Van Gogh paintings.
Future work should pursue several directions. First, both channels should be validated on larger and more diverse datasets spanning different stimulus types (natural scenes, visual search, reading), to establish whether the observer-versus-stimulus discrimination pattern reported here generalises beyond art viewing. Second, the Pupil–Duration association should be tested with an independent measure of arousal or cognitive load (e.g., task-evoked pupillary response, skin conductance, or subjective workload rating) to determine whether the shared variance is genuinely cognitive or partly artifactual. Third, a luminance-corrected version of the pupil channel should be computed once per-fixation luminance values are available, to disentangle psychosensory and light-response contributions to the sequential structure. Fourth, extending the channel input from
to the compound state
would define a second-order channel that, by the data-processing inequality, captures at least as much sequential dependence as reported here. Whether this additional structure is practically detectable requires a larger dataset: the second-order amplitude channel (
transition matrix) is too sparse for reliable G-test evaluation in the present data [
7]. Finally, additional channels—such as saccade velocity or blink rate — could further enrich the multi-channel framework.