1. Introduction
The human ability to perceive depth is not only a physiological necessity but also a rich perceptual phenomenon that has long intrigued artists and philosophers. While binocular vision plays a well-documented role in spatial perception, a great deal of visual depth is also inferred from monocular cues such as light, shadow, occlusion, and linear perspective
Goldstein (
2001);
Howard and Rogers (
2012). These cues are of particular relevance in the realm of visual art, where a two-dimensional surface must suggest a three-dimensional world. Artists have historically exploited these perceptual mechanisms to engage the viewer in a dynamic visual experience. In the context of our study, it is useful to distinguish three related notions of depth perception. First, we use
normal stereopsis to refer broadly to the human ability to recover depth from all available visual information, including both binocular and monocular cues. Second,
binocular stereopsis designates the more specific case in which depth arises from binocular disparity, that is, the systematic differences between the retinal images in the two eyes. Third,
phantom stereopsis describes a class of depth experiences that can occur even in the absence of binocular disparity, when occlusion relationships, shading, and other monocular cues are sufficient to evoke a compelling sense of spatial organization
Nakayama and Shimojo (
1990). It is within this latter domain of phantom stereopsis that the present work is situated. Among these techniques, the use of subtle shading and occlusion, what has come to be referred to as Da Vinci stereopsis, offers a compelling case for the interplay between illusion and interpretation
Nakayama and Shimojo (
1990). In this article, we use the term Da Vinci stereopsis in the sense proposed by Nakayama and Shimojo, namely as a monocular depth effect that relies on specific occlusion and shading configurations rather than binocular disparity. Rather than depending on binocular disparity, Da Vinci stereopsis relies on the careful orchestration of occlusion, shadow, and spatial ambiguity within a single pictorial view. Our analysis focuses on how such cues are embedded in the painted surface and how they may give rise to viewer-dependent depth experiences. Leonardo da Vinci’s ability to evoke spatial presence through tonal transitions and visual ambiguity has inspired centuries of artistic inquiry. His application of methods such as
sfumato and atmospheric perspective introduced a new kind of engagement: one in which the image seems to shift or breathe depending on the position of the observer.
Our previous work
Wen and Khatibi (
2016) explored this perceptual effect by attempting to quantify depth illusion in a classical painting using manual annotation and simple geometric relationships. While this provided preliminary evidence of measurable depth phenomena in painted surfaces, the method relied heavily on subjective point selection and lacked fine-grained visual precision.
In this study, we expand upon that foundation by introducing a refined image-based analysis method designed to capture subtle perceptual shifts in pictorial space. Through a series of photographs taken from slightly varied viewpoints, we use image alignment and pixel-level difference analysis to reveal how classical artistic techniques, particularly shadowing and tonal gradients, can evoke depth that changes with the viewer’s position.
By bridging computational tools and art historical concepts, we aim to offer a framework that respects the interpretive richness of Renaissance art while contributing empirical insight into its perceptual dynamics. This work seeks to demonstrate that illusionistic depth is not only a matter of technique but also a dialogue between viewer, artwork, and visual cognition.
2. Related Work
The illusion of depth in classical painting has long captivated both artists and scholars. Far from being a modern phenomenon, the exploration of spatial perception on a flat surface lies at the core of Renaissance artistic innovation. Among the most celebrated figures in this tradition is Leonardo da Vinci, whose works exemplify the use of visual techniques that evoke a profound sense of three-dimensionality. Central to these techniques are three artistic devices that have been extensively studied:
sfumato,
chiaroscuro, and
linear perspective Brown (
1998);
Clark (
1989);
Kemp (
1981).
2.1. Sfumato and the Evocation of Space
Sfumato, derived from the Italian “sfumare,” meaning “to fade out”, is a technique that produces soft transitions between tones and colors, eschewing hard edges or outlines. Da Vinci’s use of sfumato, most famously in the Mona Lisa, blurs the boundary between subject and atmosphere, allowing forms to emerge gently from the surrounding space. Art historians such as
Clark (
1989);
Kemp (
1981) have emphasized how this technique encourages immersive viewing: the ambiguity of form invites the observer’s eye to continually reinterpret spatial relationships, especially when the viewer’s position shifts.
This phenomenon has been described as “dynamic engagement,” where the painting seems to subtly transform as one moves, producing the sensation of a living presence rather than a static image. Such perceptual effects resonate with contemporary interest in visual ambiguity and viewer-centered interpretation.
2.2. Chiaroscuro and the Modeling of Form
Chiaroscuro, the interplay of light and shadow, has long been used to create the illusion of volume and depth. While often associated with Baroque painters like Caravaggio, da Vinci employed a more delicate and naturalistic form of chiaroscuro. His shadowing was not dramatic but continuous, revealing subtle curvature and depth in the human form. As highlighted in recent visual analyses
Jalandoni et al. (
2024);
Jue et al. (
2016), this method contributes not only to spatial realism but also to psychological expressiveness. The careful modulation of light invites a deeper emotional and perceptual reading of the figure.
2.3. Perspective and Viewer Involvement
The mathematical organization of pictorial space, particularly through linear perspective, constitutes a fundamental component of Renaissance strategies for constructing depth. Leonardo’s architectural studies and works such as The Last Supper demonstrate a sophisticated command of perspective as a means of structuring space and directing the viewer’s gaze. In these works, perspective functions not merely as a geometrical system, but as a compositional device that shapes how the image is apprehended as the observer’s position changes.
Art-historical scholarship has long emphasized that this spatial construction is closely tied to the viewer’s active participation.
Clark (
1989) describes how the soft transitions and lack of sharply defined contours in Leonardo’s mature works encourage a continual reconfiguration of forms by the moving eye, while
Kemp (
1981) has shown that Leonardo’s compositional schemes often presuppose a viewer whose position is not fixed. More recently,
Vartanian and Chatterjee (
2013) have suggested that such perceptual instability may play a central role in aesthetic engagement, prompting the viewer to actively negotiate spatial and formal relationships within the image.
Taken together, these accounts point toward an understanding of Renaissance painting in which depth and spatial meaning emerge through an interaction between pictorial structure and viewer movement. The present study engages with this perspective by focusing on how changes in viewpoint relate to variations in tonal structure, offering a basis for examining viewer involvement at the level of observable image behaviour rather than interpretation alone.
2.4. Perception and Illusion in Contemporary Analysis
Modern theories in visual perception and neuroaesthetics have also examined the relationship between visual ambiguity and viewer engagement. Concepts such as phantom occluders and Da Vinci stereopsis explore how single-eye cues, like unpaired, can evoke depth perception
Nakayama and Shimojo (
1990). Scholars including
Vartanian and Chatterjee (
2013) argue that tonal ambiguity stimulates active mental reconstruction by the viewer, heightening aesthetic experience.
Several artworks discussed in the art-historical literature illustrate how such phenomena may manifest in practice. Examples include Jacopo Tintoretto’s Finding of the Body of St. Mark, in which the apparent spatial arrangement of figures and architecture shifts as the viewer moves laterally, and the much debated Salvator Mundi, attributed by some to Leonardo or his workshop, where the positions of the mouth and eyes have been reported to appear subtly mobile under changing viewpoints. These cases exemplify the kind of viewer-dependent instability that concerns us here and provide a qualitative counterpart to the quantitative analysis undertaken in this study.
While these ideas have typically been discussed in theoretical terms, our research applies digital tools to analyze how classical painting techniques generate measurable perceptual effects. In doing so, we aim to bridge art historical inquiry with empirical observation, offering a fresh perspective on an enduring question: how does a static image come alive in the eye of the beholder?
3. Materials and Methods
3.1. Artwork and Observational Setup
This study centers on the classical painting known as the youngest Madonna Nativity (Nativity portrait), held in the Kulenović Collection in Karlskrona, Sweden. The painting was selected for its atmospheric treatment of light, soft tonal transitions, and the perceptual ambiguity that many observers report when viewing it from different positions.
The Nativity portrait depicts a Madonna-and-Child subject rendered with pronounced chiaroscuro and softly graded shadows, characteristics that lend the work a visually ambiguous spatial structure. According to prior analyses (e.g.,
Wen and Khatibi (
2016)), the painting exhibits viewer-dependent depth effects that are noticeable even under naked-eye observation. The work is in a stable state of conservation within the Kulenović Collection, and its tonal modeling and shadow transitions remain well preserved—properties that are central to the stimulus of phantom-type Da Vinci stereopsis cues observed in previous experimental studies.
Our choice of this painting is motivated not by any specific attributional connection to Leonardo but by the documented presence of depth-like perceptual effects generated through shadow-based modelling. Earlier measurements on this same portrait
Wen and Khatibi (
2016) showed that its shadows produce measurable point-displacement and line-ratio variations across viewpoints, consistent with phantom-type Da Vinci stereopsis phenomena. These properties make the Nativity portrait a suitable test case for the present computational analysis of viewer-dependent tonal change.
All photographs were taken under deliberately controlled and carefully stabilised illumination conditions. While the exact lighting equipment was not recorded, attention was paid to ensuring that the light falling on the painting remained as uniform as possible across all viewpoints and that no directional reflections or specular highlights appeared on the surface. Exposure settings on the camera were kept constant throughout the acquisition sequence. As an additional precaution, we later verified the homogeneity of the captured images through histogram comparison (see
Section 4.3), which showed only minimal global tonal variation between views. Although viewers in a gallery typically move both laterally and in depth, our experiment isolates lateral displacement in order to maintain constant viewing distance and to focus exclusively on the effects of horizontal viewpoint change on the painting’s perceived depth structure.
3.2. Framing the Gaze: Image Preparation and Region Selection
Each photograph was first converted to grayscale using a perceptually tuned transformation to isolate tonal structure over color. We then extracted a consistent region of interest from all five images, focused on the sitter’s face and upper body, where form, shading, and gaze converge. Because the camera shifted laterally between exposures, each image was manually cropped with adjusted offsets to ensure the same facial region was framed, compensating for parallax and viewpoint change (
Figure 1).
3.3. Viewpoint Registration and Alignment
To enable a true comparison of perceptual subtleties, the cropped facial regions from each viewpoint were brought into precise correspondence. The alignment followed an intensity–based strategy: rather than anchoring to hard anatomical landmarks, each image sought its nearest echo in the tonal fabric of the next, shifting gently until the play of light and shadow resonated most closely across the pair. In effect, we let the painting’s own luminance patterns guide the overlap, an approach well established in image registration research, where correspondence is inferred from the distribution of grayscale intensities rather than from explicit features
Zitová and Flusser (
2003).
This alignment proceeded sequentially (first to second view, second to third, and so on), mirroring the natural tempo of a viewer’s lateral movement before the work. Under the hood, the correspondence criterion follows the same family of mutual-information-driven methods that have become standard in contemporary registration toolkits, prized for their ability to reconcile small changes in appearance while preserving overall structure
Insight Toolkit (ITK) Documentation (
2025). The result is a shared visual frame in which the slightest shifts in shading, depth cues, or gaze become legible, unclouded by incidental differences in camera position, so that the painting’s subtle spatial choreography can be read with clarity (see
Figure 2 for example of viewpoint alignment).
3.4. Mapping Visual Change Across Viewpoints
With the images aligned, we examined how the painting’s visual fabric transforms as the observer shifts position. For each consecutive pair of views, we computed a pixel–wise difference map, revealing where tonal relationships had altered between vantage points. These raw differences were then normalized, allowing perceptually significant changes to stand out, regions of greater variation appearing as luminous passages, signalling where form, shadow, or highlight respond most to the change in gaze.
Yet in their unrefined state, such maps can be crowded with incidental texture and surface noise, obscuring the larger compositional gestures. To address this, each difference image was refined using the
L0 gradient minimization technique introduced by
Xu et al. (
2011). This method, originally conceived for image smoothing, delicately suppresses fine–scale detail while preserving the decisive boundaries that define form. In doing so, it mirrors the way the human eye privileges edges and contours over minute texture when interpreting spatial depth. For comparison, the difference maps without smoothing showed elevated high-frequency texture and weaker edge continuity;
L0 smoothing reduced these artifacts while preserving boundary salience, consistent with
Xu et al. (
2011).
The result is a series of clarified change–maps in which the principal transitions of light and shadow emerge with heightened legibility, revealing the subtle choreography by which the painting’s depth and presence shift as one moves before it. In simple terms, the CPF maps summarise how local tonal variations accumulate across views, highlighting broader regions of perceptual activity. Rather than proposing a new kind of depth, the method makes visible the regions where the existing tonal modelling is most sensitive to changes in vantage point.
3.5. Emergent Structural Phenomena
Beyond the isolated moments of difference between single viewpoints, it is the accumulation of these changes that reveals the painting’s deeper choreography. To capture this unfolding visual rhythm, we constructed what we call Cumulative Perceptual Fields (CPF). Each CPF is created by summing the normalized difference maps from two consecutive intervals, allowing small, local shifts to gather into broader waves of perceptual activation.
In our analysis, three distinct fields emerged:
CPF-A (First Resonance Field): Formed from the earliest transitions (View 1 to 2, and 2 to 3), this field illuminates the first stirrings of perceptual activity. Here, luminous traces gather along the contours of the face and in the gentle gradients of shadow, as if the image is sensing the viewer’s approach.
CPF-B (Mid-Sequence Amplification): Drawn from the central progression (View 2 to 3, and 3 to 4), this field reveals a swelling in the painting’s visual voice. Tonal contrasts intensify, edges assert themselves, and certain regions seem to pulse with heightened presence, suggesting a midpoint climax in the dialogue between artwork and observer.
CPF-C (Late-Sequence Emergence): Composed from the final transitions (View 3 to 4, and 4 to 5), this field marks a gentle dispersal of attention. Depth cues shift outward, peripheral forms reawaken, and the composition seems to guide the gaze toward departure, offering a lingering afterimage of presence.
The tripartite division into CPF–A, CPF–B, and CPF–C is defined directly by the temporal structure of the image sequence: each field sums two adjacent difference maps (Views 1–2 and 2–3; 2–3 and 3–4; 3–4 and 4–5, respectively). To ensure that these stages correspond to meaningful changes rather than arbitrary groupings, we computed summary statistics (mean and standard deviation of normalized intensity change) within each field. CPF–B shows the largest overall variance, indicating that the central portion of the sequence does indeed correspond to a peak in depth-related activity, flanked by lower-intensity “early” and “late” stages.
Seen together, these fields trace an arc of engagement: an awakening, an intensification, and a release. The painting appears to breathe in slow motion, its surface animated not by pigment alone, but by the subtle theatre of shifting perspectives. In this way, the CPF patterns do more than record change, they reveal the painting’s hidden capacity to respond, to invite, and to transform with the movement of the gaze. For readers less familiar with digital image analysis, the Cumulative Perceptual Fields can be understood as summaries of “where things happen” in the image sequence. By adding together successive difference maps, we track how small, localised changes accumulate into broader zones of activity, much as an art historian might mentally register where a painting seems most alive when viewed from different angles.
4. Results
4.1. Visual Alignment and Continuity Across Views
The alignment process successfully registered the sequence of five photographic perspectives, allowing for a consistent region of interest to be compared across views. The transitions between aligned image pairs were smooth and coherent, demonstrating that the lateral movement of the camera did not introduce distracting distortions or misalignments. The aligned sequence is shown in
Figure 3, illustrating the consistency of the facial region across viewpoints. This ensured that any visual differences between images could be meaningfully attributed to the perceptual effects created by the painting itself, rather than to flaws in the photographic method. To verify that alignment effects reflect perceptual change rather than residual geometry, we computed a similarity index between each pair of consecutive views. Across the sequence, mean normalized cross-correlation increased from
to
after alignment, indicating a consistent improvement in tonal correspondence within the facial region.
4.2. Difference Maps and Perceptual Shifts
To explore how the image content evolves with the viewer’s position, we generated pixel-wise difference maps between each successive pair of aligned photographs. These maps visualize the local changes in tonal values and luminance structure across views. These effects are visualized in
Figure 4, where normalized difference maps highlight perceptual shifts between successive views. Although the subject remains static, the difference maps reveal systematic variations, particularly in areas rich in soft shading and gradient transitions.
The brighter regions signify stronger perceptual shifts, suggesting areas where the visual content appears to “move” or change subtly as the observer’s perspective shifts. Interestingly, these fluctuations are not uniformly distributed: they concentrate around the subject’s facial contours, folds in garments, and regions where light and shadow blend most gently.
4.3. Depth Illusion Across the Flat Surface
Though the painting is entirely two-dimensional, the registered difference maps suggest that the visual experience of the painting changes depending on where the viewer is situated. The most dynamic perceptual changes occur in regions associated with depth construction, around facial features, drapery folds, and transitional shadows.
These findings support longstanding claims that classical painters, consciously or intuitively, embedded visual cues that activate spatial perception. In this case, the data supports the notion of phantom stereopsis, a type of monocular depth effect produced not by binocular disparity but by cues such as shading and occlusion. The viewer, when moving across the frontal plane of the painting, experiences these cues differently, creating the illusion that the image subtly reshapes itself.
4.4. Patterns in Tonal Variation
Histograms of intensity differences further confirmed that changes were localized rather than global. While overall brightness and contrast remained stable, indicating consistent lighting conditions during photography, specific regions showed measurable shifts in grayscale intensity. These tonal variations, though small in magnitude, were perceptually significant.
The most pronounced differences occurred between the second and third images in the sequence, likely due to the alignment of key facial shadows with viewer movement. Such subtle variations in visual rhythm may explain why some viewers report a heightened sense of depth or presence when viewing classical works like this one in person rather than in reproduction. The cumulative perceptual patterns across the series are presented in
Figure 5, where emergent depth-related activity becomes visible.
5. Discussion
The results of this study offer empirical support for a long-standing observation in art history and visual perception: that certain classical paintings appear to shift, deepen, or “breathe” as the viewer moves. What has often been described as an illusion of three-dimensionality, achieved through techniques like sfumato and chiaroscuro, is shown here to correlate with measurable, localized changes in pixel intensity across slightly different viewing angles.
Rather than treating the painting as a static image, our approach frames it as a dynamic perceptual object. The subtle tonal transitions embedded in the work invite continuous visual interpretation, especially when observed from multiple positions. These changes are not mere artifacts of lighting or camera perspective; they correspond to compositional strategies that direct the viewer’s gaze and simulate depth.
It is important to stress that our measurements are based on a digital camera and pixel-wise comparisons, not on direct recordings of human neural activity. The camera serves here as a stable, repeatable sensor that reveals where the image structure changes across viewpoints; it does not replicate the full complexity of visual processing, which also involves binocular integration, eye movements, memory, and attentional factors. For this reason we interpret the difference maps and Cumulative Perceptual Fields as indications of where the painting affords depth-related interpretations, rather than as a direct model of the underlying neurophysiology of depth perception.
Our findings are particularly relevant to the concept of Da Vinci stereopsis, wherein depth is inferred not through binocular cues but through the careful interplay of occlusion, shadow, and spatial ambiguity. While this effect has often been discussed theoretically or anecdotally, our difference maps and intensity analyses provide a framework for visualizing and quantifying it.
Importantly, the areas of most pronounced perceptual variation align with zones of soft shading, gentle transitions, and structural overlap, all core features of Leonardo’s visual vocabulary. The viewer’s shifting perspective seems to activate these pictorial zones, producing the illusion that the painted subject has presence, dimensionality, and even a form of responsiveness.
We note that the patterns reported here are sensitive to the spatial window used for analysis and to small variations in exposure across photographs. To assess this sensitivity, we performed control tests in which the cropped region was shifted by pixels (px), corresponding to a subtle lateral displacement within the facial area. Under these conditions, the Cumulative Perceptual Field (CPF) distributions changed in magnitude but preserved their spatial loci, indicating that the principal regions of perceptual activation remained stable even when the observational window was slightly perturbed.
A central limitation of the present study is that it focuses on a single painting. While this case was chosen for its suitability to our method, the results cannot by themselves establish that the reported phenomena are unique to Leonardesque or even Renaissance practice. A more systematic assessment would require applying the same procedure to a corpus of works spanning different periods and techniques, including paintings with sharper contours, more graphic modelling, or non-figurative compositions. We regard the present analysis as a proof of concept that motivates such future comparative studies.
Preliminary exploratory tests on a small number of works painted with harder contours and less pronounced tonal gradients suggest that the resulting difference maps exhibit weaker and more spatially fragmented CPF patterns. These informal observations are not sufficient for firm conclusions, but they support the hypothesis that the continuous modelling associated with sfumato may be especially effective at generating coherent, viewer-dependent depth signals.
Additionally, we applied histogram matching across views to harmonize global brightness and contrast levels prior to comparison. This procedure reduced broad tonal drift without altering the localized peaks of activity, suggesting that the observed phenomena are not artifacts of minor exposure fluctuations. Together, these checks indicate that the results are robust to modest acquisition variation while remaining grounded in the intrinsic tonal structure of the painting.
This observation resonates with art historical theories of “dynamic engagement,” which argue that paintings are not passive representations but perceptual experiences that unfold over time. Our use of digital tools demonstrates that such experiences can now be studied systematically, bridging the gap between subjective interpretation and objective analysis.
6. Conclusions
This study has demonstrated that classical painting techniques—particularly those associated with Leonardo da Vinci—can produce measurable, viewer-dependent illusions of depth. By capturing and analyzing a series of images taken from slightly different perspectives, we were able to identify how tonal transitions and shadow structures contribute to the perception of space within a flat image.
Using a registration-based approach and pixel-wise difference mapping, we visualized subtle perceptual shifts that align with long-held observations about Renaissance visual strategies. These shifts, concentrated in areas rich with sfumato and chiaroscuro, support the notion that such paintings contain latent spatial cues activated through movement and gaze.
More broadly, this research suggests that computational tools can serve as a valuable ally in the interpretation of historical artworks, not as replacements for traditional methods, but as complements that reveal what is otherwise difficult to observe. The intersection of visual perception, digital analysis, and art historical inquiry offers exciting possibilities for future scholarship.
At the same time, our conclusions are intentionally modest: they concern the behaviour of one carefully chosen work under controlled conditions and should be read as opening, rather than closing, questions about the generality of viewer-dependent depth illusions in painting. Extending the method to multiple works and to different artistic traditions will be essential to determine how widely these effects occur and how they vary with pictorial technique.
In an age when many encounters with art happen digitally, understanding how paintings behave across different viewing conditions may help preserve not only their visual content, but also the perceptual richness that gives them life.