1. Introduction
In industrial production, the operational state of critical machinery—such as turbines, generators, and precision manufacturing equipment—is intrinsically linked to subtle vibrational signatures. These micro-vibrations serve as vital indicators of equipment health, early warnings of potential faults (e.g., imbalance, misalignment, and bearing wear), and essential data for structural dynamics analysis [
1]. Similarly, in civil engineering, large-scale structures such as bridges, skyscrapers, and offshore platforms undergo continuous, often imperceptible, geometric deformations and vibrations under operational loads (e.g., traffic, wind, waves) and environmental influences [
2,
3]. Accurately measuring these micro-amplitude, potentially high-frequency vibrations is paramount for structural health monitoring, safety assessment, and predictive maintenance [
4].
Traditional non-contact vibration measurement techniques, including laser Doppler vibrometers (LDVs), fiber Bragg grating (FBG) sensors, and photoelectric position sensors [
5], offer high precision but often suffer from limitations such as point-based sensing (lacking spatial resolution), complex setups, high costs, and sensitivity to environmental interference. The advent and rapid progress in computer vision have catalyzed the emergence of vision-based vibration measurement as a highly flexible, non-contact, and full-field alternative. Originating from digital photogrammetry, this approach extracts structural motion by tracking image features (e.g., edges, corners, and artificial markers) or analyzing pixel intensity variations across video sequences [
6]. Similar comparative studies on image-based vibration extraction have highlighted the sensitivity of feature-based methods to noise and lighting variations [
7]. This method boasts significant advantages, including non-invasive measurement, high spatial resolution, applicability to complex geometries (such as beams, shells, and enclosures), and, crucially, the absence of mass-loading effects, making it ideal for delicate or micro-scale structures [
8].
Conventional vision-based vibration measurement encounters inherent limitations when applied to micro-amplitude vibrations. These limitations manifest primarily in three key areas: (1) Amplitude Sensitivity Constraint: The spatial resolution of standard area-scan cameras, typically limited by pixel sizes exceeding 1 μm, fundamentally restricts the reliable detection of sub-pixel or nanometer-scale displacements. Overcoming this limitation necessitates sophisticated super-resolution techniques, which are computationally intensive and susceptible to noise amplification. (2) Temporal Bandwidth Limit: The frame rate of area-scan cameras, typically constrained to hundreds or a low number of frames per second, imposes a significant bottleneck on the measurable frequency range. Consequently, capturing high-frequency vibrations requires specialized ultra-high-speed cameras, often entailing prohibitive costs. (3) Computational Burden and Artifact Sensitivity: Full-frame vision methods—including feature tracking [
9], optical flow [
10], and two-dimensional phase analysis—require the processing of large spatial neighborhoods, which results in substantial computational cost and reduced real-time feasibility. Tracking-based approaches are vulnerable to occlusion, drift, and noise-induced trajectory errors, often producing visual artifacts such as “ghosting” or amplified noise. Although Eulerian Video Magnification (EVM) [
11] and its phase-based derivatives [
12,
13,
14,
15] avoid explicit feature tracking, their reliance on two-dimensional spatial decompositions (e.g., Complex Steerable Pyramids [
16] and Riesz transforms [
17]) still imposes considerable computational overhead and may degrade robustness in noisy micro-vibration scenarios.
Line-scan cameras overcome the temporal bandwidth limitation inherent in area-scan systems [
18]. By capturing a single pixel line at rates exceeding 10 kHz, they enable the detection of high-frequency vibrations that are beyond the reach of conventional cameras, making them ideal for monitoring high-speed machinery or structural resonances. Processing this 1D signal stream also drastically reduces computational complexity compared to full-frame analysis, facilitating real-time processing.
However, resolving the high-frequency sampling issue exposes the second, more formidable algorithmic bottleneck: spatial sensitivity. In rigid structures, high-frequency vibrations invariably correspond to microscopic spatial displacements (often at the micrometer scale), rendering them entirely imperceptible to the naked eye and extremely challenging for standard sub-pixel edge detection algorithms. To visually reveal and quantitatively measure these subtle dynamics, Video Motion Magnification algorithms have been developed. Eulerian Video Magnification (EVM) linearly amplifies intensity variations over time; however, it tends to significantly amplify background noise and introduce blurring artifacts at high magnification factors. To suppress noise, Phase-Based Video Magnification (PVM) [
19] was proposed. Traditional PVM utilizes 2D Complex Steerable Pyramids (CSPs) to isolate and magnify the local spatial phase, successfully preserving structural edges.
While 2D PVM is highly effective for standard area-scan videos, it exhibits a computational and structural discrepancy when applied to line-scan camera data. Applying computationally heavy, multi-directional 2D spatial pyramids to the 1D spatial domain of line-scan sensors is mathematically redundant and undermines the inherent high-speed advantages of the line-scan architecture. The specific novelty of our approach lies in the mathematical reformulation of phase-based magnification from a 2D spatial pyramid to a 1D continuous wavelet framework. This transition enables direct local phase isolation along the 1D spatial axis, effectively reflecting physical acquisition characteristics of line-scan sensors.
To overcome these intertwined hardware and software limitations, this paper proposes a novel high-fidelity measurement framework that integrates line-scan imaging with a customized Complex Morlet Wavelet Phase-Based Video Magnification (CMW-PVM) algorithm. Employing the 1D continuous Complex Morlet Wavelet, the proposed method optimally isolates the instantaneous local phase directly along the 1D spatial axis. This streamlines the computational architecture while inheriting the superior anti-blurring characteristics of phase processing.
The primary contributions of this paper are summarized as follows:
A streamlined 1D CMW-PVM mathematical framework is established, specifically tailored for the high-speed spatiotemporal data architecture of line-scan sensors, eliminating the computational redundancy of traditional 2D pyramid methods.
Rigorous theoretical simulations demonstrate that the proposed method significantly extends the linear magnification range (up to ) while maintaining exceptional structural edge fidelity (FSIM ), even under extreme background noise conditions (SNR = 10 dB).
A comprehensive experimental validation was conducted using a laser Doppler vibrometer (LDV) as a reference, demonstrating that the system achieves high-precision kinematic reconstruction with a relative error of only 1.65% at 10 Hz. The method effectively resolves microscopic displacements (≈10 μm) at 100 Hz, effectively extending the sampling bandwidth of vision-based systems without requiring ultra-high-speed area-scan hardware. While the high sampling rate of line-scan imaging effectively mitigates aliasing for high-frequency vibrations, it should be clarified that the method adheres to fundamental sampling theory rather than violating the Nyquist–Shannon limit.
The remainder of this paper is organized as follows.
Section 2 reviews the related work concerning traditional EVM and PVM methods.
Section 3 details the mathematical derivation of the proposed CMW-PVM framework.
Section 4 presents the quantitative simulation analysis.
Section 5 discusses the experimental validations across multiple frequency bands, followed by the conclusions in
Section 6.
3. Proposed Methodology
As discussed in
Section 2, while the global Fourier transform diagonalizes translation perfectly, its fundamental limitation lies in its inability to isolate spatially varying local motions. Conversely, existing 2D multi-scale spatial pyramids successfully extract local phase but suffer from severe computational redundancy and topological mismatch when applied to the 1D high-speed spatial arrays generated by line-scan cameras.
To bridge this mathematical and computational gap, we propose a streamlined, 1D-optimized framework: Complex Morlet Wavelet Phase-Based Video Magnification (CMW-PVM). Employing the 1D Complex Morlet Wavelet (CMW), we decompose the structural profile into complex-valued basis functions that are optimally localized simultaneously in spatial position and scale (wavenumber). Crucially, each complex wavelet coefficient encodes a local amplitude and a local instantaneous phase within a finite spatial neighborhood determined by the scale . Analogous to the global Fourier phase, temporal variations in this highly localized phase correspond directly to localized sub-pixel mechanical displacements. The core innovation of CMW-PVM lies in applying temporal band-pass filtering and spatial amplification exclusively to this local phase signal directly along the 1D spatial axis. This enables robust magnification of spatial motion patterns while avoiding the computational burden of 2D directional filtering.
To bridge this mathematical and computational gap, we propose a streamlined, 1D-optimized framework: Complex Morlet Wavelet Phase-Based Video Magnification (CMW-PVM). Employing the 1D Complex Morlet Wavelet (CMW), we decompose the structural profile into complex-valued basis functions that are optimally localized simultaneously in spatial position and scale (wavenumber). Crucially, each complex wavelet coefficient encodes a local amplitude and a local instantaneous phase within a finite spatial neighborhood determined by the scale . Analogous to the global Fourier phase, temporal variations in this highly localized phase correspond directly to localized sub-pixel mechanical displacements. The core innovation of CMW-PVM lies in applying temporal band-pass filtering and spatial amplification exclusively to this local phase signal directly along the 1D spatial axis. This elegantly enables robust magnification of intricate spatial motion patterns while completely bypassing the computational burden of 2D directional filtering.
3.1. Complex Morlet Wavelet Transform
The Complex Morlet Wavelet Transform (CMWT) is a complex-valued, over-complete linear transform for 1D signals. It decomposes a signal into coefficients
corresponding to basis functions localized in position
b and spatial scale
. These basis functions are self-similar, generated through translations and dilations of a mother wavelet
[
22]. Specifically, the analytic Complex Morlet Wavelet mother function is defined in the time domain as a complex sinusoid modulated by a Gaussian envelope (see
Figure 1A):
where
is the scale parameter that controls the spatial width of the wavelet (inversely related to frequency). A larger corresponds to a broader wavelet and a lower frequency.
is the central frequency of the mother wavelet. For practical purposes, we set , which ensures that the wavelet is analytic (i.e., its Fourier transform is negligible for negative frequencies when ).
This wavelet is designed to separate the real (even, cosine-based) and imaginary (odd, sine-based) components:
As illustrated in
Figure 1B, a spatial translation of the mother wavelet basis corresponds directly to a linear progression in the local phase
, which establishes the fundamental mathematical link between phase variation and sub-pixel mechanical displacement.
The Gaussian window ensures spatial localization, while the complex sinusoid provides frequency localization within that window. The Continuous Wavelet Transform (CWT) of a 1D signal
at scale
and position
b, using the Complex Morlet Wavelet, is given by the inner product (convolution):
where
denotes the complex conjugate of the mother wavelet
. The coefficient
encodes the local amplitude
and the local phase
at position
b and scale
:
The local amplitude represents the concentration of energy of the signal component near b that oscillates near the frequency . The local phase captures the instantaneous phase offset of this oscillatory component, which is directly related to the local structure and position within the Gaussian window.
3.2. Complex Morlet Wavelet Phase-Based Video Magnification (CMW-PVM)
The core of our Phase-Based Video Magnification approach (CMW-PVM) is the amplification of local motion through phase. The amplification is achieved by adjusting the phase of the complex wavelet coefficients while maintaining the amplitude. We now formalize the CMW-PVM approach for the specific case of 1D translational motion of a diffuse object under constant illumination.
Figure 2 illustrates the proposed CMW-PVM architecture, with the main steps outlined as follows:
Wavelet decomposition: For each time
t, compute the CWT of each line
using the Complex Morlet Wavelet
on relevant scales
:
Continuous phase extraction: Extract the local phase map
from the complex coefficients:
Temporal phase difference: Compute the localized phase difference relative to a reference frame (e.g.,
):
Phase amplification: Amplify the phase difference by the factor
:
Phase factor reconstruction: Construct a modified complex coefficient that incorporates the amplified phase shift while preserving the original local amplitude:
This step effectively rotates the coefficient in the complex plane. Theoretically, for an analytic wavelet, a local spatial translation manifests as a linear phase shift . By augmenting the instantaneous phase to , the subsequent inverse transform maps this total phase shift back into the spatial domain, directly yielding the magnified profile .
Inverse Wavelet Transform: Reconstruct the magnified spatial signal
at each time
t using the Inverse Continuous Wavelet Transform (ICWT). For an overcomplete transform like the CMWT, this is achieved by summation over scales and positions using the real-valued synthesis wavelet or directly via
where
is a constant depending on the wavelet,
p is an exponent related to the reconstruction norm (often
), and
extracts the real part. In practice, discrete approximations and specific reconstruction filters are used to ensure numerical stability and reproducibility.
To ensure the reproducibility of the proposed CMW-PVM method (as shown in Algorithm 1), the specific implementation choices and computational environment are detailed as follows. The temporal band-pass filtering stage utilizes a 4th-order Butterworth filter to isolate the target motion frequency
. The cut-off frequencies are adaptively set as
Hz, providing a 40 Hz pass-band centered at the fundamental oscillation to suppress out-of-band sensor noise. For signal reconstruction, the Inverse Continuous Wavelet Transform (ICWT) incorporates the admissibility constant
specific to the Complex Morlet kernel to ensure energy conservation during the phase-to-intensity mapping.
| Algorithm 1 CMW-PVM pipeline implementation details | |
| Require: 1D line-scan sequence , magnification factor , wavelet scale , central frequency , band-pass frequency range | |
| Ensure: Magnified spatiotemporal signal | |
| // Step 1: Preprocessing | |
| for each time step do | |
| | ▹ Normalize intensity to |
| | ▹ Zero-padding for FFT efficiency |
| end for | |
| // Step 2: Complex Morlet Wavelet Decomposition | |
| for each time step t do | |
| | ▹ Using Equation (14) |
| | ▹ Extract local amplitude |
| | ▹ Extract local phase |
| end for | |
| // Step 3: Phase Manipulation and Noise Decoupling | |
| | ▹ Compute phase difference |
| | ▹ Butterworth or Ideal filter |
| | ▹ Apply magnification factor |
| // Step 4: Reconstruction | |
| for each time step t do | |
| | ▹ Modify phase only |
| | ▹ Inverse CWT reconstruction |
| end for | |
| return | |
4. Simulation Evaluation and Discussion
To assess the efficacy of the proposed Complex Morlet Wavelet phase video magnification (CMW-PVM) method, we theoretically examined the limitations of critical parameters to provide empirical guidance, subsequently developing simulations for comparison with traditional Eulerian video magnifier (EVM) and phase video magnifier (PVM) methods.
4.1. Parameters and Boundary Conditions
All motion magnification approaches possess fundamental boundary conditions beyond which amplification generates severe artifacts. For Eulerian Video Magnification (EVM) [
11], the underlying first-order Taylor approximation fails when spatial displacements
are large. Consequently, its amplification factor
is strictly bounded by the dominant spatial wavelength
:
Furthermore, EVM linearly amplifies intensity noise, severely limiting its practical magnification scale.
Similarly, Phase-Based Video Magnification (PVM) [
19] is governed by the phase wrapping limit to prevent aliasing artifacts (i.e.,
). Given the phase-displacement relationship
, its physical constraint translates to
. Both EVM and PVM reveal a critical vulnerability: higher spatial frequencies (smaller
) impose stricter magnification ceilings. Additionally, PVM’s reliance on global Fourier spatial homogeneity often causes phase distortions and linear noise amplification when processing localized complex motions.
To mitigate these global constraints, the proposed CMW-PVM establishes mathematically localized boundaries for its core parameters.
4.1.1. Limitation for Gain Coefficient
The gain coefficient
dictates the degree of amplification of motion. In CMW-PVM, excessive
will induce non-physical phase wrapping beyond the principal range
. Based on the local phase approximation
derived in
Section 3.2, the necessary condition to prevent phase wrapping is
Substituting the local approximation yields the precise theoretical upper bound for
:
Here,
(radians/pixel) denotes the maximum spatial frequency at which the wavelet exhibits the highest sensitivity to local structural changes, while
(pixels) is the maximum expected displacement. Unlike global methods, this formulation indicates that the applicable gain limit is adaptively determined by highly localized spatial textures and regional displacements.
A comprehensive comparison of theoretical constraints and noise sensitivities is summarized in
Table 1. Although EVM and PVM exhibit theoretical gain bounds dependent on global wavelengths, their practical effectiveness is severely capped by early noise-driven saturation. In contrast, CMW-PVM inherently incorporates spatial localization (
) and limits phase noise propagation. This results in an amplification factor that scales robustly within a predictable and physically interpretable saturation limit, significantly outperforming conventional global methods.
4.1.2. Time–Frequency Locality Constraint for the Wavelet Scale
The wavelet scale governs the fundamental trade-off between spatial localization and frequency resolution in the CMW-PVM method. An optimal must capture the target motion while simultaneously suppressing high-frequency noise and avoiding spatial blurring. This establishes a definitive operational interval.
The lower bound for
ensures that the wavelet avoids excessive sensitivity to high-frequency noise. Specifically, the effective carrier frequency of the wavelet
must not exceed the characteristic spatial frequency of the target movement
. This prevents the wavelet from predominantly responding to high-frequency interference:
According to Equation (
6), the central frequency
is typically set to 6. Therefore, the lower bound is strictly dictated by the frequency matching condition:
The CMW-PVM method estimates motion based on the locally stationary motion hypothesis, assuming the displacement field is approximately uniform within the wavelet’s spatial support region. If is excessively large, its spatial window will encompass areas with diverse motion patterns, leading to blurred motion and phase distortion.
To maintain local validity, the effective spatial support width of the wavelet (
) must not exceed the scale at which the motion experiences significant variation (
):
This yields a concise spatial locality constraint:
In general, the complete theoretical constraint interval for the scale parameter
is defined by two intrinsic spatial scales of the target motion.
This dual-scale boundary elegantly balances time–frequency resolution. It provides explicit theoretical guidance for practical parameter selection: the lower bound ensures spectral matching to suppress noise, while the upper bound preserves spatial locality to maintain the linearity of the phase-motion relationship.
4.2. Simulation Framework and Performance Evaluation Metrics
To ensure a rigorous and fair benchmark, we selected the 1D implementations of EVM and PVM as baselines. It is important to note that although these methods are widely known for 2D area-scan applications, their mathematical foundations—namely, the first-order Taylor expansion for EVM and the Fourier shift theorem for PVM—are fundamentally derived from 1D signal analysis. Therefore, the 1D-based versions of these algorithms are the most appropriate for this specific line-scan architecture. In contrast, modern learning-based methods [
13] are inherently designed for 2D spatial feature extraction, making them unsuitable for our 1D spatiotemporal signal stream.
Based on the theoretical constraint analysis of EVM, PVM, and CMW-PVM, the derived boundaries for amplification coefficients and phase continuity provide essential guidance for experimental design. To quantitatively assess the proposed method, we constructed a controlled 1D simulation model. This framework generates the ground truth at the pixel-level, enabling the precise quantification of magnification performance, robustness, and operational limits.
4.2.1. Signal Generation and Motion Modeling
The simulation utilized a Gaussian intensity profile to represent a smooth edge, providing a continuous gradient for sub-pixel estimation.
To simulate sub-pixel displacements
with high kinematic fidelity, we employed a high-order cubic spline interpolation scheme. Specifically, for each time step
t, a continuous cubic spline function
was constructed from the discrete reference signal. This function ensured second-order derivative continuity (
), which is critical for preventing artificial phase jumps during wavelet decomposition. The shifted signal was then generated by re-sampling the spline at the translated coordinates
. Compared to linear interpolation, this high-order approach minimizes intensity quantization errors and ensures that the simulated motion is perceived as a smooth, physically-consistent translation rather than a numerical artifact. The simulation and experimental processing pipelines were developed using Python 3.9.0. Key libraries include NumPy (v2.0.2) for high-performance complex-valued matrix operations and OpenCV-Python (v4.13.0.92) for spatiotemporal data handling. In the simulation benchmark (
Section 4), the magnification factor
was swept from 1 to 50 with a step of 1, and the input noise levels were varied between SNR = 10 dB and 30 dB to rigorously evaluate gain linearity and robustness boundaries.
To simulate real-world sensor limitations, the raw sequence was corrupted with additive white Gaussian noise. The resulting spatiotemporal map (
Figure 3C) exhibits characteristic intensity fluctuations used to test the robustness of phase-based extraction. To evaluate magnification accuracy, a ground-truth signal
was defined by applying the target amplified displacement
to the original profile (
Figure 3D).
This model provides a reference with precisely known values and a clear geometric meaning. To ensure statistical significance, we sweep multiple noise levels (e.g., varying input SNR) and report averaged results across multiple random seeds. For a fair comparison, all methods (EVM, PVM, and CMW-PVM) operate on the same noisy input and utilize the same temporal band-pass filter centered at . This ensures that performance variances arise solely from the spatial representation—specifically, the distinction between Eulerian intensity, global Fourier phase, and local wavelet phase.
4.2.2. Performance Evaluation Metrics
To comprehensively evaluate the algorithms, we employed two complementary families of metrics: motion magnification accuracy and video fidelity. These metrics quantify the precision of the motion extraction and the linearity of the amplification:
RMSE measures the root mean square error between the estimated amplified displacement
and the ground truth
.
The HDI is the ratio of energy in the higher harmonics to the fundamental frequency . This reflects spurious artifacts introduced when phase-continuity or linearity assumptions fail.
The PSNR evaluates reconstruction fidelity by calculating the ratio between the maximum possible power of the signal and the Mean Squared Error (MSE) between the amplified result and the ground truth
:
where
L is the dynamic range of the pixel values. A higher PSNR indicates superior pixel-level precision and reduced distortion.
In the context of motion magnification, excessive amplification often induces structural distortion or edge blurring (ringing artifacts). Standard metrics like the Peak Signal-to-Noise Ratio (PSNR) evaluate global pixel-wise errors but fail to align with the human visual system’s perception of structural fidelity. To rigorously quantify the structural preservation capacity of the proposed CMW-PVM algorithm, the Feature Similarity Index Measure (FSIM) was adopted as the primary evaluation metric.
The FSIM is explicitly designed to assess image quality based on two low-level features that are crucial for structural perception: phase congruence (PC) and gradient magnitude (GM) [
23]. The evaluation compares the reference (unmagnified or ground truth)
and the distorted (magnified) image
. Its mathematical model is as follows:
where
and
denote the phase congruency of
and
, respectively, and
. The gradient magnitudes of
and
are represented by
and
, respectively.
and
are positive constants, and
a and
b are two constants.
4.3. Performance Benchmark and Comparative Analysis
Drawing on aforementioned theoretical constraints and the experimental design, this section aims to evaluate and compare the performance of three methods: the linear Eulerian video magnifier (EVM), the phase-based video magnifier (PVM), and the proposed Complex Morlet Wavelet-based phase video magnifier (CMW-PVM). Using the quantitative metrics defined in
Section 4.2, the analysis focuses on the accuracy of the amplification, noise robustness, and distortion characteristics of each method under various simulation conditions. These evaluations serve to validate the superiority and applicable scope of the CMW-PVM approach.
The spatiotemporal diagrams (
Figure 4A–D) illustrate the motion reconstruction capabilities of each method.
Figure 4A represents the amplified signal of the ground truth. The EVM result (
Figure 4B) exhibits severe signal attenuation and discontinuity, a known limitation when the product of the amplification factor and displacement exceeds the spatial wavelength. PVM (
Figure 4C) successfully recovers the oscillation pattern but suffers from significant phase noise and background artifacts. In contrast, the CMW-PVM output (
Figure 4D) shows the highest correspondence to the ground truth, maintaining a high contrast-to-noise ratio and a smooth reconstruction of the wavefront.
The time-domain displacement analysis (
Figure 4E) further quantifies these observations. The EVM significantly underestimates the displacement amplitude due to its restricted linear approximation range. Although PVM tracks periodic motion more effectively, it introduces visible fluctuations at the oscillation peaks. Conversely, CMW-PVM yields a trajectory nearly indistinguishable from the ground truth. This precision is attributed to the localized multi-resolution properties of the Complex Morlet Wavelet, which ensures robust phase estimation under noisy conditions.
To assess spectral fidelity, the fast Fourier transform (FFT) spectra of the estimated displacements are compared (
Figure 4F). Although all methods identify the primary 100 Hz frequency, EVM exhibits a distinct secondary peak at the third harmonic (300 Hz). This harmonic leakage and waveform clipping indicate severe non-linear distortion, confirming that the motion has exceeded the Taylor series expansion constraints. Similarly, the overall magnitude reduction in both EVM and PVM indicates a breakdown of algorithmic saturation at
. In contrast, CMW-PVM yields a clean, high-magnitude peak at 100 Hz with negligible harmonics, successfully avoiding the non-linear artifacts that plague conventional methods.
These simulation results validate the theoretical boundaries established in
Section 4.1. Under typical operational conditions, CMW-PVM exhibits superior linearity, spectral fidelity, and noise suppression. To comprehensively delineate these advantages, parametric experiments were conducted across a broader spectrum of noise levels and magnification factors, utilizing quantitative image- and displacement-domain metrics.
4.4. Comprehensive Analysis of Algorithm Performance Boundaries and Robustness
This section comprehensively evaluates the performance boundaries of the algorithms across three critical dimensions: gain linearity, visual fidelity, and noise robustness.
4.4.1. Dynamic Magnification Range and Gain Linearity
In motion magnification tasks, the effective dynamic range is fundamentally determined by the algorithm’s capacity to preserve gain linearity—the fidelity with which the actual output magnification approximates the theoretical magnification factor . To evaluate this, we systematically investigated the saturation boundaries and amplitude errors of the three algorithms with a standardized noise floor (SNR = 20 dB).
As illustrated in
Figure 5A, EVM deviates from the theoretical baseline at a very early stage (
), corroborating its well-known sensitivity to noise and truncation errors with the Taylor first-order approximation. By shifting to the phase domain, PVM successfully extends the operational range. However, it plateaus near
. This indicates that under noisy conditions, the large, localized phase shifts induced by high magnification factors easily trigger global phase wrapping, causing the global Fourier transform to yield unstable or noisy phase estimations.
In contrast, CMW-PVM closely tracks the theoretical baseline in a significantly broader operational window, with saturation only emerging beyond . The spatiotemporal localization characteristic of the Complex Morlet Wavelet effectively isolates noise and prevents the global phase distortions that cripple PVM.
The corresponding amplitude error dynamics (
Figure 5B) further expose the underlying stability differences. As summarized in
Table 2, CMW-PVM maintains kinematic fidelity within an order of magnitude superior to traditional methods with high magnification regimes. While traditional Eulerian and global phase-based techniques are fundamentally constrained by early saturation and severe amplitude degradation, CMW-PVM keeps the amplitude percentage error strictly suppressed below 3% throughout the tested range, establishing its reliability for large-scale micro-motion extraction.
4.4.2. Fidelity Degradation and Motion Accuracy Analysis
Although increasing is mathematically feasible, it inevitably introduces background noise and structural artifacts. This section quantitatively compares visual fidelity and kinematic accuracy with extreme magnification regimes ().
Figure 6A,B illustrate the degradation trends of the PSNR and FSIM for
. As expected, EVM suffers the most severe degradation, with the lowest overall PSNR values and FSIM scores stagnating at an unacceptable baseline due to linear noise amplification.
PVM provides a noticeable improvement in the low range but deteriorates significantly as the magnification intensifies. Its FSIM curve drops sharply beyond . This structural collapse is driven by the violation of global phase continuity: heavily amplified localized motions cause phase wrapping in the global Fourier basis, physically manifesting as spatial ringing artifacts. In contrast, CMW-PVM strictly confines phase manipulation within localized spatiotemporal wavelet supports, effectively preventing noise propagation. Consequently, its FSIM remains stable within a high-quality threshold (0.87–0.91) even at , along with the most compact PSNR interquartile ranges.
The precision of kinetic measurement (
Figure 7A) reveals a distinct performance hierarchy. The rapid divergence of EVM’s RMSE confirms the breakdown of its linear approximation model. PVM exhibits a linear RMSE growth trajectory, indicating that global phase processing accumulates errors when handling large-scale motions. Conversely, CMW-PVM maintains a near-zero RMSE baseline (<0.1 mm) with extremely narrow 95% confidence intervals, demonstrating highly reproducible displacement estimates.
Spectral analysis of these errors (
Figure 7B) shows that both EVM and PVM exhibit significant harmonic leakage as
increases (e.g., surge energy
for EVM and increase total harmonic energy for PVM). In contrast, the harmonic energy of CMW-PVM remains negligible (below
dB). By avoiding global phase discontinuities, CMW-PVM confines the motion signal within the linear phase regime, ensuring that the amplified output is free from spurious frequency artifacts.
4.4.3. Noise Evolution Mechanisms and Robustness Analysis
In practical engineering applications, video inputs are inevitably contaminated by variable sensor noise. To evaluate algorithmic robustness, this section simulates continuous input SNRs from 30 dB (clean) to 10 dB (extreme noise) at fixed magnification levels ().
Figure 8A–C illustrate the evolution of displacement RMSE. As input SNR decreases, all methods exhibit an upward trend in error, exacerbated by larger magnification factors. EVM consistently presents the highest baseline error. PVM shows a moderate increase in error at lower magnifications, but its trajectory becomes critically steep at
as the SNR drops, indicating that the global Fourier phase is highly fragile and prone to catastrophic unwrapping failures under heavy noise conditions. However, CMW-PVM exhibits a remarkably flat growth rate. Even under the compound stress of
and
, its maximum RMSE is well below 0.2 mm. This suggests that the localized spatiotemporal support of the Complex Morlet Wavelet acts as a dynamic band-pass filter, preventing random noise from corrupting the kinematic trajectory.
The output PSNR heat maps (
Figure 8D–F) provide a holistic view of signal preservation. EVM (
Figure 8D) is dominated by cool colors with steep degradation gradients. Although PVM (
Figure 8E) improves globally by 5 to 7 dB, its vertical degradation at magnification levels remains significant. The CMW-PVM heat map (
Figure 8F) is distinguished by extensive warm regions, which achieve up to 30.1 dB under optimal conditions. More importantly, its horizontal and vertical degradation rates are remarkably gentle. By decoupling the motion phase from global noise via localized wavelet analysis, CMW-PVM mitigates the severe noise amplification effects that constrain traditional algorithms.
4.5. Summary of Simulation Results, Methodological Performance, and Limitations
Extensive controlled simulations quantitatively validated the theoretical advantages of the proposed CMW-PVM framework. Compared to traditional EVM and PVM, the proposed method demonstrates a significantly broader dynamic magnification range (maintaining strict linearity up to ), exceptional structural fidelity (FSIM ), sub-pixel kinematic accuracy, and robust noise immunity under severe dual-stress conditions.
The selection of baselines was strategically limited to the 1D analytical kernels of EVM and PVM. Recent 2D phase-based and learning-based techniques were excluded as they primarily prioritize visual magnification tasks, focusing on artifact suppression and perceptual realism. While these 2D methods can produce visually impressive magnified videos, they lack a deterministic and linear relationship between the input magnification factor and the reconstructed physical displacement. In contrast, our 1D CMW-PVM is designed for quantitative metrology, where the localized phase-to-motion mapping ensures that the magnified output maintains a strict, traceable correspondence to the actual vibration amplitude.
Despite its performance, several limitations of the 1D CMW-PVM must be acknowledged. First, the 1D architecture restricts motion acquisition to a single spatial axis, making it unsuitable for complex scenes requiring full 2D modal analysis or multi-directional non-rigid deformations. Second, concerning the sampling theorem, we clarify that the system resolves high-frequency dynamics not by “violating” Nyquist–Shannon limits but by leveraging the high physical sampling rate (exceeding 10 kHz) of line-scan sensors. This hardware advantage ensures that vibrations remain well within the sensor’s folding frequency, effectively avoiding the aliasing artifacts prevalent in conventional 2D vision systems.
Furthermore, while robust against Gaussian noise, the method’s performance under conditions of extreme illumination fluctuations or non-periodic transient impulses requires further evaluation to ensure industrial generalizability. We deployed the algorithm on actual high-speed line-scan video sequences to verify its metrological accuracy.
5. Experimental Verification and Discussion
5.1. Experimental Setup and Data Acquisition
To comprehensively evaluate the practical performance of the proposed CMW-PVM algorithm, a controlled laboratory experiment was conducted using a precision linear motion platform.
Instead of traditional contact-based sensors, a Polytec OFV-5000 high-performance single-point laser Doppler vibrometer (LDV) (Polytec GmbH, Waldbronn, Germany) was employed as the reference standard for ground-truth measurement. It boasts exceptionally high measurement resolution and a massive dynamic range, ensuring reliability without introducing any mass-loading effects to the moving structure. In this experiment, the motion extraction results processed by our vision-based algorithm were directly compared against the highly accurate reference signals acquired by the LDV.
Figure 9 depicts the experimental setup, comprising the schematic representation of the optical arrangement (
Figure 9a) and the tangible physical apparatus (
Figure 9b). The target plate was firmly attached to a motorized linear stage to produce regulated dynamic displacements. To ensure a thorough comparative analysis, the target’s vibration was concurrently recorded using both the LDV and the proposed vision-based sensor (line-scan camera).
The LDV optical head was strategically positioned to send a laser beam onto the target, thereby capturing its instantaneous displacement. Simultaneously, the line-scan camera, fitted with an industrial optical lens, was securely affixed to an optical vibration-isolation table. The camera was positioned roughly m from the target, with its optical axis oriented to guarantee that the target’s motion edge consistently remained within the 1D imaging sensor line. This dual-sensor non-contact system ensured that both devices saw the identical dynamic event without physical interference.
In the data collection procedure, the Polytec OFV-5000 LDV’s sampling frequency was established at 1000 Hz to obtain high-fidelity reference signals. Similarly, the line-scan camera continuously recorded the spatial profile at a line rate of 1000 Hz, achieving a spatial resolution of 2048 pixels per line. A specific high-contrast structural edge was employed on the target plate to enable precise optical-to-physical unit conversion (camera calibration). This 1000 Hz sampling rate provided a Nyquist folding frequency of 500 Hz, ensuring that vibration components up to several hundred Hertz could be captured without temporal aliasing. To establish the ground-truth displacement reference, the LDV velocity signal was numerically integrated over time and subsequently filtered using a high-pass filter to eliminate any low-frequency baseline drift.
5.2. Spatiotemporal Validation and Kinematic Accuracy
Because the line-scan camera acquired a 1D spatial array consecutively, the kinematic evolution was encoded in a spatiotemporal () image, where the horizontal and vertical axes represent time t and spatial position x, respectively.
Figure 10 visually compares the
cuts of the structural target edge before and after the application of the CMW-PVM algorithm. In the raw unmagnified slice
(
Figure 10a), the high-contrast edges of the target manifest themselves as perfectly straight horizontal bands. This geometric linearity visually confirms that the physical mechanical vibration of the platform is strictly confined within the sub-pixel regime, rendering the dynamic displacement completely imperceptible to standard optical observation.
In stark contrast, after being processed by the proposed CMW-PVM framework at a magnification factor of
, the latent micro-vibration is explicitly revealed. As depicted in the magnified
slice (
Figure 10b), the previously straight edges are transformed into highly visible continuous sinusoidal trajectories. More importantly, close inspection of the wavefronts (see the zoom-in insets) reveals that CMW-PVM faithfully preserves the structural sharpness of the spatial edges. Despite the significant spatial shift, no noticeable ringing artifacts or blurring degradation are introduced, which validates the theoretical findings of high structural fidelity (FSIM) simulated as described in
Section 4.4.2.
To quantify the kinematic fidelity of the system, the dynamic displacement trajectory was extracted from the magnified sequence using a sub-pixel edge detection algorithm (normalized by ) and compared with the LDV ground truth.
Figure 11A,B present the time-domain overlay and the corresponding fast Fourier transform (FFT) spectrum at a fundamental excitation frequency of 10 Hz. As depicted in the time-domain plot, the vision-extracted trajectory (red dashed line) exhibits high synchronization with the LDV ground truth (blue solid line). The vision system recorded a displacement amplitude of
mm compared to the LDV’s
mm, resulting in a relative error of
. This correlation confirms that CMW-PVM introduces minimal kinematic distortion during the 1D spatial magnification process for standard vibration frequencies.
Further spectral validation in
Figure 11B shows a dominant peak at 10 Hz. The overlapping spectral magnitudes indicate that the energy distribution of the vibration is faithfully preserved. In particular, consistent with the simulation analysis in
Section 4.4.2, a minor third-order harmonic component (
at 30 Hz) is observable in the vision-based spectrum. However, its magnitude remains over 20 dB below the fundamental frequency, indicating that the non-linear distortion is well-controlled.
A critical limitation of traditional vision-based modal analysis using area-scan cameras is the temporal aliasing induced by low frame rates (typically 30–60 fps), which restricts the measurable frequency bandwidth to below 15–30 Hz according to the Nyquist–Shannon sampling theorem. Furthermore, high-frequency vibrations in rigid structures often correspond to diminutive spatial displacements, presenting a dual challenge of temporal resolution and spatial sensitivity.
To demonstrate the bandwidth measurement capability of the proposed line-scan CMW-PVM framework, structural excitation was increased to 100 Hz. As illustrated in
Figure 11C, the physical vibration amplitude at 100 Hz drops significantly to approximately
mm. Despite the ten-fold increase in frequency and the reduction in the displacement scale, the proposed system successfully reconstructs the rapid sinusoidal oscillations. The vision-based sensor measured an amplitude of
mm compared to the LDV reference of
mm, maintaining respectable accuracy with a relative deviation of approximately
. This deviation is attributed to the decreased signal-to-noise ratio (SNR) at higher frequencies and smaller amplitudes, yet the waveform integrity remains intact.
The frequency spectrum in
Figure 11D correctly identifies the singular peak at 100 Hz without aliasing artifacts. It is important to clarify that this capability does not represent a violation of the Nyquist–Shannon sampling limit; rather, it is achieved by utilizing the high physical line rate of the 1D sensor (configured at 2 kHz in this experiment), which ensures the 100 Hz signal is oversampled and remains well within the folding frequency. This result proves that the system avoids the practical aliasing bottlenecks that plague conventional 30 fps area-scan sensors, effectively extending the operational bandwidth for vision-based metrology.
However, structural health monitoring in industrial settings often involves environmental uncertainties beyond high-frequency demands. While the CMW-PVM framework demonstrates high metrological fidelity under controlled laboratory lighting, its resilience to adverse conditions—specifically illumination fluctuations—requires dedicated evaluation. Since the CMW-PVM algorithm extracts motion from localized phase angles rather than raw pixel intensities, it is theoretically expected to maintain kinematic stability even when the image contrast degrades. To rigorously investigate the engineering robustness of this phase-based architecture, a quantitative stress test under intentionally degraded, low-contrast lighting is presented in the following section.
5.3. Robustness Under Adverse Lighting Conditions
To evaluate the engineering robustness of the CMW-PVM framework beyond controlled laboratory settings, a comparative stress test was conducted. We subjected the system to the same high-frequency configuration validated in
Section 5.2 (100 Hz excitation at 0.020 mm amplitude) but intentionally degraded the ambient illumination to simulate a low-contrast, shadow-occluded industrial scenario.
The theoretical resilience of the CMW-PVM algorithm fundamentally lies in its phase-based architecture. As established in
Section 3.1, structural vibration is exclusively encoded in the localized phase angle
, while sudden lighting drops primarily attenuate the signal’s intensity (amplitude
).
Figure 12A visually demonstrates this stability: despite the severe perceptual degradation and increased graininess in the raw
spatiotemporal slices under low-light conditions, the fundamental sinusoidal trajectory remains discernible for phase extraction.
Quantitative comparisons between the normal baseline (from
Section 5.2) and the low-light stress test are summarized in
Figure 12B,C. In the normal environment, the system achieved a measurement of 0.0182 mm (relative error of 9.1%). Under the degraded lighting, the measured amplitude slightly shifted to 0.0170 mm. While the increased background noise introduced minor low-frequency drift—raising the relative amplitude error to 15%—the absolute kinematic error remained tightly bounded at the micrometer scale (≈3 μm). Crucially, as shown in the spectral comparison (
Figure 12C), the dominant 100 Hz peak remained perfectly distinct and accurately aligned in both scenarios. This evidence proves that the CMW-PVM localized phase mechanism provides a robust, field-ready solution capable of maintaining metrological integrity in unconstrained, poorly illuminated environments.
Furthermore, while the CMW-PVM framework excels at magnifying periodic vibrations with high precision, its robustness in capturing non-periodic, transient impulses (e.g., impact loading or sudden structural shifts) requires further investigation. Future work will focus on adaptive filtering strategies and potential 1D-to-2D architectural extensions to ensure broader applicability in more diverse and unconstrained industrial environments.
6. Conclusions
This paper presents a 1D-optimized, analytical optical measurement framework designed to accurately capture high-frequency, micro-amplitude structural vibrations. By synergistically combining the ultra-high temporal resolution of line-scan imaging with the robust spatial localization of the proposed Complex Morlet Wavelet Phase-Based Video Magnification (CMW-PVM) algorithm, we address the inherent bandwidth and traceability limitations of conventional 2D vision systems. The primary contributions and findings of this study are summarized as follows.
1D-Analytical Metrology Framework: We proposed a novel CMW-PVM framework that establishes a deterministic, analytical mapping between the localized Complex Morlet Wavelet phase and physical structural displacement. By decoupling motion phase from global noise via spatiotemporal localization, the system ensures high kinematic linearity and metrological traceability, complemented by the rigorous derivation of operational boundaries () to ensure algorithmic stability.
Superior Robustness in Simulation: Rigorous parametric simulations revealed that CMW-PVM outperforms traditional 1D implementations of EVM and PVM. It extends the dynamic magnification range to and maintains a sub-pixel RMSE baseline with minimal harmonic distortion, demonstrating remarkable noise immunity at SNR dB.
Exceptional Kinematic Accuracy: Physical experiments benchmarked against a high-precision laser Doppler vibrometer (LDV) confirmed the algorithm’s metrological fidelity. At a baseline frequency of 10 Hz, the vision-extracted displacement achieved high synchronization with the ground truth of the LDV, yielding a relative error of only .
Broadband and Micro-Displacement Resolvability: The 100 Hz high-frequency experiment explicitly demonstrated the superiority of the line-scan architecture. By leveraging the high physical sampling rate of the 1D sensor, the system effectively avoids the practical aliasing artifacts that typically constrain 30–60 fps area-scan cameras, allowing for the reconstruction of microscopic displacements (≈20 μm) well within the physical Nyquist limits of the hardware.
In conclusion, the proposed line-scan CMW-PVM system provides a highly reliable, cost-effective, and broadband alternative to traditional point-based sensors for industrial structural health monitoring. Beyond vibration metrology, the 1D CMW-PVM framework also holds potential for broader sensing applications that rely on subtle temporal or structural variations. For instance, its noise-resilient localized phase features could support data-driven condition-monitoring pipelines used in air-handling units, where real operational datasets have recently become available [
24]. Moreover, the method may complement semi-labeled AHU fault-detection datasets [
25] by providing micro-dynamic descriptors that enrich feature representations for mechanical health assessment. In parallel, the rapid development of low-cost optical acquisition platforms [
26] suggests promising opportunities for deploying CMW-PVM on compact, field-ready hardware. While the current implementation is restricted to 1D spatial analysis and remains sensitive to extreme illumination fluctuations, addressing these limitations represents a clear direction for enhancement. Future work will focus on improving environmental robustness and deploying the algorithm onto embedded hardware (e.g., FPGA) to achieve real-time, edge-computing-based predictive maintenance for rotating machinery.