Next Article in Journal
Beating the Page Limit: Minimizing the Size of Documents
Next Article in Special Issue
Multitemporal Geodetic and TLS Survey of the Bridge ‘Ponte della Costituzione’ in Venice for High-Precision Deformation Monitoring
Previous Article in Journal
Experimental Study on Wind-Induced Vibration of Single-Axis Solar Tracker
Previous Article in Special Issue
Machine Learning-Based Classification of Vibration Patterns Under Multiple Excitation Scenarios for Structural Health Monitoring
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Vibration Measurement Data Enhancement Approach Based on Variational Autoencoders for Structural Health Monitoring

Dipartimento di Ingegneria dei Sistemi e delle Tecnologie Industriali, Università di Parma, Parco Area delle Scienze 181/a, 43124 Parma, Italy
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2026, 16(10), 4844; https://doi.org/10.3390/app16104844
Submission received: 11 April 2026 / Revised: 8 May 2026 / Accepted: 11 May 2026 / Published: 13 May 2026
(This article belongs to the Special Issue State-of-the-Art Structural Health Monitoring Application)

Abstract

Structural Health Monitoring (SHM) increasingly relies on data-driven approaches to detect structural changes under environmental and operational variability, yet the limited availability and imbalance of baseline data remain critical challenges. This study proposes a novel framework for vibration-based SHM that combines Convolutional Neural Networks and Variational Autoencoders to model structural response in the frequency domain through Cross-Spectral Matrices. The methodology includes a tailored data representation based on Cholesky factorisation, a CNN-VAE architecture with structural constraints to ensure data consistency, and an Enhanced Loss Function designed to improve sensitivity to modal characteristics. The trained model is used both as a generative tool to produce realistic synthetic data and as a feature extractor through latent variable distributions. Validation on an experimental truss structure subject to thermal variability shows that the model accurately reproduces the statistical distribution of natural frequencies and spectral features, while generating plausible synthetic responses. The proposed approach enables baseline enhancement through data balancing and supports effective damage detection using both modal features and latent space indicators. These results demonstrate that the framework can improve the robustness of vibration-based SHM systems and can be integrated with existing frequency domain monitoring techniques, offering a practical data-driven solution for real-world applications.

1. Introduction

A fundamental practice for ensuring the safety, reliability and longevity of buildings, civil infrastructures, aerospace structures or similar is Structural Health Monitoring (SHM), which can be defined as the process of implementing different strategies for novelty and damage detection through periodic measurements and analyses [1]. This is a multidisciplinary field that has evolved from qualitative visual inspections to modern SHM based on advanced automated data analysis and statistical pattern recognition. This was made possible through the development of sensor technology, wider interconnection between systems, and high-speed computing, amongst others. In recent years, other crucial players have been approaches based on Artificial Intelligence(AI), and consequently on Machine Learning and Deep Learning, that are fostering and enabling a strong transition from manual feature definition and assessment to automated and data-driven techniques [2,3,4,5]. It is worth mentioning the improvement of unsupervised learning and the diffusion of physics-informed AI applications, which implement the data-driven learning process while respecting physical principles [6,7]. An important branch of the wide field of SHM consists of vibration-based techniques [8,9,10], which are the subject of interest for this article.
The Fundamental Axioms of Structural Health Monitoring introduced in [11] define the general principles that should drive the realisation of any SHM technique and system. In particular, Axiom II affirms that damage assessment requires a comparison between two system states. In other words, any SHM approach requires the observation of a reference healthy state, the so-called baseline, that is compared with the current state. This comparison is realised by means of structural features extracted from measured data using, for instance, Operational Modal Analysis (OMA) [12,13,14] or other techniques [15]. However, practical implementation must face challenges related to environmental and operational variations (EOVs) and the scarcity of labelled data of damaged states. Typically, the structure under analysis is observed and characterised in a healthy state, under environmental and operational (EO) conditions, in order to understand the typical variations in the monitoring features [16]. However, the inability to control these external factors may lead to a biased baseline towards more frequently observed conditions. In some applications, it is possible to define and adopt monitoring features that are sensitive to damage but insensitive to EOV [17,18]. Otherwise, if a clear relationship between influencing variables and monitored features is identified, a compensation can be performed to reveal the residual signal cleaned up from these interferences [19,20,21,22]. In [23], the authors review different approaches to address EOV in SHM, identifying three categories of mitigation strategies: direct baseline compensation, adaptive and multi-baseline techniques, and reference-free methods, including those involving transfer learning and domain adaptation. In particular, the “data-driven virtual baseline synthesis” approach is mentioned as a method that mitigates the need for exhaustive references by reconstructing “virtual references for any given state” from a limited set of physical measurements. The key aspect of this method is the availability of a model that reliably captures EO dependencies in order to ensure diagnostic accuracy.
This article proposes a framework based on the integration of Variational Autoencoders (VAEs) [24] and Convolutional Neural Networks (CNNs) [25] for modelling dynamic structural responses in the frequency domain. The approach presented here involves a multi-purpose model that performs unsupervised feature learning based on measured structural response and, at the same time, can generate synthetic data to enhance various SHM techniques (e.g., data augmentation, virtual baselines [23]). A founding decision for this approach is to work with vibration structural responses expressed in the form of Cross-Spectral Matrices (CSMs) that lead to different advantages. In fact, the frequency domain representation acts as a preliminary feature extraction step, by reducing the dimensionality while preserving information related to structural dynamics, also according to [26]. Furthermore, it enables sample averaging to reduce the random fluctuations in the estimates of CSMs. Conversely, many other SHM approaches based on comparable models, such as Sparse Autoencoders (SAEs), operate in time domain [27,28] and do not exploit these advantages. The reason for adopting a CNN-VAE architecture is twofold. The first is that CNNs demonstrate high effectiveness in learning the vibration patterns [29,30], by capturing spectral continuity and the relationships between sensor spectra in a hierarchical manner. VAEs support this process of feature learning by enabling an advantageous probabilistic latent representation of input data [31,32]. The second reason is the potential of VAEs as powerful generative models for producing synthetic data. In fact, they learn the empirical joint distribution from experimental training data, making it possible to use the CNN-VAE model as a surrogate model in the field of SHM [33] and potentially for uncertainty quantification [34] based on the probabilistic framework of VAEs. In general, generative models are one of the emerging class of approaches for SHM [35].
This kind of empirical models are often fundamental for the implementation of SHM approaches since they are trained on a baseline dataset to capture the variability of the healthy structure response due to EOV. Then, they are fed with new monitoring data and measures based on reconstruction error, or changes in latent features are defined for novelty or anomaly detection. In [36,37], classic Autoencoders (AEs) are trained on frequency domain data of the healthy structure and the reconstruction error is used as a damage-sensitive feature. Similar approaches are proposed using SAEs [27,28,38], while other approaches use the monitoring of latent space of VAEs for condition monitoring [39] or SHM [40].
The key advantage of the approach presented in this paper compared to other approaches in the literature is the development of a single CNN-VAE model with dual functionality: feature learning and synthetic data generation. The original contributions of this article include:
  • A Cholesky factorisation [41] approach is adopted to encode CSMs into a format that brings multiple advantages in learning and generation.
  • The CNN-VAE model incorporates a tailored hard constraint to ensure the generation of positive (semi-)definite CSMs.
  • Development of custom Enhanced Loss Function (LF) for VAE training to promote the sensitivity of the CNN-VAE to modal features. This improves the accuracy of synthetic data and the use of latent distribution parameters as state/damage-sensitive indicators.
  • Use of CNN-VAE as an empirical surrogate model capable of generating realistic synthetic data samples. This capability can be integrated with many SHM techniques and is employed here for baseline balancing, mitigating statistical biases caused by non-uniform sampling of environmental conditions.
This paper demonstrates how this approach can be exploited to enhance long-term vibration monitoring through validation on a laboratory truss structure subject to environmental thermal variability and artificial damages. Anyway, these kinds of models are intended to be used for typical vibration-based SHM applications such as buildings [42], bridges [28], aerospace applications [5] and others. Moreover, the CNN-VAE model has the potential to be easily integrated with other SHM approaches, such as OMA, when used as a generative model for frequency domain data synthesis.
The paper is structured as follows. Section 2 describes the experimental setup of the laboratory structure and the measurement campaign with both healthy and artificially damaged structures. This section also briefly recalls the basic concepts of OMA. Section 3 introduces the proposed methodology step-by-step, from vibration frequency domain data organisation as model input to the CNN-VAE architecture and the training approach. This section also describes how to deal with the latent variables and the methods adopted to generate synthetic data. Section 4 shows the training of the CNN-VAE model specifically for the experimental benchmark structure and provides an evaluation of the performance achieved. Section 5 shows the results obtained using the same CNN-VAE model for two different SHM approaches. The first is used as a generative model for baseline balancing with respect to EOV combined with OMA. The second is the use as feature extractor applied to SHM, where latent variables are used to define a novelty index. Finally, Section 6 discusses the results obtained and Section 7 draws the main conclusions on the methodology presented.

2. Experimental Setup and Operational Modal Analysis

2.1. Benchmark Structure and Measurement Setup

The benchmark of the study is a simple truss structure installed in the laboratory at the University of Parma for SHM purposes (Figure 1) [43]. The structure is a classical two-dimensional Pratt truss girder with an overall span of 4.20 m and a height of approximately 0.80 m. Both bottom ends of the structure are linked to the ground supports with roller-type connections made by Teflon elements, tightened with a clamp system to allow for thermal dilatations. The diagonal and vertical members are aluminium rods with a hexagonal cross-section (12 mm across flats) and lengths of 500 or 700 mm, while the horizontal members consist of two steel H-beams, 4.30 m long and 1.1 cm thick. To improve structural stability and adjust the natural frequencies, seven rectangular masses of 15 kg each are placed at the bottom nodes of the truss. The total mass of the structure is about 300 kg. The entire structure is kept vertical by four struts located at one-third and two-thirds of the total span. The interaction between the struts and the upper beam is achieved through four stainless steel cylindrical elements which act as rollers, allowing for only vertical movement.
The truss structure is located in a laboratory environment where natural ambient excitation is not enough for vibration measurements and reliable modal parameter extraction with OMA. For this reason, an electrodynamic PCB Piezotronic model 2007E shaker (Depew, NY, USA) is mounted on the truss to excite the system in the vertical direction using a band-limited controlled white noise input signal [43]. The input control strategy ensures that the actual random excitation force has as flat a spectrum as possible. In fact, output-only OMA assumes a broadband excitation of the structure, ensuring that all structural modes in that frequency range are excited and can be detected [44]. Vibration output of the structure is monitored using six low-cost triaxial MEMS accelerometers (model TE4030) with a sensitivity of 1000 mV/g and a measurement range of ±2 g. Each sensor has a mass of about 50 g, which is negligible with respect to the total mass of the structure. The number of sensors has been deliberately kept low to replicate a realistic monitoring scenario. Environmental temperature is measured with four type K thermocouples placed in different positions, covering the whole geometry of the structure in order to highlight possible thermal gradients. All signals are generated and acquired through a National Instruments data acquisition system. The sampling rate is 2048 Hz. The modules’ built-in anti-aliasing filters prevent aliasing. To monitor and characterise the dynamic behaviour of the structure and the environmental conditions, a custom software based on LabVIEW 2020 manages the data acquisition system. The software automatically executes excitation and transducer measurements every 30 min. Each acquisition lasts 10 min. Raw time data from each test are stored in binary files, which are imported into MATLAB 2024b for the processing phase. The thermocouple signals are temporally and spatially averaged to monitor the environment temperature, thus giving a single general value for each test. For the structure response, Welch’s method is chosen to estimate auto-powers and cross-powers that form the CSM of the acceleration signals using the Hanning window, 50% of overlap and a block size of 10 s, thus having a spectral resolution of 0.1 Hz. For all the studies presented in this paper, the frequency range of analysis is 14–270 Hz where the modes of interests are located [18], corresponding to F = 2560 spectral lines, while only the accelerations in the vertical direction are considered, thus having N sig = 6 signals. Therefore, the structure response CSM is G C N sig × N sig × F , which denotes the whole 3D CSM. The 2D complex covariance matrix of size N sig × N sig for each frequency in the discrete spectrum f k is indicated by G ( f k ) .

2.2. Modal Parameters Extraction

A typical approach for SHM is to perform OMA, using only the vibration output of the structure [44,45], and monitor the modal parameters over time. In this work, the modal parameters extraction is based on a Frequency Domain Decomposition (FDD) algorithm [46,47,48]. For each temporal frequency f or angular frequency ω = 2 π f , this technique uses the structure output CSM G ( ω ) and its Singular Value Decomposition (SVD):
G ( ω ) = U ( ω ) S ( ω ) U H ( ω )
where
  • S ( ω ) is the matrix of singular values, i.e., a diagonal matrix with real and positive entries sorted in descending order s 1 ( ω ) s 2 ( ω ) s N ( ω ) ;
  • U ( ω ) is the matrix of singular vectors [ u 1 ( ω ) , u 2 ( ω ) , , u N ( ω ) ] that form a unitary matrix (i.e., U U H = I ).
In this case the SVD has this form due to the hermicity of the CSM, i.e., G ( ω ) = G H ( ω ) Both singular values and singular vectors are a function of the frequency. The first singular value s 1 ( ω ) reaches a local maximum and becomes dominant close to the natural frequency ω m of a structural mode. In such cases, the CSM G ( ω ) is approximately a rank one matrix, and the first singular vector u 1 ( ω m ) is a good estimate of the mode shape with unitary normalisation. Therefore, s 1 ( ω ) can be used as a Mode Indicator Function (MIF) to detect the modes and track the natural frequencies.
In practice, the spectral data are only known at discrete frequencies depending on the frequency resolution. The most straightforward approach to identify a structural mode is the Peak Picking approach, where the frequencies of the maxima of s 1 and the corresponding u 1 are taken respectively as modal frequency and mode shapes. However, the finite spectral resolution is an issue for SHM purposes, where a fine tracking of ω m is necessary. Therefore, in this work, natural frequency estimation is performed using the approach suggested in [49]. This is based on the approximation of the peak in the discrete singular value spectrum with the behaviour of a single degree of freedom (SDOF) system:
s 1 ( ω ) ω ω m 2 c m i ω λ m = 2 c m η m + i ( ω ω d m )
where λ m = η m + i ω d m is the pole of the SDOF system, ω d m is the damped frequency, η m is the real part of the pole and c m is a real coefficient. It follows that the natural frequency can be computed as ω m = η m 2 + ω d m 2 . The function in Equation (2) is fitted to the discrete singular value spectrum around a modal peak to estimate the modal frequency ω m by means of Least Squares Fitting. For conciseness, this approach is referred to in this paper as Peak Fitting (PF). It has been proven that this approach overcomes the limitations imposed by the spectral resolution [49]. The FDD has been adopted in this paper, together with PF, due to its effectiveness and simple implementation. It is worth clarifying that other frequency domain modal analysis techniques can be used in combination with the CNN-VAE model presented in this paper. The main requirement is that the technique uses a CSM as input, or similar complex matrices such as the matrix of Frequency Response Functions used in Experimental Modal Analysis [50].

2.3. Measurement Campaign

The first part of the measurement campaign was dedicated to the structure in a healthy condition and involved about three months of monitoring, from April to early July 2023, for approximately 3000 measurements taken during this period that serve as the baseline data. The main EO factor acting on the structure is the ambient temperature ranging from 18 to 30 °C, as shown in Figure 2. All CSMs from vibration response measurements are normalised to have the same overall level, as explained later on in the paper (Section 3.2, Equation (3)). The baseline measurements have been processed, as described in the previous sections, to estimate the natural frequencies by means of PF and the mode shapes to set up a standard SHM approach. The sample distribution for the spectrum of the maximum singular value s 1 ( f k ) is depicted in Figure 3. This is used as an MIF to estimate the natural frequencies f m using the PF technique described in Section 2.2. For SHM purposes, M = 5 modes are selected. Figure 4 shows the distribution of natural frequencies for the structure in a healthy condition. According to the MIF and natural frequency distributions, the monitoring bands B m reported in Table 1 have been identified. The most meaningful changes due to structural conditions and EOV are expected in these spectral bands.
The second part of the measurement campaign is performed in a damaged condition to test the SHM method. These measurements are carried out by creating three different types of artificial damage in different nodes:
  • Added mass: 1 kg mass is placed on the structure.
  • Bolt Torque: the torques of all bolts at a given node are changed with respect to the nominal torque of 12 Nm.
  • Bolt Removal: only a fraction of the bolts are left at a given node with respect to the complete junction with six bolts.
A complete overview of the combination of damage type and location on the structure is provided in Table 2. After the tests of each artificial damage type, the structure is restored to the healthy state to check its behaviour. These tests are named “Check A/B/C” in the rest of the paper. Figure 5 shows the temperature ranges for each experimental measurement condition, proving that the range observed during the baseline period covers all artificial damage and check conditions. Hence, no unobserved temperatures occurred during the various damage tests.

3. Methodology

The approach presented in this paper proposes modelling the frequency domain structural response, in terms of CSM, using a CNN-VAE architecture with a twofold objective:
  • To obtain a generative model capable of producing synthetic data consistent with the experimental observations;
  • To extract representative features from the measured spectral data.
In fact, CNNs’ capabilities are used to learn spectral features of healthy structure from baseline data, while VAEs are powerful generative models that produce realistic synthetic samples “similar” to training data. Essentially, the features extracted are sensitive to the most influencing factors in the reference conditions, accounting for structural dynamics and EOV, and can be used as “monitoring features”. In addition, these control the generation of synthetic structure responses that mimic the variability of real measurements for the healthy structure. The flowchart depicted in Figure 6 shows how to integrate a CNN-VAE model with vibration-based SHM. The model can be used to implement an SHM approach directly relying on the features learned from the baseline dataset. Otherwise, traditional SHM techniques based on frequency domain data can take advantage of possible baseline augmentation strategies, exploiting the CNN-VAE as a generative model.
This section of the paper is dedicated to the general description of the building blocks of the proposed approach. A short introduction of VAEs and their classic implementation is presented, since it is the same as the implementation adopted here. Then, the approach used to encode the CSM into a format more suitable for modelling and generation is then presented. The last sections deal with the custom Enhanced LF introduced for training, the structure chosen for the CNN-VAE model and, finally, the approach to encode input data in features and how to use them for synthetic data generation. Although the following sections present and discuss one aspect at a time, the choice of some parameters and ideas intersect with each other. In the remaining part of the paper, this methodology is applied to the case study presented in Section 2.1, but the approach developed and presented here can also be applied to analogous cases for modelling structure frequency domain output in terms of CSMs.

3.1. Variational Autoencoders (VAEs)

In the branch of deep learning, AEs [51] are a type of neural network that is capable of two core tasks: encode high-dimensional input data into a low-dimensional latent representation and decode it to produce a reconstruction of the input data. Therefore, an AE is a dual network made up of an Encoder, whose task is to learn an efficient representation of the input data, and a Decoder that is responsible for generating copies faithful to the original data using the latent representation. The deterministic latent representation of input data learnt by AEs can be used both for dimensionality reduction or feature extraction. Moreover, these models are widely used in SHM applications for anomaly/novelty detection by monitoring the latent space or the reconstruction error of new observed data [26,36,39,52].
A generalisation of the standard AE is the VAE [24,53]. This paradigm of deep neural network retain the same general architecture of standard AEs, but introduces a key difference: a probabilistic framework. In fact, VAEs treat the latent variables as random variables that are represented by probability distributions, rather than deterministic values. The probabilistic formulation yields two main advantages. On the one hand, the latent space is continuous and more suitable for meaningful interpolation. On the other hand, new samples similar to the input data can be generated by means of random sampling from the latent distribution. This makes the generation capabilities of VAEs far superior to those of standard AEs, making them one of the main types of generative models. Other models, such as SAEs [54], can learn efficient and more interpretable representation in latent space but they suffer the same AE issues as generative models. Conversely, Generative Adversarial Networks (GANs) [55] are powerful generative models, but they do not naturally extract features from input data. Hence, the final choice of model type falls on VAEs, essentially because they effectively combine dual capabilities in a unique model. In fact, VAEs learn a latent representation of input data and concurrently act as a generative model.
The typical implementation of VAEs is depicted in Figure 7, and it is the same as the implementation adopted in this work. The Encoder is a function that maps the input data X into a set of random latent variables Z, each with its own probability distribution. In this case, each latent variable is represented by a Gaussian Z N ( μ , σ 2 ) . According to the reparametrisation trick [24], the randomness of each latent variable is provided by an external random generator that samples values ε from a standard Gaussian, i.e., ε N ( 0 , 1 ) , while the actual Encoder outputs correspond to the parameters of each latent distribution, i.e., means μ and variances σ 2 . In this way, each input X is deterministically encoded into a vector of means μ and variances σ 2 for all the latent variables representing each particular input into a multivariate Gaussian in the latent space. Random samples for each latent variable are drawn by means of proper scaling and offset according to the probability distribution parameters provided by the Encoder, i.e., Z = μ + ε · σ . Thereafter, the values of latent variables Z are passed through the Decoder network, which produces the final VAE output Y, which is an approximation of the input X. Therefore, the reparametrisation trick decouples the randomness from the trainable Encoder/Decoder network parameters, thus allowing for the calculation of gradients for network training, making the encoding of X into probability distribution parameters a deterministic process.

3.2. From Spectral Data to VAE Input

The vibration output of the structure is available in the frequency domain using pre-processing, mentioned in Section 2.1, to estimate the 3D CSM G of measured signals for a discrete set of frequencies. In order to achieve effective data modelling with VAEs, the input data is organised relying on the following considerations:
1.
A proper CSM is a complex covariance matrix with some fundamental mathematical properties [41]. In fact, any CSM is a Hermitian matrix, must have a real positive diagonal and must be a definite positive matrix. Therefore, any model must generate synthetic data that fulfils the properties of a valid CSM. Moreover, the hermicity of G ( f ) implies that part of the matrix is redundant and provides no additional information.
2.
The structure response is dominated by the resonances in some bands, while the others have a much lower level, even three to five orders of magnitude below the resonance peaks (Figure 3). The main risk is that the lower zones of the spectra are treated as noise and not as a meaningful part of the data. Anyway, this may strongly penalise both the data reconstruction accuracy and the sensitivity of latent features to those bands if not handled properly.
3.
The structure response magnitude depends on the level of excitation, thus meaning that the actual level can vary between different measurements.
In order to address both these theoretical and practical aspects, the raw CSMs G estimated from measurements are pre-processed by means of a set of preliminary transformations for more robust and effective modelling with VAE.
The first operation, which addresses the last of the aspects mentioned in the list above, is the normalisation of the measured CSMs considering the whole spectrum. This operation is performed by multiplying each CSM by a factor in order to have the same overall Root-Mean-Square (RMS) for all measurements. The RMS value is computed using the amplitude of all the elements in the 3D CSM G i , j ( f k ) as follows:
RMS ( G ) = 1 N 2 F i , j , k | G i , j ( f k ) | 2
The uniformity of the overall level among all VAE input data makes the model “equally” sensitive to all measurements in the baseline and keeps as the most meaningful aspect the “shape” of the structure response over frequency and over sensors.
The next step is a cornerstone of the proposed approach since it enables an advantageous organisation of the CSM G ( f ) as VAE input and the control of the CSM properties generated by the model. Therefore, the first issue mentioned at the beginning of this section is resolved using the Cholesky decomposition of G ( f k ) , which is a form of matrix factorisation that represents the covariance matrices as follows:
G ( f k ) = L ( f k ) L ( f k ) H .
Therefore, the CSMs can be represented for each frequency as the product of a lower complex triangular matrix L ( f k ) , named the Cholesky factor, and its conjugate transpose. This is the ideal candidate for preparing the CSM as VAE input for multiple reasons:
  • The decomposition is unique and holds the same information of the original CSM;
  • The Cholesky factor is a lower triangular matrix with no redundancy in its entries;
  • The diagonal elements of the Cholesky factor are real and positive ( L i i > 0 ), if the original CSM is a positive definite matrix;
The first two properties are ideal to guarantee that the information remains the same but it is condensed in the lower triangular part of the matrix, since the triangular part above the diagonal is null. The third property makes it straightforward to impose the condition of having a positive definite CSM G ( f k ) , since it is sufficient to constrain all diagonal elements of the Cholesky factor to be real and positive ( L i i ( f k ) > 0 ). The application of this constraint will be discussed later on in this section in the definition of the Decoder network. Finally, the diagonal elements of the Cholesky factors are the conditioned standard deviations on all preceding variables. These represent the remaining standard deviation in a variable after the linear dependence on previous variables has been removed, by starting from L 11 ( f k ) = G 11 ( f k ) [41,56]. Therefore, as a further benefit of the Cholesky factorisation, the terms in L ( f k ) are proportional to standard deviations, i.e., proportional to the square root of variances in G ( f k ) ; hence, their dynamics is reduced analogously to the off-diagonal terms.
The next step is the reshaping of the 3D complex matrix into a 2D real matrix. Firstly, the lower triangular part of the Cholesky factor L is vectorised and the null entries of the upper part are ignored. Then, since there are still complex numbers because of off-diagonal terms, the vector is made real by splitting into adjacent elements the real and imaginary parts:
L v ( f k ) = L 11 ( L 21 ) ( L 21 ) L 22 ( L 21 ) ( L 21 ) L N sig N sig ( f k ) .
Therefore, a single frequency of spectral data G ( f k ) has been encoded in a real row vector L v ( f k ) of length C = N 2 . The matrix serving as the VAE input is obtained by stacking all of these vectors, forming a real two-dimensional matrix X R F × C as follows:
X = L v ( f 1 ) L v ( f 2 ) L v ( f k ) L v ( f F ) .
Each row of the data “image” X contains the response of the structure at a single discrete frequency f k and contains the relationship between the response of the structure measured by the different sensors represented in terms of elements of the Cholesky factor. Clearly, the order of the sensors, and consequently the order of the Cholesky elements, must be consistent throughout the measurements and the analyses.
After the organisation of the VAE input according to Equation (6), a final operation can be necessary depending on the characteristics of the experimental data (structure modes, sensor locations, etc.). In fact, as mentioned at the beginning of this section, the amplitude of the structure response spectrum is dominated by its resonances in some bands, while other regions of the spectrum exhibit a much lower level. This can lead to a model that learns faithfully only the peaks of the frequency response, while ignoring the other bands. This issue is partially mitigated by the Cholesky factorisation, but it may not be sufficient for an effective modelling. Therefore, a non-linear scaling function of the data, serving as compression of data dynamics, is proposed. The following function can be applied to each element of VAE input X , denoted by X k c , to scale original data:
x scaled = sign ( x orig ) x orig 1 / r
where sign ( · ) denotes a function that returns 1 for positive values and 1 for negative values. This expression keeps the sign of the original value x orig but uses the fractional exponent to rescale its amplitude depending on the parameter r. Conversely to a simple r th root, this can be fine-tuned and adapted to any dataset by varying the parameter r, in conjunction with the choice of the normalisation applied. Clearly, for integer odd values of r, it boils down to a regular r th root. Unlike a non-linear scaling based on logarithms, this circumvents the singularity for null values. The scaled values x scaled can be brought back to the original scale by applying the inverse function
x orig = sign ( x scaled ) x scaled r .
All the operations described in this section are used to prepare raw spectral data for the training of the VAE, i.e., the baseline, but also for new monitoring data provided to the model.

3.3. Training of VAE

The training of the VAEs is performed by maximising an objective function, i.e., the Evidence Lower Bound (ELBO) [53]. The optimisation is achieved by minimising an LF, defined as the negative of ELBO, that combines two terms:
  • Reconstruction Loss L r e c encourages the Decoder to accurately reconstruct the input from latent representation. This is typically implemented as the Sum of Squared Error (SSE) between the Encoder input X and the Decoder output Y , often referred to as “ L 2 -Loss”.
  • KL Loss L K L encourages the Encoder to map all the inputs into a distribution similar to a chosen prior distribution, favouring a compact and continuous latent space, and acting as regulariser for the model. This term is the Kullback–Leibler divergence between the actual latent distribution and the standard Gaussian N ( 0 , I ) , chosen here as prior.
The Reconstruction Loss L r e c is computed as
L rec = k F c C ( Y k c X k c ) 2 = X Y F 2
where the element-wise squared error between the Decoder output Y k c and Encoder input X k c is computed and then summed, thus resulting as equal to the Frobenius norm of the difference between output and input matrices. The second term of LF can be analytically expressed in terms of μ and σ 2 for each latent variable, given the choice of standard Gaussian N ( 0 , I ) as prior distribution; hence, the KL Loss L K L is
L K L = 1 2 i N lat 1 + log ( σ i 2 ) μ i 2 σ i 2 .
The typical LF for training a VAE can be expressed as
L VAE = L rec + β L KL
where β 0 is a hyper-parameter that allows us to control the balance between the accuracy of reconstruction and the regularisation of the latent space. As β increases, regularisation of the latent space becomes more important in training, but an excessive regularisation worsens the accuracy of reconstruction. This variant is referred to as “ β -VAE” [57,58], when β = 1 Equation (11) boils down to the classic VAE LF. The training is performed iterating on subsets of the entirety of the training data, typically referred to as “batches”. Therefore, each term of the LF is computed for each input in the batch and averaged over the batch size. An “epoch” is when the training over batches covers the whole set of training data. The reparametrisation trick enables the backpropagation, bypassing the stochastic part of the encoding–decoding process, and allows us to update network parameters and minimise L VAE .
The Reconstruction Loss of Equation (9) is the basic starting point to train a model that provides a general approximation of input data. In fact, the classic L 2 -Loss endorses the reduction in “big” errors. Considering the type of data used in this paper, some portions of the spectra are prominent with respect to others and are automatically privileged in the training process. Therefore, the resonant modes are naturally favoured in this process, even though they do not have the same amplitude. Conversely, as already mentioned in the previous section, this penalises the weakest region of the spectra. Different countermeasure contributes to mitigate this aspect such as the Cholesky decomposition and the non-linear scaling of Equation (7). Despite these countermeasures, modal peaks still have different heights and can be learned by the VAE with different accuracies, since no specific information about the natural frequencies and mode shapes are introduced in the LF. However, for SHM purposes, the modes are all equally important, independently of their magnitude in the structure response spectra.
In SHM contexts, it can be advantageous to extend the LF by incorporating additional terms to improve the learned representations [6,7,34]. The aim is to make them physically consistent, interpretable or simply more accurate on some particular characteristics. This approach adds a “soft” constraint to the training process that guides the model to learn not only a mere accurate reconstruction of the data. In fact, it can be exploited to focus on specific characteristics of the data but also to respect other requirements, without imposing a rigid constraint. The approach adopted in this paper is to exploit some typical concepts from OMA and FDD to “promote learning of important characteristics” of data useful for SHM, such as natural frequency and mode shapes, without ignoring the whole structure response.
As mentioned in Section 2.2, the spectrum of the biggest singular value of G ( f k ) is used as the MIF to detect modes and monitor the natural frequencies. For this reason, the LF is extended with two components that guarantee the fidelity of the MIF and mode shapes between experimental input data and VAE output. However, performing SVDs for each training iteration meaningfully increases the computational burden. Therefore, a fast estimation of the MIF is performed in the training phase, by exploiting two mathematical considerations. Firstly, the Frobenius norm of the Cholesky factor L has the following property:
L ( f k ) F 2 = trace S ( f k ) ,
where S ( f k ) is the singular value matrix from the SVD of G ( f k ) (Equation (1)). Secondly, the squared Frobenius norm L F 2 is a good estimate of s 1 ( f k ) in the proximity of the modes since the latter is dominant in those bands. According to VAE data organisation (Equations (5) and (6)), the approximate MIF can be computed in the training phase by
MIF ( X , f k ) = c = 1 C X k c 2
where MIF ( X , f k ) is a vector of length F; in fact, each element corresponds to a discrete frequency of spectral data. If no scaling is applied, Equation (13) is exactly equivalent to Equation (12); otherwise, the amplitude is squeezed by the non-linear scaling (Equation (7)). This information is useful to formulate the additional loss terms that improve the effectiveness of the VAE for SHM, allowing it to focus on the most important spectral features of structure response. This is implemented focusing on the monitoring bands B m (Table 1) and using a weighting strategy over frequency defined ad hoc to compensate different heights of the MIF for each mode peak:
w ( X , f k ) = 1 / max f j B m MIF ( X , f j ) if f k B m , 0 otherwise .
The weighting vector w ( X , f k ) basically normalises the MIF ( X , f k ) by making the maximum peaks equal to one in each band, while it cancels out the rest of the spectrum. Hence, the MIF Loss term L MIF is implemented as the weighted SSE between the input and output MIFs:
L MIF = k F w ( X , f k ) · ( MIF ( Y , f k ) MIF ( X , f k ) ) 2 .
Intuitively, this term makes the VAE training more sensitive to the location of the peak on the spectrum in the input data, rather than its amplitude. For this reason, L MIF guides the VAE to finely follow the peak position on the spectral axis.
A similar approach is adopted to endow the training process with specific sensitivity to modal shapes. In fact, as mentioned in Section 2.2, on the resonance the structure response is dominated by the mode shape. For this reason, the Shape Loss L SHP is defined as
L SHP = k F c C w ( X , f k ) · ( Y k c X k c ) 2 ,
where the weighting vector w ( X , f k ) balances the response amplitude for each monitoring band, thus making each mode shape equally important in the training.
The custom Enhanced LF used to train the VAE on CSM data is
L VAE = L rec + β KL L KL + β MIF L MIF + β SHP L SHP
where β MIF and β SHP are two additional non-negative hyperparameters that control the importance of these two additional custom terms in the LF. The terms L MIF and L SHP level out the sensitivity of the model to natural frequency shifts and variations in mode shapes, giving each mode the same “weight”. The benefits of these additional terms in the enhanced LF are shown and discussed in the rest of the paper. The CSM positivity on the Cholesky factor is not introduced in the modelling approach as a “soft” constraint since it does not prevent the occurrence of negative CSMs. Instead, a “hard” constraint is forced by means of the Decoder architecture and is discussed in the next section.

3.4. Encoder and Decoder of CNN-VAE

This section describes in detail the architecture chosen for the Encoder and Decoder based on a CNN. CNNs are a widely used class of deep learning models that allow for the hierarchical extraction of features with different levels of abstraction, while reducing the number of parameters compared to fully connected networks. The key factors that guided the choice of this architecture originate from the organisation of the structure response into a rectangular 2D real matrix, according to Equation (6), and the characteristics of spectral data. The “spatial” patterns on these pseudo-images can be effectively learnt by CNNs. In this specific case, the two dimensions represent the relationship between contiguous frequencies, along the row axis, and structural patterns, along the column axis. In fact, structural dynamic responses exhibit spectral continuity related to modal behaviour, while the structural response over sensors is encoded by the elements of Cholesky decomposition of the CSM G . The strength of CNNs when dealing with this kind of input data is their ability to learn and generate spectral and structural patterns while maintaining a relatively small number of trainable parameters. The hierarchical feature extraction provided by successive convolutional layers makes it possible to achieve a progressive aggregation of local spectral information into higher-level representations that describe dynamic behaviour of the monitored structure. Moreover, CNNs introduce a form of invariance with respect to translations along the frequency axis. In fact, they can still recognise and generate similar patterns even if they are shifted in the spectrum. This is useful because modal peaks vary and shift due to EOV. The schemes of both the Encoder and the Decoder are depicted in Figure 8.
The Encoder has a 2D input with a single channel that is followed by a series of convolutional blocks. Each of these blocks is formed by a 2D convolutional layer followed by a LeakyReLU activation function. The number of consecutive blocks of this type controls the receptive field of the CNN and the depth of the network. In this application, all these blocks share the same parameters to reduce the number of hyperparameters, which are number of filters N filt , kernel size K F × K C , stride ( S F , S C ) and LeakyReLU slope for negative values a. A stride S F > 1 can be adopted to progressively reduce the size of the frequency axis (rows) down to the last convolutional layer. The number of reduced frequencies as output from the last convolutional block is denoted by F r e d and depends on the number of convolutional layers and the stride itself. The combination of these two parameters controls the ability of the CNN to learn features from a larger frequency range. Thereafter, the convolutional feature map of size F r e d × C × N filt is flattened by a fully connected layer that returns the parameters of the Gaussian distribution, μ i and log σ i 2 , for each latent variable. In this implementation, the Encoder returns the logarithm of the variance log σ i 2 , instead of variance σ i 2 itself, since it is more convenient for numerical stability. Finally, the sampling layer puts in practice the reparametrisation trick and generates Gaussian random samples Z i N ( μ i , σ i 2 ) from each individual latent variable distribution.
The Decoder has a structure that mirrors that of the Encoder, but with some crucial differences. The input layer receives latent variable samples Z , then a fully connected layer coupled with a reshape layer mixes and projects them back into a space of dimension F r e d × C × N filt . At this point, the Decoder also has a series of convolutional blocks, each made up of a LeakyReLU activation function and a 2D transposed convolution layer, which progressively restores the resolution along the frequency axis, reversing the compression performed by the Encoder. In fact, these share the same parameters of convolutional blocks of the Encoder. The structural differences with the Encoder are found in the next layers. The space of size F × C × N filt coming out of the previous block is squeezed back into a space of dimension F × C × 1 by a “Summing Layer” created with a simple 2D convolution, having a single 1 × 1 kernel. This recreates a pseudo-image of the same size and shape of the input X . At this point, one of the key elements of this model comes into play. In fact, the “hard” positivity constraint on Cholesky factors takes place by means of a custom layer, named “Selective ReLU”. This layer simply applies a ReLU function only on those columns corresponding to the diagonal elements of the Cholesky factor, according to Equation (6). This impedes negative values on the Cholesky factor diagonals for each frequency and imposes a structural constraint of non-negativity on the Decoder output. Finally, a non-linear activation function (hyperbolic tangent) produces the Decoder output Y , keeping all values bounded in the range [ 1 , 1 ] for better stability.
The proposed architecture of the CNN-VAE is tailored to model this kind of spectral data. The anisotropic structure of input data and convolutions for the Encoder and Decoder give the user the possibility of tuning kernel sizes and strides independently for the frequency axis and the cross-sensor axis, according to the characteristics of specific data. For instance, the kernel size along the frequency axis can be chosen according to the frequency resolution and bandwidth of the spectral characteristics of the specific case under study. Similarly, the size of the kernel can be chosen to capture and represent the relationship between sensors. Moreover, the combination of Cholesky factorisation with Selective ReLU activation on its diagonal components removes the redundancy inherent to the hermicity of CSMs and, at the same time, imposes a structural hard constraint to generate Cholesky factors associated with (semi-)definite positive CSMs.

3.5. Latent Variables and Data Synthesis

This final section on the methodology used in this paper focuses on how to retrieve the latent representation of measured data using the Encoder and the approach adopted to generate synthetic data using the Decoder. According to the scheme in Figure 8, two kinds of outputs can be retrieved from the Encoder for a specific given input X ; one is deterministic and the other is random [59]. In the following, the subscripts D and R stand respectively for “deterministic” and “random” to specify which type of Encoder output is considered, while the Encoder function is denoted by the operator ENC ( · ) . The deterministic output consists of the parameters of the latent variable distributions as
P X = P μ , X P σ 2 , X = ENC D ( X ) ,
where P X is the vector of distribution parameters related to a specific input X . This vector is split in two parts, P μ and P σ 2 , respectively, the vectors of the means μ i and variances log σ i 2 :
P μ , X = μ 1 μ i μ N lat X P σ 2 , X = log σ 1 2 log σ i 2 log σ N lat 2 X .
Instead, the Encoder random output is the vector of samples Z for each latent variable, given the input X . This is indicated as
Z X = Z 1 Z i Z N lat X = ENC R ( X ) .
In this case, the samples are generated by feeding P X into the sampling layer of the Encoder that applies the reparametrisation trick. Therefore, the sample Z X represents a random value extracted from the latent distribution resulting from the mapping of the input X into the latent space.
An approach used in this paper is to mimic the usage of a deterministic standard AE, while keeping the advantages of a VAE in the latent space regularisation due to its probabilistic nature. In fact, it is useful to have deterministic mapping of a given input X into the latent space. This is directly achieved considering the distribution parameters returned by the Encoder P X (Equation (18)), instead of the random samples Z X (Equation (20)). Despite the differences with the deterministic latent variables of a standard AE, μ and log σ 2 are also parameters uniquely associated with a specific input X that can be used as monitoring features for SHM. This aspect will be discussed later on in the paper.
Synthetic data can be generated as the Decoder output for a generic latent vector Z as input. This is written as
Y = DEC ( Z ) ,
where the operator DEC ( · ) indicates the Decoder function. Depending on the content of the vector Z , different synthetic data can be generated. From the considerations above about the deterministic or random representation in the latent space, different approaches can be defined for synthetic data generation through the Decoder. Hence, a process named here as “Deterministic Generation” (DGEN) with the CNN-VAE is defined as
P X = P μ , X P σ 2 , X = ENC D ( X ) Y D , X = DEC ( P μ , X ) .
Basically, this process involves mapping a specific input X into the latent space, bypassing the sampling layer of the Encoder and feeding the Decoder directly with the deterministic latent vector of the mean values Z = P μ , X . In this way, the CNN-VAE generates a deterministic representation Y D , X of a desired input X that can be used, for instance, in model performance evaluation.
Conversely, one of the key advantages of the VAE framework is the possibility to randomly sample the latent space and produce new and unseen synthetic data. This process, named here as “Random Generation” (RGEN), with the CNN-VAE is defined as
Z X = ENC R ( X ) Y R , X = DEC ( Z X ) .
This approach enables the CNN-VAE to generate random variations Y R , X of a specific input X . This approach can be used for synthetic data generation useful for SHM purposes. The Decoder output, i.e., a synthetic Cholesky factor, can be restored to the original scale by means of Equation (8) and the synthetic CSM is obtained by means of Equation (4).

4. Spectral Data Modelling with CNN-VAE

4.1. Model Definition and Training

Using the methodology introduced in the previous section, a CNN-VAE model has been developed for the specific experimental case presented in Section 2. The architecture adopted for Encoder and Decoder networks was introduced in Section 3.4 and summarised in the scheme of Figure 8. The steps involved in defining the specific model are described below. It is worth noting that the various choices were validated by verifying afterwards that they were appropriate for the purpose of the model.
The first step was the definition of the number of consecutive convolutional blocks for both the Encoder and Decoder. In this case, the number was set to three, as this depth is considered sufficient for feature extraction. In fact, increasing the number of consecutive convolutional blocks also increases the receptive field of the CNN and enables the learning of more general features from the data. Conversely, a lower number of convolutional blocks would result in a model that is more sensitive to local features in the “pseudo-images”, rather than to global ones. Thereafter, the choice of a stride S F = 2 along the frequency axis allows for a progressive reduction in dimensionality and enables the extraction of patterns sensitive to a wider frequency range. After three convolutional layers, the frequency axis is reduced by a factor of 2 3 , thus resulting in a reduced frequency axis of length F r e d = 320 . Instead, the stride over column axis S C = 1 is adopted to keep its dimension over convolutional layers since it is already much smaller than the frequency axis. This combination of number of convolutional layers and stride is also advantageous because it reduces the number of learnable parameters with respect to an Encoder and Decoder with the same structure but with a single convolutional block. This happens since the size of fully connected layers present in both Encoder and Decoder are drastically reduced.
At this stage, a series of preliminary analyses was performed to set up the remaining hyperparameters. All the preliminary analyses were carried out using a subset of 300 measurements uniformly taken from the baseline (i.e., one tenth of the whole baseline, 50% for training and 50% for testing) and using the Classic LF (Equation (11)) without the custom terms. For the training on the small subset we used a batch size of 20, a number of epochs of 100 and a learning rate of 10 3 . The VAE networks were trained using the ADAM optimisation algorithm [60] to minimise the LF. The first analysis was conducted to set the parameter r for non-linear scaling of input data (Equations (7) and (8)), using a first tentative CNN-VAE. The minimum value that produced satisfying results was r = 3 . Thereafter, the remaining key parameters were studied: number of convolutional filters N filt , kernel sizes K F and K C , and the latent space dimension N lat . Two preliminary parametric analyses were conducted to understand the advantageous combinations that allow us to reduce the Reconstruction Loss (Equation (9)) and provide accurate reconstruction of structure response. Firstly, a rough grid of the four parameters was explored to identify the most favourable order of magnitude for each parameter. Secondly, once a plausible range was identified for each parameter, a Bayesian Optimiser was used to explore in detail the parameter subspace. The verdict of this parametric study is that a “high number of small convolutional filters” is adequate for a low Reconstruction Loss. This outcome observed on a smaller subset of data was confirmed on the whole training set used for the final model. The values of N filt = 48 , K F = 3 , and K C = 4 were chosen. From the same analysis, it emerged that a latent dimension N lat > 10 does not significantly improve the model performance in terms of Reconstruction Loss. In fact, by increasing the latent space dimension, the information is more diluted across the latent space variables, making them less informative. For this reason, N lat = 7 is chosen.
The definitive model was trained using the whole baseline, which was split in two equal groups; 50% of the data was used for training and the remaining 50% was used for testing. All the parameters used to define the model and for its training are summarised in Table A1, while the layers of the Encoder and Decoder are summarised in Table A2. Two models were trained: one using the Classic LF of Equation (11) and one using the Enhanced LF of Equation (17). To privilege the accuracy of the reconstruction, a value of β KL < 1 was chosen, while the values of β MIF and β SHP were tuned empirically, evaluating the performance of the trained CNN-VAE model in terms of accuracy and sensitivity to the modal behaviour of the structure. Figure 9 depicts a summary of the values for the total LF and each individual term, comparing the average on the training set with the average on the testing set. Both models show very similar results on both datasets, which may indicate that the models are not overfitting. The Reconstruction Loss L r e c assumes almost identical values for both models, thus meaning that the additional terms in the Enhanced LF do not degrade the model performance in terms of accuracy. The values set for β MIF and β SHP result in similar contributions of L KL , L MIF and L , SHP to the Enhanced LF (between 10 and 15%).
Finally, some notes about the development platform and computational burden. The models were realised with MATLAB 2024b using a desktop computer for training equipped with an Intel Core i9 processor 14900KF, RAM 64 GB, and GPU NVIDIA T1000 8 GB. Both the Classic and Enhanced LF CNN-VAE models required about 6 h to be trained using the specified parameters and training dataset. Once trained, the model can be run on less powerful hardware, such as a laptop without a GPU (Intel Core i7-1255U and 24 GB RAM). In this case, encoding a measured CSM in latent space takes approximately 0.11 s, also including Cholesky factorisation, reshape and scaling. The complete process of data generation from a measured CSM takes approximately 0.17 s (same for DGEN or RGEN).

4.2. Model Performance Evaluation

The performance of the CNN-VAE model is evaluated here before it is used for SHM purposes. The analysis presented here focuses on capabilities as a generative model for studying the ability to replicate baseline data and the level of accuracy. With this purpose, each measurement X in the baseline is replicated using the CNN-VAE model by means of Equation (22) ( Y D , X ), thus obtaining the DGEN synthetic baseline. Similarly, Equation (23) is used to generate 10 samples ( Y R , X ) from each experimental measurement in the baseline and have the RGEN synthetic baseline. The quantitative comparison between synthetic and measured data is realised in the original data scale, using Equation (8) to restore the amplitude.
The basic performance analysis regards the natural frequencies f m , estimated by means of PF introduced in Section 2.2, on experimental measurements and synthetic data. Figure 10 compares the result of measured data with DGEN synthetic data, also showing the difference between the CNN-VAE trained with the Classic LF and the Enhanced LF. The model trained with the Classic LF represents the most frequent behaviour but ignores rare “peripheral” behaviour. Instead, the custom terms in the Enhanced LF of Equation (17) improve the capacity of the model to better reproduce less frequent conditions. This behaviour is more evident in Figure 11, where the natural frequencies estimated on the RGEN synthetic baseline are shown. Here, the advantage of the Enhanced LF is further highlighted, especially for rare conditions. A quantitative comparison is provided in Table 3 that reports the Root-Mean-Squared-Error between baseline measurements and synthetic data of natural frequencies f m estimates.
The next analysis is extended to the performance on the full spectrum in terms of the MIF computed as in Equation (13). Again, both synthetic baseline datasets DGEN and RGEN are compared with the experimental structure response in the baseline period. Figure 12 depicts the distribution of the Absolute Relative Error over frequency for the MIF spectrum. These results shows that the error is typically an order of magnitude lower than the MIF spectrum amplitude of measured data. As further quantitative evidence, the average value of Absolute Relative Error computed over frequency is 0.15 for DGEN data and 0.16 for RGEN data. Therefore, this means that not only is the DGEN structure response an accurate representation of measured data, but the RGEN data are plausible “variations” of structure spectral response observed in the baseline period. In this sense, the accuracy of the CNN-VAE model can be considered adequate for synthetic data generation. The usage of this tool for SHM purposes is discussed in the next section.

5. Spectral Data Enhancement with CNN-VAE Model for SHM

The previous section demonstrated that the CNN-VAE model can replicate the measured vibration response of the structure with sufficient accuracy, when properly trained on a dataset of reference experimental CSMs. This section presents two different uses of the CNN-VAE model that enhance the information extracted from the baseline dataset. The approaches described here are rooted in the main characteristics of VAEs. The generative ability to produce synthetic unseen data with characteristics similar to those of experimental ones is a key point. The other is the ability of feature learning from raw data. Therefore, vibration measurements can be enhanced for SHM purposes, exploiting the potential of the CNN-VAE model.

5.1. Generative Model: Baseline Balancing

The first approach uses the generative capability of the CNN-VAE model. Some applications may require the generation of new data with the same characteristics of the existing measured data to obtain some advantages in further processing. The idea is not to inject artificial, new information in the baseline, but rather to optimise the information that is already available. An example of this is described in the following. The structure state is typically affected by EOV, during both the reference and the monitoring periods, that may potentially give rise different issues. In fact, EO factors induce variations in monitored features that may affect a novelty or damage index. A possible issue may arise from the baseline data considering that EOV cannot be controlled during the measurements. In fact, this may produce an imbalanced baseline with respect to some influencing factors, since some conditions are more common than others. This can cause a bias in some SHM approaches towards the most frequent states.
In the experimental case presented in Section 2, the main environmental factor is the temperature, ranging from 18 to 30 °C during the baseline measurements. However, not all temperatures are evenly observed, as depicted in the histogram of Figure 2. Therefore, the baseline is dominated by temperatures in certain bins, whilst others, although recorded, occur much less frequently in the measured data. The idea proposed here is to produce synthetic data with the CNN-VAE model to balance the baseline with respect to the known influencing factors. The effect is a reduction of the bias towards frequent states due to non-uniform sampling of EO conditions. Clearly, this method can be integrated with any SHM approach based on frequency domain structure response, since the CNN-VAE model directly produces spectral data in the form of CSMs.
An example of this concept is demonstrated here using a typical SHM technique that is based on the mode natural frequencies f m , considering the modes of interest in the monitoring bands B m (Table 1). For each measurement, a vector f R 1 × M is filled with the natural frequencies estimated by means of PF. The measurements of each structure state are compared with those observed in the reference period, hence the baseline. The distance between the current state and the reference can be quantified by means of Mahalanobis Distance (MD):
MD = f μ f C f 1 f μ f T
where μ f R 1 × M and C f R M × M are the sample mean vector and the sample covariance matrix respectively, estimated on the natural frequencies observed in the baseline. This is a measure of distance for multivariate data that takes into account the possible correlation between variables and is widely used as a novelty or damage index [61].
This SHM technique is applied to the structure case of Section 2, considering all artificial damage cases reported in Table 2. The temperatures for each test condition are summarised in Figure 5. Two datasets are used. The MD is calculated as in Equation (24) for two different baselines: the experimental baseline and the balanced baseline. The natural frequencies forming the experimental baseline are depicted in Figure 4, while the balanced baseline is obtained by merging the experimental one with RGEN data properly generated to balance the temperature occurrences. This is achieved by considering the histogram in terms of counts for each bin (Figure 13). The most frequent condition is retrieved from the histogram computed on the experimental baseline, by finding the bin with the maximum height. Data synthesis with CNN-VAE is performed for each bin with a height lower than the maximum, using the following procedure:
1.
Randomly select an experimental measurement among those with their temperature in the current bin;
2.
Generate new synthetic data using the RGEN approach by means of Equation (23);
3.
Stop the process if the height of the current bin reaches the same height as the maximum bin, or if each measurement in the bin has already been replicated 10 times. Otherwise, continue the process from the step 1.
Once this procedure is complete for each bin to be increased, the RGEN synthetic data are merged with the experimental baseline to constitute the balanced baseline. In this application, about 3700 synthetic data are generated to obtain the balance. As shown in Figure 13, the difference in relative frequencies in the dataset considered as the baseline is eliminated or drastically reduced. In fact, the average bin height changes from 37% to 88% of the maximum bin height before and after the balancing, thus meaning that the temperature histogram tends to become flat (a value of 100% would indicate a perfectly flat histogram). The limit on the number of replications has been implemented to prevent measurement in bins with few occurrences from being replicated excessively. The effect of synthetic data on the new baseline of natural frequencies is an adjustment of mean and variance that compensates for the unavoidable unbalanced sampling of EOV in the experimental campaign.
At this point, the MD can be calculated with respect to both baselines, and Figure 14 summarises the results. For each condition, three sample statics of MD values are calculated: mean value, 5th percentile and 95th percentile. To ensure that results obtained from different baselines can be compared easily, all the MD values with respect to the same baseline are normalised relative to a specific threshold. In this paper, the threshold is the 95th percentile of MD values of the baseline with respect to itself. This is regarded here as the threshold for estimating whether the new structural states are novel or damaged. In this way, normalised values lower than 1 can be considered healthy, regardless of the absolute value of the MDs. The results in Figure 14a show that the MD already produces realistic indications on the experimental baseline and the state index does not deteriorate on the balanced baseline, thus meaning that the generated data are consistent with the original data. In general, 1 kg of added mass on the structure, with an initial mass of 300 kg, induces small variations on the natural frequencies. These are not always evident, depending on the location of the mass. Similarly, reducing the bolt torque in a node produces increasing effects as the tightening torque drops. Finally, the most evident changes are experienced in the case of bolt removal in a node. In this case, the dynamics of the structure is heavily affected. Anyway, a generalised advantage in damage detection is highlighted, since the statistics of MD become higher when computed with respect to the balanced baseline, while the healthy states (Check A/B/C) remain compatible with the baseline range although slightly altered. This is confirmed by Figure 14b, which shows that the average MD value increases more significantly in cases where the damage is more evident than in cases where the structure is considered healthy (e.g., bolt torque above the nominal one). This improvement in novelty/damage detection performance is paired with the advantage of knowing that there is no evident bias towards a specific EO condition, which is the main objective of the baseline balancing.
An important advantage of data synthesis with this model is that it can be easily integrated with many SHM approaches working with frequency domain structural response. For instance, the approach adopted in [18,62] involves the definition of an index that is insensitive to EOV. In fact, the natural frequencies f m of the baseline are pre-processed using Principal Component Analysis to detect and remove the component that is correlated with the temperature. Then, the MD is computed using only the components independent of the temperature, or, in general, from the known EO factors. Although this method has been shown to be effective in reducing the impact of temperature on the monitoring index, there may be some drawbacks. For example, the relationship between temperature and natural frequencies may be sampled non-uniformly. Therefore, in such cases, combining this with baseline balancing can improve the results. Otherwise, another possible drawback is that the removal of components correlated with EO factors may also delete information related to damage. In this case, the baseline balancing can be a good alternative to avoid this issue.

5.2. Feature Extraction: Latent Space Monitoring

Another possibility offered by the CNN-VAE model is the encoding of experimental spectral data X into the latent space by means of Equation (18). Therefore, the focus here is on the latent space behaviour. The parameters of the latent variable distributions P X can be used to monitor the structure state, since they are directly dependent on its spectral response. Again, the MD is used to quantify the distance between the reference dataset and the current one [63]:
MD = P X μ P C P 1 P X μ P T
where μ P R 1 × 2 N lat and C P R 2 N lat × 2 N lat are the sample mean vector and the sample covariance matrix respectively of the latent variable distribution parameters. The reference dataset adopted here is the experimental baseline.
A comparison is performed between the CNN-VAE models trained with the two different LFs, Classic and Enhanced, in order to highlight their effect on the latent space. Figure 15 summarises the results in terms of MDs (Equation (25)), using the same strategy for normalisation already adopted in Section 5.1. The model trained with the Classic LF recognises all new data as anomalous with respect to the baseline, independently from the actual structure conditions. Instead, when the model is trained with the Enhanced LF, the latent variables become more sensitive to the structural state, acting much more similarly to the natural frequencies used in common SHM techniques (Figure 14). A key difference is that the latent features are still sensitive to other changes in the structure response, thus making them more general. The drawback can also be that meaningless changes in the response may be detected. Anyway, this result further demonstrates the effectiveness of the additional terms in the Enhanced LF, in addition to the better accuracy already highlighted in Section 4.

6. Discussion

Vibration measurements are one of the fundamentals of SHM for the assessment of structure condition. In many applications, only data from the healthy state is available, i.e., the baseline that is typically subject to EOV. In such cases, a model of the structural “normal” behaviour is useful for gathering as much information as possible and gaining the maximum benefit from the available data. For this reason, this paper has introduced a CNN-VAE model that learns the joint distribution of high-dimensional structural response data. This empirical model is designed to deal with frequency domain data, in the form of CSMs, and is trained using baseline measurements of structural response to learn low-dimensional state/damage-sensitive features and generate synthetic realistic data. The methodology introduced has been applied to the vibration response measurements taken from the experimental laboratory test case that is introduced in Section 2. An original approach, relying on Cholesky factor, has been introduced in this paper to organise a 3D CSM for CNN-VAE, which brings multiple advantages to the whole method:
  • Redundancy of the CSM due to its hermicity is removed without loss of information;
  • Dynamics of the spectra is reduced, favouring the reconstruction of the regions with lower amplitude;
  • Possibility to enforce a simple constraint for the positivity of CSMs generated by the model.
For this reason, the Decoder has been designed to embed an “hard” constraint that leads to the generation of Cholesky factors related to positive definite CSMs, which is another key aspect of the methodology presented. This guarantees the fundamental mathematical properties of synthetic CSMs and ensures that they are compatible with other techniques relying on them. Lastly, a custom Enhanced LF tailored to modal behaviour of structures has been developed to train the CNN-VAE model. Its implementation only requires the identification from the baseline dataset of the bands of analysis where the mode natural frequencies are located. The results demonstrated that this strategy improves the quality of synthetic data (Section 4.2) and that the latent variables become highly sensitive to structural dynamic changes and representative of the structural condition (Section 5.2).
The capability of CNN-VAE as generative model opens the possibility to different enhancements of vibration spectral data in the context of SHM. An example is the baseline balancing shown in Section 5.1, but any similar data augmentation approach can be easily integrated with the CNN-VAE model, such as the idea of “data-driven virtual baseline synthesis” suggested in [23]. One strength of the approach presented in this paper lies in the fact that traditional SHM methods based on frequency domain structural response can benefit from data synthesis. At the same time, the low-dimensional latent representation learnt by the Encoder enables other SHM approaches. Typical anomaly detection techniques can be implemented when only data related to an undamaged structure is initially available, such as One Class Support Vector Machine (OC-SVM) [64]. Moreover, monitoring data in the latent space can be integrated with classification techniques, for damage identification, or clustering techniques. Otherwise, more advanced anomaly detection strategies relying on the spectral reconstruction error [65] or custom novelty scores [59,66] are possible.
Practical implementations of an SHM strategy that makes use of this CNN-VAE model should take into account the computational costs and time requirements of different operations, on the basis of the estimates reported in Section 4. The heavy load is the initial training of the model on baseline data, but this is operated only once or when the baseline is updated. The generation of synthetic data for the purpose of implementing data augmentation and enhancement techniques combined with SHM approaches is significantly less computational resource-intensive. Typically, this is not a real-time operation, but the latent space monitoring can also be implemented in near-real-time. Indeed, encoding a CSM into latent space takes a fraction of a second. Therefore, the time to obtain the monitoring latent features is dominated by the measurement time required and the computation time required to estimate the CSM.
Other potential applications involve combining insights into structural behaviour derived from both simulation and experimental measurements. In fact, simulations can provide information about the structural behaviour under various configurations and conditions by means of parametric analyses, while experimental tests can reveal the gap between the complexity of a real structure that is not adequately represented by a finite element model. This typically happens with complex structures due to complex geometries and materials (e.g., anisotropy and non-linearity), as it is common for aerospace applications [67] and others. In such cases, the CNN-VAE model could be trained using frequency domain data from both worlds and merging the knowledge in a unique surrogate model.

7. Conclusions

The study successfully demonstrated the effectiveness of the CNN-VAE architecture designed specifically for structural vibration data in the frequency domain with the purpose of SHM technique enhancement. The use of Cholesky factorisation for handling CSMs has proved to be an advantageous choice, as well as the introduction of the Enhanced LF for model training, revealed to be crucial in improving the sensitivity to modal characteristics. The results obtained in the laboratory case study demonstrated the effectiveness of the CNN-VAE model in its dual use. The features learnt from spectral data can be used to implement an effective SHM strategy. Concurrently, the CNN-VAE serves as a powerful generative tool for producing realistic synthetic data, thus making it possible for it to be integrated into traditional SHM approaches based on frequency domain data by means of baseline augmentation strategies or similar.
The proposed framework offers a practical solution for improving vibration-based SHM systems, since the model does not have particularly limiting requirements in terms of number of vibration sensors and structural characteristics; hence, it is suitable for different type of structures, such as buildings, bridges, and aerospace structures, among others. This CNN-VAE model is already compliant with measuring 3D vibration response, since these are simply additional sensors; therefore, this extension is straightforward. Moreover, the very same architecture of the CNN-VAE is also compatible with other types of frequency domain data, such as frequency response functions or transmissibilities.
Future works will focus on the extension of this model in terms of its applications and architecture. A finer characterisation of damage evolution over monitoring could be useful to understand the fine sensitivity of the model to gradual changes in structure dynamics. A possible direction of improvement for the current model is the shift to the paradigm of Conditioned VAE, where measurements of EO factors are provided to the Decoder, which becomes aware of the conditions of that specific input, thus making the synthesis more controllable.

Author Contributions

Conceptualisation, S.P., G.B. and M.V.; methodology, S.P., G.B. and M.V.; software, S.P. and G.B.; validation, S.P., G.B. and M.V.; writing—review and editing, G.B. and M.V.; supervision, M.V.; funding acquisition, G.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research was granted by University of Parma through the action Bando di Ateneo 2024 per la ricerca.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The dataset used in this study is not publicly available.

Acknowledgments

The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Appendix A

Table A1. Constants and parameters of input data normalisation and CNN-VAE (all values are non-dimensional, unless otherwise specified).
Table A1. Constants and parameters of input data normalisation and CNN-VAE (all values are non-dimensional, unless otherwise specified).
ParameterSymbolValue
Input Data
RowsF2560
ColumnsC N sig 2 = 36
CSM normalisation RMS ( G ) 10 3 ( m / s 2 ) 2
Non-linear scaling function exponentr3
VAE Training
Latent space dimension N lat 7
Loss KL Divergence weight β KL 0.125
Loss MIF weight β MIF 4
Loss SHP weight β SHP 0.05
Optimiser-ADAM
Learning Rate (reduced every 100 Epochs)- 10 3 , 10 4 , 10 5
Batch Size-60
Number of Epochs-300
Convolutional Layers
Number of filters N filt 48
Kernel size K F × K C 3 × 4
Stride ( S F , S C ) ( 2 , 1 )
Leaky ReLU slope for negative valuesa 0.01
Table A2. CNN-VAE architecture for modelling structure response CSM in terms of Cholesky factors.
Table A2. CNN-VAE architecture for modelling structure response CSM in terms of Cholesky factors.
Block TypeLayersOutput Size
Encoder
Input2D input F × C × 1
Convolutional Blocks (3×)2D Convolution F / 8 × C × N filt
LeakyReLU
Fully ConnectedFully Connected 2 N lat
OutputReparametrisation N lat
Sampling
Decoder
InputFeature input N lat
Project and ReshapeFully Connected F / 8 × C × N filt
Reshape
Convolutional Blocks (3×)2D Transp. Convolution F × C × N filt
LeakyReLU
Sum Channels2D Convolution ( 1 × 1 ) F × C × 1
Selective ActivationSelective ReLU F × C × 1
OutputHyperbolic Tangent F × C × 1

References

  1. Elsisi, A.; Zamrawi, A.; Emad, S. A Comprehensive Review of Structural Health Monitoring for Steel Bridges: Technologies, Data Analytics, and Future Directions. Appl. Sci. 2025, 15, 12090. [Google Scholar] [CrossRef]
  2. Azimi, M.; Eslamlou, A.D.; Pekcan, G. Data-Driven Structural Health Monitoring and Damage Detection through Deep Learning: State-of-the-Art Review. Sensors 2020, 20, 2778. [Google Scholar] [CrossRef]
  3. Spencer, B.F.; Sim, S.H.; Kim, R.E.; Yoon, H. Advances in artificial intelligence for structural health monitoring: A comprehensive review. KSCE J. Civ. Eng. 2025, 29, 100203. [Google Scholar] [CrossRef]
  4. Sarma, I.V.; Chanda, S.; Reddy, M.S. The convergence of artificial intelligence and structural health monitoring: A comprehensive review of methodologies, advancements, and future trajectories. Life Cycle Reliab. Saf. Eng. 2025, 14, 531–543. [Google Scholar] [CrossRef]
  5. Scarselli, G.; Nicassio, F. Machine Learning for Structural Health Monitoring of Aerospace Structures: A Review. Sensors 2025, 25, 6136. [Google Scholar] [CrossRef]
  6. Maria Tronci, E.; Downey, A.R.J.; Mehrjoo, A.; Chowdhury, P.; Coble, D. Physics-Informed Machine Learning Part I: Different Strategies to Incorporate Physics into Engineering Problems. In Data Science in Engineering; River Publishers: Aalborg, Denmark, 2025; Volume 10, pp. 1–6. [Google Scholar] [CrossRef]
  7. Downey, A.R.J.; Maria Tronci, E.; Chowdhury, P.; Coble, D. Physics-Informed Machine Learning Part II: Applications in Structural Response Forecasting. In Data Science in Engineering; River Publishers: Aalborg, Denmark, 2025; Volume 10, pp. 63–66. [Google Scholar] [CrossRef]
  8. Toh, G.; Park, J. Review of Vibration-Based Structural Health Monitoring Using Deep Learning. Appl. Sci. 2020, 10, 1680. [Google Scholar] [CrossRef]
  9. Avci, O.; Abdeljaber, O.; Kiranyaz, S.; Hussein, M.; Gabbouj, M.; Inman, D.J. A review of vibration-based damage detection in civil structures: From traditional methods to Machine Learning and Deep Learning applications. Mech. Syst. Signal Process. 2021, 147, 107077. [Google Scholar] [CrossRef]
  10. Narayanan, N.; Harsha, S.P. Advancements in Vibration-Based Damage Detection for Railway Tracks: A Review on Techniques and Emerging Trends in Predictive Maintenance. Int. J. Struct. Stab. Dyn. 2025. [Google Scholar] [CrossRef]
  11. Farrar, C.R.; Worden, K. Chapter 13: Fundamental Axioms of Structural Health Monitoring. In Structural Health Monitoring; John Wiley & Sons, Ltd.: Hoboken, NJ, USA, 2012; pp. 439–460. [Google Scholar] [CrossRef]
  12. Zahid, F.B.; Ong, Z.C.; Khoo, S.Y. A review of operational modal analysis techniques for in-service modal identification. J. Braz. Soc. Mech. Sci. Eng. 2020, 42, 398. [Google Scholar] [CrossRef]
  13. Hasani, H.; Freddi, F. Operational Modal Analysis on Bridges: A Comprehensive Review. Infrastructures 2023, 8, 172. [Google Scholar] [CrossRef]
  14. Ren, Y.; Bareille, O.; Lin, Z.; Huang, X.-R. Review of damage detection techniques in vibration-based structural health monitoring. Int. J. Dyn. Control 2025, 13, 99. [Google Scholar] [CrossRef]
  15. Zhang, C.; Mousavi, A.A.; Masri, S.F.; Gholipour, G.; Yan, K.; Li, X. Vibration feature extraction using signal processing techniques for structural health monitoring: A review. Mech. Syst. Signal Process. 2022, 177, 109175. [Google Scholar] [CrossRef]
  16. Lucà, F.; Pavoni, S.; Manzoni, S.; Vanali, M. Time Reliability of Empirical Models for the Prediction of Building Parameters: The Case of Palazzo Lombardia. In Proceedings of the European Workshop on Structural Health Monitoring; Rizzo, P., Milazzo, A., Eds.; Springer: Cham, Switzerland, 2023; pp. 186–194. [Google Scholar]
  17. Cross, E.J.; Manson, G.; Worden, K.; Pierce, S.G. Features for damage detection with insensitivity to environmental and operational variations. Proc. R. Soc. A Math. Phys. Eng. Sci. 2012, 468, 4098–4122. [Google Scholar] [CrossRef]
  18. Berardengo, M.; Lucà, F.; Manzoni, S.; Pavoni, S.; Vanali, M. A PCA/Natural Frequency-Based Approach for Damage Detection: Implementation on a Laboratory Structure Subjected to Environmental Variability. In Special Topics in Structural Dynamics & Experimental Techniques; River Publishers: Aalborg, Denmark, 2025; Volume 5, pp. 133–139. [Google Scholar] [CrossRef]
  19. Battista, G.; Pavoni, S.; Vanali, M. Compensation of Thermal Effects on Tiltmeter Measurements with Moving Least Squares. IEEE Trans. Instrum. Meas. 2024, 73, 7501914. [Google Scholar] [CrossRef]
  20. Mariani, S.; Kalantari, A.; Kromanis, R.; Marzani, A. Data-driven modeling of long temperature time-series to capture the thermal behavior of bridges for SHM purposes. Mech. Syst. Signal Process. 2024, 206, 110934. [Google Scholar] [CrossRef]
  21. Kamali, S.; Palermo, A.; Marzani, A. Virtual baseline to improve anomaly detection of SHM systems with non-stationary data. Mech. Syst. Signal Process. 2025, 224, 111968. [Google Scholar] [CrossRef]
  22. Crognale, M.; Rinaldi, C.; Ciambella, J.; Gattulli, V. Machine learning–enhanced structural health monitoring of the steel–glass exedra under year-round environmental loading. J. Low Freq. Noise Vib. Act. Control 2026. [Google Scholar] [CrossRef]
  23. Rezazadeh, N.; De Luca, A.; Perfetto, D.; Salami, M.R.; Lamanna, G. Systematic critical review of structural health monitoring under environmental and operational variability: Approaches for baseline compensation, adaptation, and reference-free techniques. Smart Mater. Struct. 2025, 34, 073001. [Google Scholar] [CrossRef]
  24. Kingma, D.P.; Welling, M. Auto-Encoding Variational Bayes. arXiv 2022, arXiv:1312.6114. [Google Scholar] [CrossRef]
  25. Taye, M.M. Theoretical Understanding of Convolutional Neural Network: Concepts, Architectures, Applications, Future Directions. Computation 2023, 11, 52. [Google Scholar] [CrossRef]
  26. Rai, A.; Ierimonti, L.; Giglioni, V.; Tomassini, E.; Ubertini, F.; Venanzi, I. A multivariate autoencoder-based anomaly detection for post-earthquake diagnosis of masonry structures using long-term accelerometer and ambient temperature recordings. J. Build. Eng. 2026, 117, 114784. [Google Scholar] [CrossRef]
  27. Amaral, R.; Gentile, C.; Barbosa, F.; Cury, A. Structural novelty detection based on sparse autoencoders and control charts. Struct. Eng. Mech. 2022, 81, 647–664. [Google Scholar] [CrossRef]
  28. Pirrò, M.; Gentile, C. Sparse Auto-Encoder Networks to Detect and Localize Structural Changes in Metallic Bridges. Buildings 2026, 16, 802. [Google Scholar] [CrossRef]
  29. Pamplona Berón, L.E.; De Simone, M.C.; de Falco, D.; Guida, D. Machine Learning-Based Classification of Vibration Patterns Under Multiple Excitation Scenarios for Structural Health Monitoring. Appl. Sci. 2026, 16, 2107. [Google Scholar] [CrossRef]
  30. Chen, H.; Wan, H.P.; Weng, Z.H.; Zhou, Y.; Xu, H.C. An approach for reconstruction of missing SHM data using AO-CNN-BiLSTM-Attention. J. Civ. Struct. Health Monit. 2026, 16, 16. [Google Scholar] [CrossRef]
  31. Bacsa, K.; Liu, W.; Abdallah, I.; Chatzi, E. Structural Dynamics Feature Learning Using a Supervised Variational Autoencoder. J. Eng. Mech. 2025, 151, 04024106. [Google Scholar] [CrossRef]
  32. Wan, H.P.; Zhu, Y.K.; Luo, Y.; Todd, M.D. Unsupervised deep learning approach for structural anomaly detection using probabilistic features. Struct. Health Monit. 2025, 24, 3–33. [Google Scholar] [CrossRef]
  33. Dadras Eslamlou, A.; Huang, S. Artificial-Neural-Network-Based Surrogate Models for Structural Health Monitoring of Civil Structures: A Literature Review. Buildings 2022, 12, 2067. [Google Scholar] [CrossRef]
  34. Tang, S.; Yin, X.; Wu, W. Uncertainty Quantification Analysis of Dynamic Responses in Plate Structures Based on a Physics-Informed CVAE Model. Appl. Sci. 2026, 16, 1496. [Google Scholar] [CrossRef]
  35. Di Mucci, V.M.; Cardellicchio, A.; Ruggieri, S.; Nettis, A.; Renò, V.; Uva, G. Artificial intelligence in structural health management of existing bridges. Autom. Constr. 2024, 167, 105719. [Google Scholar] [CrossRef]
  36. Gunes, B.; Gunes, O. An Unsupervised Hybrid Approach for Detection of Damage with Autoencoder and One-Class Support Vector Machine. Appl. Sci. 2025, 15, 4098. [Google Scholar] [CrossRef]
  37. Okur, F.Y.; Altunişik, A.C.; Kalkan Okur, E. A Novel Approach for Anomaly Detection in Vibration-Based Structural Health Monitoring Using Autoencoders in Deep Learning. Struct. Control Health Monit. 2025, 2025, 5602604. [Google Scholar] [CrossRef]
  38. Pirrò, M.; Gentile, C. Detection and localization of anomalies in bridges using accelerometer data and sparse auto-encoders. Dev. Built Environ. 2025, 23, 100715. [Google Scholar] [CrossRef]
  39. Komorska, I.; Puchalski, A. Condition Monitoring Using a Latent Space of Variational Autoencoder Trained Only on a Healthy Machine. Sensors 2024, 24, 6825. [Google Scholar] [CrossRef]
  40. Pakzad, S.S.; Masoodi, A.R. Early damage detection in bridges using a variational autoencoder–based hybrid unsupervised learning framework. Sci. Rep. 2025, 16, 1389. [Google Scholar] [CrossRef] [PubMed]
  41. Horn, R.A.; Johnson, C.R. Matrix Analysis, 2nd ed.; Cambridge University Press: New York, NY, USA, 2012. [Google Scholar]
  42. Lucà, F.; Gholibeiki, H.; Manzoni, S.; Mola, E.; Pavoni, S.; Vanali, M. Principal Component Analysis of Monitoring Data of a High-Rise Building: The Case Study of Palazzo Lombardia. In Data Science in Engineering; River Publishers: Aalborg, Denmark, 2023; Volume 10, pp. 35–42. [Google Scholar] [CrossRef]
  43. Pavoni, S.; Vanali, M.; Manzoni, S.; Lucà, F.; Berardengo, M. Comparison of Experimental and Operational Modal Analysis Results on Long-Term Monitoring of a Laboratory Truss Girder Subjected to Environmental Variability. In Proceedings of the 10th International Operational Modal Analysis Conference (IOMAC 2024); Rainieri, C., Gentile, C., Aenlle López, M., Eds.; Springer: Cham, Switzerland, 2024; pp. 315–325. [Google Scholar]
  44. Rainieri, C.; Fabbrocino, G. Output-only Modal Identification. In Operational Modal Analysis of Civil Engineering Structures: An Introduction and Guide for Applications; Springer: New York, NY, USA, 2014; pp. 103–210. [Google Scholar] [CrossRef]
  45. Brincker, R.; Zhang, L.; Andersen, P. Modal identification of output-only systems using frequency domain decomposition. Smart Mater. Struct. 2001, 10, 441. [Google Scholar] [CrossRef]
  46. Brincker, R.; Ventura, C.; Andersen, P. Damping Estimation by Frequency Domain Decomposition. In Proceedings of the IMAC 19, Kissimmee, FL, USA, 5–8 February 2001; pp. 698–703. [Google Scholar]
  47. Rainieri, C.; Fabbrocino, G. Operational Modal Analysis of Civil Engineering Structures: An Introduction and Guide for Applications; Springer: New York, NY, USA, 2014. [Google Scholar] [CrossRef]
  48. Hasan, M.D.A.; Ahmad, Z.A.B.; Leong, M.S.; Hee, L.M. Enhanced frequency domain decomposition algorithm: A review of a recent development for unbiased damping ratio estimates. J. Vibroeng. 2018, 20, 1919–1936. [Google Scholar] [CrossRef]
  49. Zhang, L.; Wang, T.; Tamura, Y. A frequency–spatial domain decomposition (FSDD) method for operational modal analysis. Mech. Syst. Signal Process. 2010, 24, 1227–1239. [Google Scholar] [CrossRef]
  50. Ewins, D.J. Modal Testing; Research Studies Press: Hertfordshire, UK, 2000. [Google Scholar]
  51. Schmidhuber, J. Deep learning in neural networks: An overview. Neural Netw. 2015, 61, 85–117. [Google Scholar] [CrossRef]
  52. Nesackon Abraham, J.; Tran, M.Q.; Jayaraj, J.S.; Matos, J.C.; Valluzzi, M.R.; Dang, S.N. Unsupervised Learning-Based Anomaly Detection for Bridge Structural Health Monitoring: Identifying Deviations from Normal Structural Behaviour. Sensors 2026, 26, 561. [Google Scholar] [CrossRef]
  53. Tomczak, J.M. Deep Generative Modeling; Springer International Publishing: Cham, Switzerland, 2024. [Google Scholar] [CrossRef]
  54. Ng, A. Sparse Autoencoder. Lecture Notes, Stanford University CS294A. 2011. Available online: http://web.stanford.edu/class/cs294a/sparseAutoencoder.pdf (accessed on 10 May 2026).
  55. Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Networks. arXiv 2014, arXiv:1406.2661. [Google Scholar] [CrossRef]
  56. Brandt, A. Noise and Vibration Analysis: Signal Analysis and Experimental Procedures; Wiley: Hoboken, NJ, USA, 2023. [Google Scholar] [CrossRef]
  57. Higgins, I.; Matthey, L.; Pal, A.; Burgess, C.; Glorot, X.; Botvinick, M.; Mohamed, S.; Lerchner, A. beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. In Proceedings of the International Conference on Learning Representations, Toulon, France, 24–26 April 2017. [Google Scholar]
  58. Burgess, C.P.; Higgins, I.; Pal, A.; Matthey, L.; Watters, N.; Desjardins, G.; Lerchner, A. Understanding disentangling in β-VAE. arXiv 2018, arXiv:1804.03599. [Google Scholar] [CrossRef]
  59. Steinmann, L.; Migenda, N.; Voigt, T.; Kohlhase, M.; Schenck, W. Variational Autoencoder based Novelty Detection for Real-World Time Series. In Proceedings of the 2021 3rd International Conference on Management Science and Industrial Engineering, New York, NY, USA, 2021; pp. 1–7. [Google Scholar] [CrossRef]
  60. Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar] [CrossRef]
  61. Farrar, C.R.; Worden, K. Chapter 12: Data Normalisation. In Structural Health Monitoring; John Wiley & Sons, Ltd.: Hoboken, NJ, USA, 2012; pp. 403–438. [Google Scholar] [CrossRef]
  62. Berardengo, M.; Lucà, F.; Vanali, M.; Annesi, G. Short-Training Damage Detection Method for Axially Loaded Beams Subject to Seasonal Thermal Variations. Sensors 2023, 23, 1154. [Google Scholar] [CrossRef]
  63. Weil, M.; Weijtjens, W.; Devriendt, C. Autoencoder and Mahalanobis distance for novelty detection in structural health monitoring data of an offshore wind turbine. J. Phys. Conf. Ser. 2022, 2265, 032076. [Google Scholar] [CrossRef]
  64. Hejazi, M.; Singh, Y.P. One-Class Support Vector Machines Approach to Anomaly Detection. Appl. Artif. Intell. 2013, 27, 351–366. [Google Scholar] [CrossRef]
  65. Li, G.; Wen, Z.; Xie, X. Unsupervised Anomaly Detection Based on CNN-VAE with Spectral Residual for KPIs. In Proceedings of the 2022 IEEE 24th International Conference on High Performance Computing & Communications; 8th International Conference on Data Science & Systems; 20th International Conference on Smart City; 8th International Conference on Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC/DSS/SmartCity/DependSys), Hainan, China, 18–20 December 2022; pp. 1307–1313. [Google Scholar] [CrossRef]
  66. Daoud, A.M.; Elkomy, O.M.; Khedr, W.I.; Hosny, K.M. A new hybrid model for improving outlier detection using combined autoencoder and variational autoencoder. Sci. Rep. 2025, 15, 43387. [Google Scholar] [CrossRef]
  67. Milewski, M.; Wróbel, J.; Kierzkowski, A.; Vališ, D. Experimental and numerical modal analysis of an unmanned aerial vehicle’s composite wing. Simul. Model. Pract. Theory 2025, 142, 103106. [Google Scholar] [CrossRef]
Figure 1. Truss structure installed for SHM research at University of Parma vibration laboratory. (a) Picture of laboratory truss and measurement setup. (b) Scheme of test setup showing structure node numbers (circled numbers) and location of accelerometers, thermocouples and shaker.
Figure 1. Truss structure installed for SHM research at University of Parma vibration laboratory. (a) Picture of laboratory truss and measurement setup. (b) Scheme of test setup showing structure node numbers (circled numbers) and location of accelerometers, thermocouples and shaker.
Applsci 16 04844 g001
Figure 2. Temperature measured during the baseline monitoring period from April to early July 2023 and the histogram with relative frequencies.
Figure 2. Temperature measured during the baseline monitoring period from April to early July 2023 and the histogram with relative frequencies.
Applsci 16 04844 g002
Figure 3. Maximum singular value ( s 1 ) spectrum distribution of the benchmark structure response. A histogram is computed independently for each discrete frequency from the baseline data. Grey regions: bands of analysis B m .
Figure 3. Maximum singular value ( s 1 ) spectrum distribution of the benchmark structure response. A histogram is computed independently for each discrete frequency from the baseline data. Grey regions: bands of analysis B m .
Applsci 16 04844 g003
Figure 4. Natural frequencies f m of the baseline dataset for the benchmark structure. Plots on the diagonal depict the histograms for each natural frequency.
Figure 4. Natural frequencies f m of the baseline dataset for the benchmark structure. Plots on the diagonal depict the histograms for each natural frequency.
Applsci 16 04844 g004
Figure 5. Temperature intervals for each test condition. The central dot is the mean values and the interval endpoints and dashed lines are the 5th and the 95th percentiles. The histogram on the left shows the histogram of relative frequencies over the baseline as a reference. Background colors represent different structure conditions. Green: baseline and healthy checks. Red: added mass. Yellow: bolt torque variation. Blue: bolt removal.
Figure 5. Temperature intervals for each test condition. The central dot is the mean values and the interval endpoints and dashed lines are the 5th and the 95th percentiles. The histogram on the left shows the histogram of relative frequencies over the baseline as a reference. Background colors represent different structure conditions. Green: baseline and healthy checks. Red: added mass. Yellow: bolt torque variation. Blue: bolt removal.
Applsci 16 04844 g005
Figure 6. Flow chart for training of the CNN-VAE model and its integration with vibration-based SHM.
Figure 6. Flow chart for training of the CNN-VAE model and its integration with vibration-based SHM.
Applsci 16 04844 g006
Figure 7. A scheme of a typical VAE architecture. The Encoder processes the input and returns the parameters of probability distributions for each latent variable. The Decoder reconstructs the input data from random samples of the latent variable. An external random sample generator ϵ decouples the stochastic part from the rest of the network (reparametrisation trick).
Figure 7. A scheme of a typical VAE architecture. The Encoder processes the input and returns the parameters of probability distributions for each latent variable. The Decoder reconstructs the input data from random samples of the latent variable. An external random sample generator ϵ decouples the stochastic part from the rest of the network (reparametrisation trick).
Applsci 16 04844 g007
Figure 8. Encoder and Decoder network scheme. Both networks can be setup with any number of convolutional blocks in series.
Figure 8. Encoder and Decoder network scheme. Both networks can be setup with any number of convolutional blocks in series.
Applsci 16 04844 g008
Figure 9. Average value of Loss terms compared between CNN-VAE models trained using the Classic LF and the Enhanced LF. Top plot shows the absolute values. Centre plot shows the weighted values. Bottom plot shows the contribution in percentage of each weighted value to the total Loss Value.
Figure 9. Average value of Loss terms compared between CNN-VAE models trained using the Classic LF and the Enhanced LF. Top plot shows the absolute values. Centre plot shows the weighted values. Bottom plot shows the contribution in percentage of each weighted value to the total Loss Value.
Applsci 16 04844 g009
Figure 10. Comparison of natural frequencies f m in the baseline dataset between measured data (in green) and DGEN data (in red). Plots on the diagonal depict the histograms for each natural frequency. (a) CNN-VAE trained with Enhanced LF. (b) CNN-VAE trained with Classic LF.
Figure 10. Comparison of natural frequencies f m in the baseline dataset between measured data (in green) and DGEN data (in red). Plots on the diagonal depict the histograms for each natural frequency. (a) CNN-VAE trained with Enhanced LF. (b) CNN-VAE trained with Classic LF.
Applsci 16 04844 g010aApplsci 16 04844 g010b
Figure 11. Comparison of natural frequencies f m in the baseline dataset between measured data (in green) and RGEN data (in purple). Plots on the diagonal depict the histograms for each natural frequency. (a) CNN-VAE trained with Enhanced LF. (b) CNN-VAE trained with Classic LF.
Figure 11. Comparison of natural frequencies f m in the baseline dataset between measured data (in green) and RGEN data (in purple). Plots on the diagonal depict the histograms for each natural frequency. (a) CNN-VAE trained with Enhanced LF. (b) CNN-VAE trained with Classic LF.
Applsci 16 04844 g011aApplsci 16 04844 g011b
Figure 12. Distribution of Absolute Relative Error of the MIF spectrum over the baseline dataset. A histogram is computed independently for each discrete frequency. Grey regions: bands of analysis B m . (a) Measured versus DGEN synthetic baseline dataset. (b) Measured versus RGEN synthetic baseline dataset.
Figure 12. Distribution of Absolute Relative Error of the MIF spectrum over the baseline dataset. A histogram is computed independently for each discrete frequency. Grey regions: bands of analysis B m . (a) Measured versus DGEN synthetic baseline dataset. (b) Measured versus RGEN synthetic baseline dataset.
Applsci 16 04844 g012
Figure 13. Histogram of temperatures for experimental baseline (grey bars) and for balanced baseline (orange bars).
Figure 13. Histogram of temperatures for experimental baseline (grey bars) and for balanced baseline (orange bars).
Applsci 16 04844 g013
Figure 14. Mahalanobis Distance of mode natural frequencies for different damage conditions using the experimental baseline (in black) or the balanced baseline (in orange) as the reference dataset. Background colors represent different structure conditions. Green: baseline and healthy checks. Red: added mass. Yellow: bolt torque variation. Blue: bolt removal. (a) Comparison of MD statistics between different baselines. The interval endpoints are the 5th and 95th percentiles, while the central dot is the mean value. The threshold depicted as the dashed black line is the 95th percentile of MDs with respect to the baseline itself. All MDs calculated with respect to the same baseline are normalised by this value. (b) Increment of MD mean using a balanced baseline compared to using the experimental baseline.
Figure 14. Mahalanobis Distance of mode natural frequencies for different damage conditions using the experimental baseline (in black) or the balanced baseline (in orange) as the reference dataset. Background colors represent different structure conditions. Green: baseline and healthy checks. Red: added mass. Yellow: bolt torque variation. Blue: bolt removal. (a) Comparison of MD statistics between different baselines. The interval endpoints are the 5th and 95th percentiles, while the central dot is the mean value. The threshold depicted as the dashed black line is the 95th percentile of MDs with respect to the baseline itself. All MDs calculated with respect to the same baseline are normalised by this value. (b) Increment of MD mean using a balanced baseline compared to using the experimental baseline.
Applsci 16 04844 g014aApplsci 16 04844 g014b
Figure 15. Mahalanobis Distance of latent space distribution parameters using the experimental baseline as the reference. For each structure condition, the interval endpoints are the 5th and 95th percentiles, while the central dot is the mean value. The threshold depicted as the dashed black line is the 95th percentile of MDs with respect to the baseline itself. All MDs are normalised by this value. Background colors represent different structure conditions. Green: baseline and healthy checks. Red: added mass. Yellow: bolt torque variation. Blue: bolt removal. (a) Enhanced LF. (b) Classic LF.
Figure 15. Mahalanobis Distance of latent space distribution parameters using the experimental baseline as the reference. For each structure condition, the interval endpoints are the 5th and 95th percentiles, while the central dot is the mean value. The threshold depicted as the dashed black line is the 95th percentile of MDs with respect to the baseline itself. All MDs are normalised by this value. Background colors represent different structure conditions. Green: baseline and healthy checks. Red: added mass. Yellow: bolt torque variation. Blue: bolt removal. (a) Enhanced LF. (b) Classic LF.
Applsci 16 04844 g015
Table 1. Bands of analysis B m to monitor natural frequencies of structural modes. Lower and upper frequencies of the intervals are expressed in hertz.
Table 1. Bands of analysis B m to monitor natural frequencies of structural modes. Lower and upper frequencies of the intervals are expressed in hertz.
B 1 B 2 B 3 B 4 B 5
Lower2555177189196
Upper3060183194202
Table 2. Artificial damage tests. The available combinations of damage type and location are marked with “√”.
Table 2. Artificial damage tests. The available combinations of damage type and location are marked with “√”.
Damage TypeDamage Location
N1N4N5N6N9N10N11N13
Mass1 kg
Bolt Torque30 Nm
20 Nm
10 Nm
6 Nm
Bolt Left2 Bolts
4 Bolts
Table 3. Root-Mean-Squared-Error of natural frequencies f m between measured data in the baseline and synthetic data. The comparison involves data generated by CNN-VAE models trained with both Enhanced and Classic LF. All values are expressed in hertz.
Table 3. Root-Mean-Squared-Error of natural frequencies f m between measured data in the baseline and synthetic data. The comparison involves data generated by CNN-VAE models trained with both Enhanced and Classic LF. All values are expressed in hertz.
f 1 f 2 f 3 f 4 f 5
Enhanced VAEDGEN0.0250.0610.0590.140.078
RGEN0.0290.0690.0630.140.081
Classic VAEDGEN0.0320.0850.0660.140.092
RGEN0.0400.0970.0730.150.097
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Battista, G.; Pavoni, S.; Vanali, M. A Vibration Measurement Data Enhancement Approach Based on Variational Autoencoders for Structural Health Monitoring. Appl. Sci. 2026, 16, 4844. https://doi.org/10.3390/app16104844

AMA Style

Battista G, Pavoni S, Vanali M. A Vibration Measurement Data Enhancement Approach Based on Variational Autoencoders for Structural Health Monitoring. Applied Sciences. 2026; 16(10):4844. https://doi.org/10.3390/app16104844

Chicago/Turabian Style

Battista, Gianmarco, Stefano Pavoni, and Marcello Vanali. 2026. "A Vibration Measurement Data Enhancement Approach Based on Variational Autoencoders for Structural Health Monitoring" Applied Sciences 16, no. 10: 4844. https://doi.org/10.3390/app16104844

APA Style

Battista, G., Pavoni, S., & Vanali, M. (2026). A Vibration Measurement Data Enhancement Approach Based on Variational Autoencoders for Structural Health Monitoring. Applied Sciences, 16(10), 4844. https://doi.org/10.3390/app16104844

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop