To rigorously evaluate the proposed Event-Triggered Hybrid State-Space LSTM-LNN architecture, we design a comprehensive experimental framework that encompasses data acquisition, preprocessing, model training, and comparative evaluation. The experimental design is motivated by the need to validate the model’s ability to address state drift under prolonged seismic excitation while maintaining high classification accuracy for seismic vulnerability state assessment of ancient halls.
5.1. Data Collection and Description
The experimental study utilizes two complementary data sources: synthetic data generated from a calibrated finite element model and real monitoring data collected from an instrumented ancient hall. The target structure is a representative timber-framed ancient hall located in Rucheng, Hunan Province, China, which has been instrumented with a multi-source structural health monitoring system.
The synthetic dataset is generated using a high-fidelity finite element model of the ancient hall, calibrated against ambient vibration tests and material property characterization. The finite element model was developed within the ABAQUS 6.14 version/Standard framework using a three-dimensional wire-frame representation of the target ancient hall.
To provide a rigorous physical basis for the synthetic data generation, the FE modelling incorporates specific fundamental assumptions, element formulations, and solution algorithms. The nonlinear time-history analyses were conducted using the implicit dynamic integration algorithm (Newmark-β method) within the ABAQUS/Standard framework. The fundamental assumptions include the validity of the plane-section hypothesis for beam–column members and the activation of large-displacement geometric nonlinearity (NLGEOM option) to accurately capture the P-Delta effects and column rocking behaviors inherent in traditional timber structures. Regarding nonlinearities, a selective approach was adopted to balance computational efficiency with physical fidelity: the timber members themselves were modeled as linearly elastic, as material yielding is rare under the considered seismic intensity levels; however, critical connection nonlinearities were explicitly included. The semi-rigid mortise–tenon joints were simulated using nonlinear spring-damper elements to capture the gap closure, slippage, and pinching effects, the specific parameters of which are detailed below.
The timber frame comprises 12 primary columns, 8 main beams, and 24 mortise–tenon joints, with all members discretized using B31 Euler–Bernoulli beam elements. Material properties for the aged Chinese fir (Cunninghamia lanceolata) were assigned based on destructive and non-destructive tests conducted on surrogate specimens extracted from structurally non-critical locations of the same hall. The mean density was determined as 420 kg/m3, with longitudinal elastic modulus = 9.8 GPa, radial modulus = 0.62 GPa, and tangential modulus = 0.48 GPa, calibrated against dynamic resonance testing results. The moisture-dependent degradation of elastic moduli was modeled through a linear reduction factor of 2.1% per 1% increase in moisture content above the 12% reference level. Mortise–tenon connections were simulated using a nonlinear spring-damper element with a modified Bouc–Wen hysteresis model, capturing the semi-rigid rotational behavior, gap closure, and pinching effects characteristic of traditional Chinese timber joints under cyclic loading. The rotational backbone curve parameters—initial stiffness = 2.35 × 103 kN·m/rad, yield rotation = 0.018 rad, and post-yield stiffness ratio α = 0.12—were identified through quasi-static cyclic tests on full-scale joint subassemblies. Rayleigh damping with coefficients = 0.28 and = 0.0095 was applied to match the first three modal damping ratios (2.1%, 2.8%, and 3.4%) derived from ambient vibration tests. The foundation was modeled as fixed supports at column bases, a simplification justified by the massive stone plinths and compacted gravel base observed during on-site excavation, which exhibited negligible rotational compliance in ambient tests. The calibrated FE model achieved first three natural frequencies of 1.82 Hz, 3.76 Hz, and 5.43 Hz, deviating from the experimental values by less than 3.1%, confirming its fidelity for subsequent nonlinear time-history analyses. The synthetic dataset was generated by subjecting the FE model to 120 scaled ground motion records from the PEER NGA-West2 database (Mw of 5.0–7.5, epicentral distances of 10–100 km, PGA of 0.05 g–0.80 g), with the structural response sampled at the same sensor locations and sampling rates as the real monitoring system to ensure consistency in the multi-source data representation. This physically anchored simulation framework ensures that the synthetic samples embody plausible damage mechanisms, including beam–column joint slippage, column rocking, and localized buckling at mortise corners, thereby providing a robust basis for pre-training the proposed hybrid neural architecture.
The finite element model incorporates the nonlinear behavior of timber connections, including semi-rigid beam–column joints and mortise–tenon connections, which are critical to the seismic response of traditional Chinese timber structures [
46]. Ground motion excitations are simulated using a suite of recorded earthquake accelerograms from the Pacific Earthquake Engineering Research (PEER) ground motion database [
47], scaled to various peak ground acceleration levels ranging from 0.05 g to 0.8 g to represent different seismic intensity levels. For each ground motion, the structural response is computed using nonlinear time-history analysis, generating synchronized acceleration, strain, displacement, temperature, and humidity signals at multiple sensor locations.
To provide a clear visual representation of the examined structural system and address the configuration details,
Figure 4 presents a 3D wireframe snapshot of the calibrated ABAQUS finite element model. The model explicitly captures the spatial topology of the ancient timber hall. Regarding the mechanical constitution, the timber members are modeled as deformable B31 Euler–Bernoulli beam elements; however, to accurately capture the column rocking behaviors and P-Delta effects inherent in traditional timber structures, the large-displacement geometric nonlinearity (NLGEOM) is activated. Furthermore, the semi-rigid nature of the 24 mortise–tenon joints (highlighted with local magnifications) is simulated via nonlinear spring-damper elements. For the boundary conditions, the column bases are modeled as fixed supports, rigorously justified by the massive stone plinths observed on-site. Beyond the structural configuration,
Figure 4 also intuitively maps the heterogeneous sensor network onto the FE model. Distinct markers detail the exact spatial installation positions of the tri-axial accelerometers (roof and mid-height columns), strain gauges (critical joints), displacement sensors (column bases and beam ends), and environmental sensors (eaves and indoors), providing a clear spatial correspondence between the numerical topology and the physical sensor layout detailed in the subsequent sections.
The real monitoring dataset comprises continuous measurements collected over a 24-month (from January 2024 to December 2025) period from the instrumented ancient hall.
For the real monitoring windows, the same EDP-based criteria are applied. The natural frequencies are extracted via operational modal analysis (OMA) from ambient vibration records at the beginning of each window. Strain ratios are computed from strain gauges at four critical beam–column joints, with the yield strain predetermined from material coupon tests on surrogate specimens. Residual displacements are obtained by averaging the terminal 5 s of displacement sensor recordings after removing the baseline drift. Windows that simultaneously satisfy all three EDP thresholds for a given class are assigned that label. In cases where EDPs fall into different classes (e.g., frequency shift indicating Class 1 but residual displacement indicating Class 2), the most severe class is assigned as a conservative measure. This protocol was applied independently by three structural engineers, achieving a Fleiss’ kappa of 0.84, indicating near-perfect inter-rater agreement.
To ensure label reliability and reproducibility, each 30-second window was assigned a four-class seismic vulnerability index (Class 0–3) following a structured expert assessment protocol that integrated four quantitative indicators: (i) relative shifts in the first three natural frequencies with respect to baseline values obtained from ambient vibration tests prior to the monitoring period; (ii) peak-to-yield strain ratios derived from strain gauges at critical beam–column joints, calibrated against material coupon tests; (iii) residual displacements recorded at column bases and main beams within the terminal 5 s of each window; and (iv) on-site visual inspection records, including crack width measurements and joint looseness grades, corresponding to the same time windows. The class boundaries were anchored to predefined numerical thresholds derived from nonlinear time-history analysis of the calibrated finite element model under incremental ground motion intensities, thereby mapping qualitative damage descriptions to quantifiable engineering demand parameters. To assess inter-rater reliability and control expert subjectivity, three independent experts with over 10 years of experience in heritage structural assessment first labeled the windows independently. The Fleiss’ kappa coefficient among the three experts was 0.84, indicating near-perfect agreement. Discrepant cases (approximately 12% of the total windows) were resolved through consensus discussion with access to raw time-history data and auxiliary sensor channels. This protocol provides a transparent, repeatable labeling framework that mitigates subjectivity and supports the supervised learning task in this study.
The sensor network includes tri-axial accelerometers installed at the roof level and mid-height columns, strain gauges attached to critical timber members, linear variable differential transformers for displacement measurement, and environmental sensors for temperature and humidity monitoring. The sampling rates vary across sensor types: accelerometers sample at 200 Hz, strain gauges at 100 Hz, displacement sensors at 50 Hz, and environmental sensors at 0.1 Hz. This heterogeneity in sampling rates presents a practical challenge that our continuous-time LNN formulation naturally accommodates through its ability to evaluate dynamics at arbitrary time points.
To construct the training and evaluation datasets, we segment the continuous monitoring records into fixed-length windows of 30 s, corresponding to the typical duration of significant seismic events. Each window is labeled with a four-class seismic vulnerability index based on expert assessment and damage inspection reports: Class 0 (no damage), Class 1 (minor damage), Class 2 (moderate damage), and Class 3 (severe damage).
To ensure quantitative rigor and reproducibility, the class boundaries for the four damage states are defined based on three engineering demand parameters (EDPs) derived from the calibrated finite element model under incremental ground motion intensities: (i) relative shift in the first three natural frequencies with respect to baseline values, ; (ii) peak-to-yield strain ratio at critical beam–column joints,; and (iii) residual displacement at column bases and main beam ends, (in mm). The threshold intervals for each class are anchored to nonlinear time-history analysis results and prior experimental studies on traditional timber joints: Class 0 (No damage): < 2%, < 0.3, < 0.5 mm; Class 1 (Minor damage): 2% ≤ < 5%, 0.3 ≤ < 0.6, 0.5 ≤ < 2.0 mm; Class 2 (Moderate damage): 5% ≤ < 12%, 0.6 ≤ < 0.9, 2.0 ≤ < 6.0 mm; Class 3 (Severe damage): ≥ 12%, ≥ 0.9, ≥ 6.0 mm. These thresholds correspond to distinct physical mechanisms: elastic reversible deformation (Class 0), onset of micro-cracking and joint slip (Class 1), significant stiffness degradation and permanent joint rotation (Class 2), and near-collapse conditions with substantial load-path redistribution (Class 3).
The synthetic dataset provides 5000 labeled windows, while the real monitoring dataset contributes 1200 labeled windows.
The 1200 real monitoring windows were labeled according to a quantitative protocol based on three engineering demand parameters: (i) relative shift of the first three natural frequencies with respect to baseline values, (ii) peak-to-yield strain ratios at critical beam–column joints, and (iii) residual displacements at column bases and main beam ends. The threshold boundaries defining the four damage classes were derived from nonlinear time-history analyses of the calibrated finite element model under incremental ground motion intensities (Class 0: frequency shift < 2%, strain ratio < 0.3, residual displacement < 0.5 mm; Class 1: 2–5%, 0.3–0.6, 0.5–2.0 mm; Class 2: 5–12%, 0.6–0.9, 2.0–6.0 mm; Class 3: >12%, >0.9, >6.0 mm). As summarized in
Table 1, the majority of real windows (78%) were classified as Class 0 under routine ambient conditions. Class 1 windows (18%) originated from two typhoon events and one nearby blasting operation, inducing recoverable elastic/micro-plastic responses. Class 2 windows (3.2%) corresponded to a foundation differential settlement event and a creep-related joint degradation case identified during routine inspection. Critically, no Class 3 (severe damage) event was observed during the 24-month(from January 2024 to December 2025) monitoring period, reflecting the low-to-moderate seismicity of the region over this specific interval. The absence of real severe damage samples is explicitly addressed in
Section 6 through supplementary validation strategies.
During the 24-month (from January 2024 to December 2025) monitoring period, seven seismic events (epicentral distance > 100 km, local magnitude M of 2.8–4.2, site PGA of 0.008 g–0.035 g) were recorded, all inducing only elastic structural responses categorized as Class 0 (no damage). Class 1 windows (18% of the real dataset) primarily originated from two typhoon events (maximum wind speed > 20 m/s) and one nearby blasting operation (PGA ≈ 0.06 g–0.09 g). Class 2 windows accounted for 38 windows (approximately 3.2%), associated with a localized foundation differential settlement event and a routine inspection window capturing cumulative creep effects at a beam–column joint. Critically, no Class 3 (severe damage) event was observed during the monitoring period, reflecting the low-to-moderate seismicity of the region over the specific two-year interval.
The datasets are split into training (70%), validation (15%), and test (15%) subsets, ensuring that windows from the same seismic event are not distributed across different subsets to prevent data leakage.
5.3. Model Configuration and Training Details
The proposed hybrid architecture is configured with the following hyperparameters, determined through a combination of Bayesian optimization and manual tuning. The LNN hidden state dimension is set to , with a hidden layer dimension of in the derivative network. The fixed-step Runge–Kutta integrator uses an internal step size of s, corresponding to 10 integration steps per 50 Hz sampling interval. The LSTM pathway employs a two-layer architecture with hidden state dimension per layer, a dropout rate of 0.3 applied between layers, and a sequence length of 1500 time steps (30 s at 50 Hz).
The event-triggered reset mechanism uses a drift accumulation dimension of , with forgetting factor and scaling parameter . The trigger threshold is initialized to and treated as a learnable parameter during fine-tuning. The attention-based fusion layer produces a fused representation of dimension , which is passed through a two-layer feedforward classification network with hidden dimension 64 and a softmax output layer for four-class classification.
Training proceeds in two stages. In the first stage, the model is pre-trained on the synthetic dataset for 200 epochs using the Adam optimizer [
49] with an initial learning rate of
, which is decayed by a factor of 0.5 every 50 epochs. The loss weights in Equation (19) are initialized as
,
, and
, and are adaptively updated during training using the gradient-based balancing approach described in
Section 4.5. Early stopping is applied based on the validation loss to prevent overfitting, with a patience of 20 epochs.
In the second stage, the pre-trained model is fine-tuned on the real monitoring dataset for 100 epochs with a reduced learning rate of . During fine-tuning, the LNN derivative network parameters are frozen for the first 20 epochs to allow the LSTM pathway and classification head to adapt to the real data distribution, after which all parameters are jointly optimized. The trigger threshold is unfrozen during fine-tuning, allowing the model to learn an optimal reset policy specific to the real monitoring conditions.
The trigger threshold τ is unfrozen during fine-tuning, allowing the model to learn an optimal reset policy specific to the real monitoring conditions.
To ensure reproducibility and clearly delineate the design of the continuous-time and discrete-time pathways within the hybrid architecture,
Table 3 details the core hyperparameters of the LNN, LSTM, event-triggered reset mechanism, and the two-stage transfer learning strategy.
This table systematically organizes the key architectural and training details of the hybrid model. In particular, the sinusoidal activation function and RK4 solver configuration for the LNN, along with the “freeze-then-unfreeze” strategy for the LNN derivative network during transfer learning, provide a clear recipe for researchers seeking to reproduce the model or adapt it to other heritage structural monitoring scenarios.
To ensure the reproducibility of the proposed methodology, all deep learning models, including the hybrid LSTM-LNN architecture and the baseline networks, were implemented using the PyTorch 2.14.0 framework, which has been widely adopted and validated in recent data-driven structural assessment studies [
50]. The continuous-time Liquid Neural Network (LNN) dynamics and Neural Ordinary Differential Equations (Neural ODEs) were solved utilizing the torchdiffeq library, leveraging the continuous-time modeling capabilities that have demonstrated superior adaptability in recent time-series forecasting and engineering applications [
51]. For the Bayesian optimization of the LSTM hyperparameters, the Optuna library was employed, following the machine learning optimization protocols established in recent structural engineering studies [
52]. Data preprocessing, including the band-pass filtering and resampling of heterogeneous sensor signals described in
Section 5.2, were conducted using SciPy v1.18.0 [
53]. The traditional machine learning baseline (SVM) and the quantitative evaluation metrics were implemented via the scikit-learn package [
54]. All models were trained and evaluated on a high-performance workstation equipped with dual Intel Xeon Gold CPUs, 256 GB of RAM, and four NVIDIA A100 (80 GB) GPUs, accelerated by the CUDA 11.8 toolkit. The source code and the pre-trained model weights will be made publicly available upon acceptance to facilitate further research in the structural health monitoring community.
5.4. Baseline Methods for Comparative Evaluation
To assess the effectiveness of the proposed architecture, we compare its performance against several baseline methods representing different paradigms in structural health monitoring and time-series classification. The baseline methods are selected to cover discrete-time recurrent models, continuous-time neural models, and traditional machine learning approaches.
The first baseline is a standard two-layer LSTM network with the same hidden state dimension and training procedure as the LSTM pathway in our proposed model, but without the continuous-time LNN component or the event-triggered reset mechanism. This baseline isolates the contribution of the LNN pathway and reset mechanism to overall performance.
The second baseline is a pure Liquid Neural Network with the same derivative network architecture and ODE solver configuration as the LNN pathway in our proposed model, but without the LSTM pathway, attention fusion, or reset mechanism. This baseline evaluates the standalone capability of continuous-time dynamics for seismic vulnerability state assessment.
The third baseline is a hybrid CNN-LSTM architecture that applies one-dimensional convolutional layers to extract local temporal features from the multi-channel sensor data before feeding them into an LSTM network. This baseline represents a state-of-the-art approach in structural damage detection that does not incorporate continuous-time dynamics or event-triggered resets.
The fourth baseline is a traditional machine learning approach based on feature extraction and classification. We extract statistical features from each sensor channel, including mean, variance, skewness, kurtosis, zero-crossing rate, and spectral centroid, and train a support vector machine (SVM) classifier with a radial basis function kernel [
55,
56]. This baseline represents the conventional approach to vibration-based damage identification.
All baseline models are trained and evaluated under identical data splits, preprocessing procedures, and evaluation metrics to ensure fair comparison. For the SVM baseline, features are extracted from the same 30-second windows used for the deep learning models, and hyperparameters are optimized using grid search with cross-validation on the training subset.
5.5. Evaluation Metrics
The performance of all models is evaluated using multiple metrics that capture different aspects of the seismic vulnerability state assessment task. The primary metric is classification accuracy, defined as the proportion of correctly classified windows in the test subset. However, given the potential class imbalance in the real monitoring dataset—where severe damage events are rare—we also report the macro-averaged F1-score, which computes the F1-score for each class independently and averages them, providing a balanced measure that is less sensitive to class distribution.
To specifically evaluate the state drift prevention capability of the proposed architecture, we introduce a drift metric that measures the divergence of the LNN hidden state from a reference trajectory over long sequences.
To quantitatively evaluate the state drift prevention capability of the proposed architecture, we define a physics-anchored drift metric that measures the divergence between the model’s reconstructed structural responses and a physically grounded reference trajectory. The reference trajectory is obtained by feeding the identical input excitation to a calibrated linearized finite element model of the ancient hall, which was validated against ambient vibration tests with first three modal frequency deviations below 3%. The LNN hidden state is projected back to the physical space (acceleration and strain at critical sensor locations) via a trained decoder, and the drift metric is computed as the mean squared error between this decoded physical response and the finite-element reference output over each 30-second window. This formulation anchors the drift assessment to independently validated structural dynamics, providing an objective and physically meaningful quantification of state divergence.
To examine the sensitivity of the reported drift values to the choice of the reset interval used in the reference model, we varied the fixed reset period
from 1 s to 30 s and recomputed the drift metric for both the pure LNN and the proposed hybrid model. The absolute drift values of the pure LNN ranged from 8.7 × 10
−3 (at
= 1 s) to 21.3 × 10
−3 (at
= 30 s), whereas the proposed model maintained consistently lower drift across all reset intervals, achieving a relative reduction of 72–81% compared to the pure LNN. The specific values reported in
Table 3 (2.18 × 10
−3 for the proposed model and 12.45 × 10
−3 for the pure LNN) were obtained with
= 5 s, a choice motivated by the structure’s fundamental period (
≈ 1.8) such that the reset interval is shorter than three times the fundamental period to preserve trajectory continuity while still allowing sufficient dynamic evolution between resets. The relative ranking and statistical significance of the drift reduction remain consistent across the entire range of reset intervals examined, confirming that the reported drift suppression advantage of the proposed event-triggered mechanism is robust to the particular choice of the reset period.
For each test window, we compute the mean squared error between the LNN hidden state at each time step and the hidden state obtained from a reference model that is periodically reset at fixed intervals. A lower drift metric indicates better state stability over prolonged sequences.
Additionally, we report the reconstruction error on the test subset, measured as the mean squared error between the input sensor signals and the model’s reconstructed outputs. This metric reflects the model’s ability to maintain an accurate internal representation of the structural dynamics, which is essential for reliable vulnerability assessment.
For the classification task, we also generate confusion matrices to visualize the distribution of prediction errors across the four vulnerability classes, providing insights into the types of misclassifications that occur and whether they are concentrated in adjacent classes (e.g., minor vs. moderate damage) or involve more severe errors.