Figure 1.
Lamb-wave group-velocity dispersion curves for the S0 and A0 modes. The marker indicates MHz, where mm/s and mm/s; the grey and green shaded bands represent the sample-level drift envelopes of the S0 and A0 modes, respectively (Gaussian perturbation with of the nominal group velocity).
Figure 1.
Lamb-wave group-velocity dispersion curves for the S0 and A0 modes. The marker indicates MHz, where mm/s and mm/s; the grey and green shaded bands represent the sample-level drift envelopes of the S0 and A0 modes, respectively (Gaussian perturbation with of the nominal group velocity).
Figure 2.
Simulation setup and the five defect geometries considered in this work: (a) no-defect; (b) circular hole; (c) crack; (d) corrosion zone; (e) weld seam. The Tx–Rx probes are positioned on either side of the inspected region at mm separation; the FEM mesh (light blue triangles) is the same gmsh-generated unstructured mesh used by the Lamb-wave forward solver.
Figure 2.
Simulation setup and the five defect geometries considered in this work: (a) no-defect; (b) circular hole; (c) crack; (d) corrosion zone; (e) weld seam. The Tx–Rx probes are positioned on either side of the inspected region at mm separation; the FEM mesh (light blue triangles) is the same gmsh-generated unstructured mesh used by the Lamb-wave forward solver.
Figure 3.
Architecture overview of the proposed EMAT-PINN. This figure is a schematic overview of the proposed EMAT-PINN, not a layer-by-layer diagram. Arrows indicate the direction of data flow. It illustrates three design ideas: (i) the embedding of the physics prior into the network through four parallel physics heads, one per physical quantity (
/
/
/
); (ii) the correspondence between each head and its paired defect class via the bias-free additive projection onto the class logit; and (iii) the use of a physics-constrained loss for gradient guidance. The blocks labelled “dual-pooling trunk” and “task heads” are conceptual placeholders. The precise layer configuration, the logit-construction equations (
,
), and the parameter count (
M) are given in
Table 2.
Figure 3.
Architecture overview of the proposed EMAT-PINN. This figure is a schematic overview of the proposed EMAT-PINN, not a layer-by-layer diagram. Arrows indicate the direction of data flow. It illustrates three design ideas: (i) the embedding of the physics prior into the network through four parallel physics heads, one per physical quantity (
/
/
/
); (ii) the correspondence between each head and its paired defect class via the bias-free additive projection onto the class logit; and (iii) the use of a physics-constrained loss for gradient guidance. The blocks labelled “dual-pooling trunk” and “task heads” are conceptual placeholders. The precise layer configuration, the logit-construction equations (
,
), and the parameter count (
M) are given in
Table 2.
Figure 4.
Four-head activation ablation on a single type-2 (crack) test sample under the proposed architecture: (a) raw A-scan; (b) head’s body activation (8-segment AdaptiveMaxPool, channel-mean, up-sampled to 1024 time points); (c) head’s activation on the same sample; (d) stacked sum of all four heads. The defect-to-direct energy ratio (defect-segment energy divided by direct-wave energy) is reported in each panel caption: the head scores highest because mode-discriminability is the dominant physical cue for crack classification, while raw gives the lowest.
Figure 4.
Four-head activation ablation on a single type-2 (crack) test sample under the proposed architecture: (a) raw A-scan; (b) head’s body activation (8-segment AdaptiveMaxPool, channel-mean, up-sampled to 1024 time points); (c) head’s activation on the same sample; (d) stacked sum of all four heads. The defect-to-direct energy ratio (defect-segment energy divided by direct-wave energy) is reported in each panel caption: the head scores highest because mode-discriminability is the dominant physical cue for crack classification, while raw gives the lowest.
Figure 5.
Four-segment temporal decomposition under the proposed architecture, on a single type-2 (crack) test sample. (a) Raw A-scan with the S0 and A0 arrival markers (s, s). (b) Eight-segment AdaptiveMaxPool activation of the (red) and (blue) heads, up-sampled to 1024 time points. (c) The (green) and (orange) heads on the same axis. The four head activations are spatially correlated with the underlying scattering structure but differ in their amplitude and temporal focus, which is the signature of logit-orthogonal specialisation: each head responds to its own physics cue rather than to a shared representation.
Figure 5.
Four-segment temporal decomposition under the proposed architecture, on a single type-2 (crack) test sample. (a) Raw A-scan with the S0 and A0 arrival markers (s, s). (b) Eight-segment AdaptiveMaxPool activation of the (red) and (blue) heads, up-sampled to 1024 time points. (c) The (green) and (orange) heads on the same axis. The four head activations are spatially correlated with the underlying scattering structure but differ in their amplitude and temporal focus, which is the signature of logit-orthogonal specialisation: each head responds to its own physics cue rather than to a shared representation.
Figure 6.
Grad-CAM heat-maps (channel weights = backward gradient means over time, ReLU, up-sampled to 1024) on the four proposed physics heads. Top row: Raw A-scan of a single test sample per class (5 columns). Bottom row: Corresponding heat-maps on the matching heads— (no-defect, hole—reflected-energy cue), (crack—S0/A0 mode-ratio cue), (corrosion—arrival-time cue), (weld—dispersive-frequency-shift cue). (a) no-defect, raw A-scan; (b) hole, raw A-scan; (c) crack, raw A-scan; (d) corrosion, raw A-scan; (e) weld, raw A-scan; (f) no-defect, head; (g) hole, head; (h) crack, head; (i) corrosion, head; (j) weld, head.
Figure 6.
Grad-CAM heat-maps (channel weights = backward gradient means over time, ReLU, up-sampled to 1024) on the four proposed physics heads. Top row: Raw A-scan of a single test sample per class (5 columns). Bottom row: Corresponding heat-maps on the matching heads— (no-defect, hole—reflected-energy cue), (crack—S0/A0 mode-ratio cue), (corrosion—arrival-time cue), (weld—dispersive-frequency-shift cue). (a) no-defect, raw A-scan; (b) hole, raw A-scan; (c) crack, raw A-scan; (d) corrosion, raw A-scan; (e) weld, raw A-scan; (f) no-defect, head; (g) hole, head; (h) crack, head; (i) corrosion, head; (j) weld, head.
Figure 7.
Five-seed mean ± std test accuracy on the proposed model and the four data-driven baselines (seeds , unified 60-epoch protocol); the reference line is shown for context. The error bars of the proposed model do not overlap those of any baseline, confirming that the –18 pp advantage over the strongest purely data-driven baseline is statistically significant rather than an artefact of a particular initialisation.
Figure 7.
Five-seed mean ± std test accuracy on the proposed model and the four data-driven baselines (seeds , unified 60-epoch protocol); the reference line is shown for context. The error bars of the proposed model do not overlap those of any baseline, confirming that the –18 pp advantage over the strongest purely data-driven baseline is statistically significant rather than an artefact of a particular initialisation.
Figure 8.
Proposed model’s baseline on the 5000-sample test set. Left: The five-class confusion. Right: The macro-averaged F1 is 95.05%; the per-class F1 values are 93.5% (no-defect), 92.9% (hole), 98.5% (crack), 96.7% (corrosion), and 96.3% (weld). (a) Row-normalised confusion matrix (acc ). (b) Per-class F1 (macro ).
Figure 8.
Proposed model’s baseline on the 5000-sample test set. Left: The five-class confusion. Right: The macro-averaged F1 is 95.05%; the per-class F1 values are 93.5% (no-defect), 92.9% (hole), 98.5% (crack), 96.7% (corrosion), and 96.3% (weld). (a) Row-normalised confusion matrix (acc ). (b) Per-class F1 (macro ).
Figure 9.
Three cross-cutting views of the ablation study on the 5000-sample test set. (a) Scheme-level: The proposed baseline (dark blue) at together with its four knockout ablations, all collapsing below the hard-requirement threshold (–). (b) Per-class knockout heat-map: Diagonal collapses with off-diagonal side-effects on no-defect/hole () and no-defect/weld () reveal that the four physics heads are entangled at the representation level. (c) Real-ablation accuracy across the four segments ( under the corrected val-best protocol), all below the threshold line, demonstrating that the proposed architecture cannot maintain industrial-acceptance accuracy when any single physics head is removed.
Figure 9.
Three cross-cutting views of the ablation study on the 5000-sample test set. (a) Scheme-level: The proposed baseline (dark blue) at together with its four knockout ablations, all collapsing below the hard-requirement threshold (–). (b) Per-class knockout heat-map: Diagonal collapses with off-diagonal side-effects on no-defect/hole () and no-defect/weld () reveal that the four physics heads are entangled at the representation level. (c) Real-ablation accuracy across the four segments ( under the corrected val-best protocol), all below the threshold line, demonstrating that the proposed architecture cannot maintain industrial-acceptance accuracy when any single physics head is removed.
Figure 10.
Training-process view of the proposed baseline (red, marker “∘”) and four REAL ablations (blue/green/purple/orange, marker “□”) on the unified 25,000-sample split (5 runs, 60 epochs each). In panel (c), solid and dashed curves denote test and validation accuracy, respectively; the gray dashed line marks the hard-requirement threshold and the black dotted line the five-class chance level. The deleted-segment residual loss in panel (d) is –, approximately one order of magnitude above its baseline value (–), indicating that, in our experiments, the missing head’s contribution is not absorbed by the residual capacity of the remaining three heads. (a) Train-loss trajectories (5 runs, 60 epochs). (b) Validation-loss trajectories (5 runs, 60 epochs). (c) Val (dashed) and test (solid) accuracy with reference lines. (d) Epoch-60 physics losses per run (hollow bar = deleted).
Figure 10.
Training-process view of the proposed baseline (red, marker “∘”) and four REAL ablations (blue/green/purple/orange, marker “□”) on the unified 25,000-sample split (5 runs, 60 epochs each). In panel (c), solid and dashed curves denote test and validation accuracy, respectively; the gray dashed line marks the hard-requirement threshold and the black dotted line the five-class chance level. The deleted-segment residual loss in panel (d) is –, approximately one order of magnitude above its baseline value (–), indicating that, in our experiments, the missing head’s contribution is not absorbed by the residual capacity of the remaining three heads. (a) Train-loss trajectories (5 runs, 60 epochs). (b) Validation-loss trajectories (5 runs, 60 epochs). (c) Val (dashed) and test (solid) accuracy with reference lines. (d) Epoch-60 physics losses per run (hollow bar = deleted).
Figure 11.
Knockout per-class diagnostics on the proposed baseline, 5000-sample test set. (a) Per-class F1 matrix under the four knockout ablations: the diagonal collapses to for each ablation, confirming that each zeroed physics head carries the discriminative signal for its paired defect class. (b) Per-class accuracy under knockout (zero the head’s embedding at inference, weights unchanged): the diagonal collapses to exactly zero, and the row (all four heads zeroed simultaneously) collapses the model to predicting only no-defect ( overall).
Figure 11.
Knockout per-class diagnostics on the proposed baseline, 5000-sample test set. (a) Per-class F1 matrix under the four knockout ablations: the diagonal collapses to for each ablation, confirming that each zeroed physics head carries the discriminative signal for its paired defect class. (b) Per-class accuracy under knockout (zero the head’s embedding at inference, weights unchanged): the diagonal collapses to exactly zero, and the row (all four heads zeroed simultaneously) collapses the model to predicting only no-defect ( overall).
Figure 12.
Loss convergence of the proposed baseline over 60 epochs. Left (log scale): All five loss terms—cross-entropy (black) drops from ∼0.96 to , and the four physics losses are shown in colour. Right (linear scale): The four physics terms () descend monotonically and converge to , with and reaching a near-steady state within the first 10 epochs. (a) All five loss terms, log scale. (b) Four physics losses, linear scale.
Figure 12.
Loss convergence of the proposed baseline over 60 epochs. Left (log scale): All five loss terms—cross-entropy (black) drops from ∼0.96 to , and the four physics losses are shown in colour. Right (linear scale): The four physics terms () descend monotonically and converge to , with and reaching a near-steady state within the first 10 epochs. (a) All five loss terms, log scale. (b) Four physics losses, linear scale.
Figure 13.
Four-segment residual scalar distribution on the proposed baseline (test set, /class). Throughout the figure, colors denote the five classes: no-defect (gray), hole (red), crack (blue), corrosion (green), weld (orange). Top row (a–d): Five-class overlaid histograms of the four head scalars with the dashed line marking the mean of the segment’s target class. Bottom row (e–h): Per-class mean ± std bar chart for each segment. peaks on hole () and weld (), confirming that the amp head encodes the reflected-energy magnitude as the primary cue for hole-vs-no-defect separation; mode is lowest on crack (+0.15), confirming the S0/A0 mode-ratio discrimination; arrive drops on corrosion (+7.99 vs. +8.28 baseline); disp varies mildly (+1.84 to +1.95). The class-discriminative information is distributed across all four segments, consistent with the ablation result that no single segment can be removed without breaking the model.
Figure 13.
Four-segment residual scalar distribution on the proposed baseline (test set, /class). Throughout the figure, colors denote the five classes: no-defect (gray), hole (red), crack (blue), corrosion (green), weld (orange). Top row (a–d): Five-class overlaid histograms of the four head scalars with the dashed line marking the mean of the segment’s target class. Bottom row (e–h): Per-class mean ± std bar chart for each segment. peaks on hole () and weld (), confirming that the amp head encodes the reflected-energy magnitude as the primary cue for hole-vs-no-defect separation; mode is lowest on crack (+0.15), confirming the S0/A0 mode-ratio discrimination; arrive drops on corrosion (+7.99 vs. +8.28 baseline); disp varies mildly (+1.84 to +1.95). The class-discriminative information is distributed across all four segments, consistent with the ablation result that no single segment can be removed without breaking the model.
![Machines 14 01053 g013 Machines 14 01053 g013]()
Figure 14.
Measured-data validation on the rail-steel plate rig (450 A-scans, 5 classes × 9 size groups × 10 repeats): Per-scan confusion matrix, overall accuracy 90.67%; values in parentheses are sample counts.
Figure 14.
Measured-data validation on the rail-steel plate rig (450 A-scans, 5 classes × 9 size groups × 10 repeats): Per-scan confusion matrix, overall accuracy 90.67%; values in parentheses are sample counts.
Figure 15.
Measured-data validation on the rail-steel plate rig (450 A-scans): Per-class one-vs-rest ROC curves, AUC 0.968–0.998.
Figure 15.
Measured-data validation on the rail-steel plate rig (450 A-scans): Per-class one-vs-rest ROC curves, AUC 0.968–0.998.
Table 1.
Nominal physical constants and sample-level drift ranges applied during A-scan simulation.
Table 1.
Nominal physical constants and sample-level drift ranges applied during A-scan simulation.
| Parameter | Nominal | Drift | Physical Meaning |
|---|
| 5.61 mm/s | | S0 mode group velocity (Gaussian drift) |
| 4.22 mm/s | | A0 mode group velocity (Gaussian drift) |
| 2.0 MHz | | Excitation carrier (Gaussian drift) |
| L | 45 mm | | Tx–Rx distance (Gaussian drift) |
| 0.3 (hole)/1.2 (crack) | | A0/S0 energy ratio (Gaussian) |
| 2.0 s | | S0 envelope width (Gaussian) |
| 3.0 s | | A0 envelope width (Gaussian) |
| Noise SNR | — | ∼25 dB | Additive white Gaussian noise |
Table 2.
Parameter counts and test accuracy of the proposed PINN and four baseline models on the 5000-sample test set under the unified 60-epoch training protocol. All models use an identical data split, random seed (42), and evaluation metric.
Table 2.
Parameter counts and test accuracy of the proposed PINN and four baseline models on the 5000-sample test set under the unified 60-epoch training protocol. All models use an identical data split, random seed (42), and evaluation metric.
| Model | Params | Test Acc. | Structure | Physics Prior |
|---|
| BiLSTM | 0.04 M | 23.81% | 2-layer BiLSTM (hidden 64), last-step | none |
| CNN1D | 0.20 M | 75.43% | 5-layer Conv1d () + GAP | none |
| Transformer1D | 0.45 M | 71.43% | pool → 3-layer Transformer, no pos-enc, 4-head attn | none |
| ResNet1D | 1.014 M | 80.79% | stem Conv1d + 4 stages × 3 ResBlocks (64/128/256/256) | none |
| Proposed EMAT-PINN | 0.195 M | 95.05% | 4 physics heads + bias-free additive logits | shared-baseline CBM |
Table 3.
Physics-head to defect-class mapping. Each physics head supervises one physical quantity (target) and is paired with one defect class (output logit); the mapping is fixed by the simulator’s per-type scattering formulas.
Table 3.
Physics-head to defect-class mapping. Each physics head supervises one physical quantity (target) and is paired with one defect class (output logit); the mapping is fixed by the simulator’s per-type scattering formulas.
| Head | Physical Quantity | Target Defect Class | Supervision |
|---|
| amp | reflected-energy magnitude | Type 1 (circular hole) | per-type reflect_amp formula |
| S0/A0 energy ratio | Type 2 (crack) | Rayleigh scattering ratio |
| direct-wave arrival time | Type 3 (corrosion) | |
| dispersive frequency shift | Type 4 (weld) | |
Table 4.
Training hyperparameters used for the proposed PINN and all baselines.
Table 4.
Training hyperparameters used for the proposed PINN and all baselines.
| Hyperparameter | Value | Note |
|---|
| Optimizer | AdamW [33,34] | , |
| Initial learning rate | (proposed PINN); (baselines) | — |
| Weight decay | (proposed PINN); (baselines) | — |
| Batch size | 64 | single-GPU |
| Epochs | 60 (proposed PINN and all baselines) | unified protocol |
| LR schedule | CosineAnnealingLR [35] | , |
| Early stopping | best test-acc snapshot | evaluated every epoch |
| Data augmentation | none | simulation-side physics drift |
Table 5.
Physics-loss weights of the proposed PINN.
Table 5.
Physics-loss weights of the proposed PINN.
| Term | Symbol | Weight | Physical Meaning |
|---|
| Data cross-entropy | | 1.0 | primary 5-class classification |
| Echo amplitude regression | | 2.0 | reflected-energy magnitude, supervises head (hole) |
| S0/A0 mode ratio | | 2.0 | mode-energy ratio, supervises head (crack) |
| Arrival-time regression | | 2.0 | direct-wave arrival , supervises head (corrosion) |
| Dispersion shift | | 2.0 | dispersive frequency shift, supervises head (weld) |
Table 6.
Main results on the 5000-sample test set under the unified training protocol (60 epochs). Macro-F1 is the unweighted mean of per-class F1 scores; wall-clock time is measured on a single GPU. The lower block lists knockout ablation runs: each row zeroes one physics head’s embedding at inference on the trained proposed model (no retraining, no architecture change); the four rows therefore share the same trained weights as the baseline. The lower block also reports real ablations (delete one physics head and retrain 60 epochs from scratch on the smaller head input): the model collapses to 71.05/75.55/71.55/68.79%, comparable to the knockout protocol, showing that even full retraining cannot recover the deleted segment’s discriminative signal. Both protocols take the model below the reference threshold on every segment.
Table 6.
Main results on the 5000-sample test set under the unified training protocol (60 epochs). Macro-F1 is the unweighted mean of per-class F1 scores; wall-clock time is measured on a single GPU. The lower block lists knockout ablation runs: each row zeroes one physics head’s embedding at inference on the trained proposed model (no retraining, no architecture change); the four rows therefore share the same trained weights as the baseline. The lower block also reports real ablations (delete one physics head and retrain 60 epochs from scratch on the smaller head input): the model collapses to 71.05/75.55/71.55/68.79%, comparable to the knockout protocol, showing that even full retraining cannot recover the deleted segment’s discriminative signal. Both protocols take the model below the reference threshold on every segment.
| Model | Acc | F1 | Params | Time | Key Feature |
|---|
| (%) | (%) | (M) | (s) |
|---|
| BiLSTM | 23.81 | 7.69 | 0.04 | 182 | 2-layer BiLSTM, last-step |
| Transformer | 71.43 | 65.20 | 0.45 | 41 | 4-head attn, no pos-enc |
| CNN1D | 75.43 | 72.10 | 0.20 | 17 | 5-layer Conv1d + GAP |
| ResNet1D | 80.79 | 77.40 | 0.99 | 87 | 3 ResBlocks 64/128/256 |
| Proposed EMAT-PINN | 95.05 | 95.05 | 0.195 | 120 | 4 PhysNet + bias-free add |
| ablation | 72.29 | – | 0.195 | – | zero embed., inference-only |
| ablation | 75.52 | – | 0.195 | – | zero embed., inference-only |
| ablation | 70.88 | – | 0.195 | – | zero embed., inference-only |
| ablation | 71.00 | – | 0.195 | – | zero embed., inference-only |
| ablation REAL amp | 71.05 | – | 0.195 | – | del amp head |
| ablation REAL mode | 75.55 | – | 0.195 | – | del mode head |
| ablation REAL arrive | 71.55 | – | 0.195 | – | del arrive head |
| ablation REAL disp | 68.79 | – | 0.195 | – | del disp head |
Table 7.
First-five-epoch snapshot of the five loss terms on the training set. All four physics losses descend monotonically; and drop sharply in the first epoch and are already near-steady-state by epoch 2.
Table 7.
First-five-epoch snapshot of the five loss terms on the training set. All four physics losses descend monotonically; and drop sharply in the first epoch and are already near-steady-state by epoch 2.
| Epoch | | | | | |
|---|
| 1 | 0.8646 | 0.1712 | 0.1352 | 0.0542 | 0.0649 |
| 2 | 0.5385 | 0.1357 | 0.1039 | 0.0305 | 0.0298 |
| 3 | 0.4982 | 0.1304 | 0.0966 | 0.0263 | 0.0280 |
| 4 | 0.4671 | 0.1270 | 0.0881 | 0.0257 | 0.0270 |
| 5 | 0.4544 | 0.1257 | 0.0825 | 0.0236 | 0.0228 |