A Conformer-Based Time–Frequency Decoupling Network for Pig Vocalization Behavior Classification
Simple Summary
Abstract
1. Introduction
2. Materials and Methods
2.1. Dataset
2.1.1. Data Collection
2.1.2. Data Characterization and Feature Analysis
2.1.3. Data Preprocessing and Noise Reduction
2.1.4. Data Partitioning and Evaluation Protocols
2.2. Method
2.2.1. ATF-Conformer Framework
2.2.2. Front-End Feature Enhancement Module
2.2.3. TF-Decoupled Conformer Encoder
2.2.4. Sequence Aggregation Module
2.2.5. Classification Head with Discriminative Loss
3. Results
3.1. Experimental Settings
3.1.1. Baseline Models
3.1.2. Experimental Environment and Hyperparameters
3.1.3. Evaluation Metrics
- Accuracy
- 2.
- Precision
- 3.
- Recall
- 4.
- F1-score
- 5.
- AUROC
- 6.
- CM
3.2. Analysis of Training Accuracy and Loss
3.3. Ablation Studies
3.4. Performance Comparison with Baseline Models
3.5. Error Structure and Confusion Analysis
3.6. Per-Class Performance Analysis
- (1)
- Cough: ATF-Conformer achieves the highest F1-score and AUROC for cough, indicating strong sensitivity to brief respiratory events under noisy barn conditions. This result suggests improved discrimination of short transient cues relative to the baseline models.
- (2)
- Scream: For scream, ATF-Conformer achieves performance comparable to or better than the strongest baseline models and attains the highest AUROC. This indicates stable recognition of high-energy abnormal vocalizations across decision thresholds.
- (3)
- Estrus: ATF-Conformer shows the best overall performance for estrus recognition. The improvement is particularly relevant for this relatively weak-energy and rhythmical class under complex farm acoustic conditions.
- (4)
- Feeding: For feeding sounds, ATF-Conformer achieves the highest overall F1-score and AUROC while maintaining a balanced precision–recall trade-off. This suggests improved robustness for a category that overlaps strongly with low-frequency barn background activity.
- (5)
- Normal: ATF-Conformer also performs best overall for the normal class, with the highest F1-score and AUROC. This result indicates more stable discrimination between routine barn acoustics and behavior-related vocal events.
3.7. Robustness and Cross-Session Generalization
3.7.1. Robustness Evaluation Under Additive Noise Perturbations
3.7.2. Cross-Session Generalization Experiment
4. Discussion
4.1. Behavioral Vocalization Characteristics and Error Patterns in Commercial Pig Barns
4.2. Robustness of Time–Frequency Modeling Under Noisy Farm Conditions
4.3. Implications for Practical Deployment in Precision Livestock Farming
4.4. Limitations and Future Perspectives
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AAM-Softmax | Additive Angular Margin Softmax |
| ATF-Conformer | Attention-Guided Time-Frequency Decoupled Conformer |
| AUROC | Area Under the Receiver Operating Characteristic Curve |
| CM | Confusion Matrix |
| Log-Mel | logarithmic Mel-frequency |
| LOSO | Leave-one-session-out |
| OOF | out-of-fold |
| OvR | one-vs-rest |
| STFT | short-time Fourier transform |
| TA | Triplet Attention |
| TF | Time-Frequency |
| TFD-Conformer | Time-Frequency Decoupled Conformer |
References
- Liao, J.; Li, H.; Feng, A.; Wu, X.; Luo, Y.; Duan, X.; Ni, M.; Li, J. Domestic pig sound classification based on TransformerCNN. Appl. Intell. 2023, 53, 4907–4923. [Google Scholar] [CrossRef]
- Chung, Y.; Oh, S.; Lee, J.; Park, D.; Chang, H.-H.; Kim, S. Automatic detection and recognition of pig wasting diseases using sound data in audio surveillance systems. Sensors 2013, 13, 12929–12942. [Google Scholar] [CrossRef] [PubMed]
- Lagua, E.B.; Mun, H.-S.; Ampode, K.M.B.; Chem, V.; Kim, Y.-H.; Yang, C.-J. Artificial intelligence for automatic monitoring of respiratory health conditions in smart swine farming. Animals 2023, 13, 1860. [Google Scholar] [CrossRef] [PubMed]
- Heseker, P.; Bergmann, T.; Scheumann, M.; Traulsen, I.; Kemper, N.; Probst, J. Detecting tail biters by monitoring pig screams in weaning pigs. Sci. Rep. 2024, 14, 4523. [Google Scholar] [CrossRef]
- Xie, Y.; Wang, J.; Chen, C.; Yin, T.; Yang, S.; Li, Z.; Zhang, Y.; Ke, J.; Song, L.; Gan, L. Sound identification of abnormal pig vocalizations: Enhancing livestock welfare monitoring on smart farms. Inf. Process. Manag. 2024, 61, 103770. [Google Scholar] [CrossRef]
- Benjamin, M.; Yik, S. Precision livestock farming in swine welfare: A review for swine practitioners. Animals 2019, 9, 133. [Google Scholar] [CrossRef]
- Dawkins, M.S. Smart farming and Artificial Intelligence (AI): How can we ensure that animal welfare is a priority? Appl. Anim. Behav. Sci. 2025, 283, 106519. [Google Scholar] [CrossRef]
- Nan, J.; Yin, Y.; Sun, W.; Zhang, Y. Pig sound analysis: A measure of welfare. Smart Agric. 2022, 4, 19–35. [Google Scholar] [CrossRef]
- Pu, P.; Wang, J.; Yan, G.; Jiao, H.; Li, H.; Lin, H. EnhancedMulti-Scenario Pig Behavior Recognition Based on YOLOv8n. Animals 2025, 15, 2927. [Google Scholar] [CrossRef]
- Shao, X.; Liu, C.; Zhou, Z.; Xue, W.; Zhang, G.; Liu, J.; Yan, H. Research on dynamic pig counting method based on improved YOLOv7 combined with DeepSORT. Animals 2024, 14, 1227. [Google Scholar] [CrossRef]
- Yang, Q.; Hui, X.; Huang, Y.; Chen, M.; Huang, S.; Xiao, D. A Long-Term Video Tracking Method for Group-Housed Pigs. Animals 2024, 14, 1505. [Google Scholar] [CrossRef]
- Ma, C.; Deng, M.; Yin, Y. Pig face recognition based on improved YOLOv4 lightweight neural network. Inf. Process. Agric. 2024, 11, 356–371. [Google Scholar] [CrossRef]
- Lei, K.; Zong, C.; Du, X.; Teng, G.; Feng, F. Oestrus analysis of sows based on bionic boars and machine vision technology. Animals 2021, 11, 1485. [Google Scholar] [CrossRef]
- Lv, Y.; Liu, Y.; Song, Y.; Wang, J.; Li, Q. Recognising Behaviorally Relevant Pig Vocalizations for Welfare Assessment via a Lightweight Deep Acoustic Model. Appl. Anim. Behav. Sci. 2026, 298, 106936. [Google Scholar] [CrossRef]
- Cai, J.; Liu, W.; Liu, T.; Wang, F.; Li, Z.; Wang, X.; Li, H. APO-CViT: A Non-Destructive Estrus Detection Method for Breeding Pigs Based on Multimodal Feature Fusion. Animals 2025, 15, 1067. [Google Scholar] [CrossRef]
- Olczak, K.; Penar, W.; Nowicki, J.; Magiera, A.; Klocek, C. The role of sound in livestock farming—Selected aspects. Animals 2023, 13, 2307. [Google Scholar] [CrossRef] [PubMed]
- Sharifuzzaman, M.; Mun, H.-S.; Ampode, K.M.B.; Lagua, E.B.; Park, H.-R.; Kim, Y.-H.; Hasan, K.; Yang, C.-J. Technological tools and artificial intelligence in estrus detection of sows—A comprehensive review. Animals 2024, 14, 471. [Google Scholar] [CrossRef]
- Yin, Y.; Tu, D.; Shen, W.; Bao, J. Recognition of sick pig cough sounds based on convolutional neural network in field situations. Inf. Process. Agric. 2021, 8, 369–379. [Google Scholar] [CrossRef]
- Shen, W.; Tu, D.; Yin, Y.; Bao, J. A new fusion feature based on convolutional neural network for pig cough recognition in field situations. Inf. Process. Agric. 2021, 8, 573–580. [Google Scholar] [CrossRef]
- Wang, B.; Qi, J.; An, X.; Wang, Y. Heterogeneous fusion of biometric and deep physiological features for accurate porcine cough recognition. PLoS ONE 2024, 19, e0297655. [Google Scholar] [CrossRef] [PubMed]
- von Borell, E.; Bünger, B.; Schmidt, T.; Horn, T. Vocal-type classification as a tool to identify stress in piglets under on-farm conditions. Anim. Welf. 2009, 18, 407–416. [Google Scholar] [CrossRef]
- Niño, J.N.R.; de Sousa, F.C.; Oliveira, C.E.A.; Coelho, A.L.d.F.; Hernandez, R.O.; Barbari, M. Systematic Review of Acoustic Monitoring in Livestock Farming: Vocalization Patterns and Sound Source Analysis. Appl. Sci. 2025, 15, 9910. [Google Scholar] [CrossRef]
- Yin, Y.; Ji, N.; Wang, X.; Shen, W.; Dai, B.; Kou, S.; Liang, C. An investigation of fusion strategies for boosting pig cough sound recognition. Comput. Electron. Agric. 2023, 205, 107645. [Google Scholar] [CrossRef]
- Shen, W.; Wang, X.; Yin, Y.; Ji, N.; Dai, B.; Kou, S.; Liang, C. Dempster Shafer distance-based multi-classifier fusion method for pig cough recognition. Int. J. Agric. Biol. Eng. 2024, 17, 245–254. [Google Scholar] [CrossRef]
- Wang, X.; Yin, Y.; Dai, X.; Shen, W.; Kou, S.; Dai, B. Automatic detection of continuous pig cough in a complex piggery environment. Biosyst. Eng. 2024, 238, 78–88. [Google Scholar] [CrossRef]
- Wu, X.; Zhou, S.; Chen, M.; Zhao, Y.; Wang, Y.; Zhao, X.; Li, D.; Pu, H. Combined spectral and speech features for pig speech recognition. PLoS ONE 2022, 17, e0276778. [Google Scholar] [CrossRef]
- Ginovart-Panisello, G.J.; Alsina-Pagès, R.M.; Sanz, I.I.; Monjo, T.P.; Prat, M.C. Acoustic description of the soundscape of a real-life intensive farm and its impact on animal welfare: A preliminary analysis of farm sounds and bird vocalizations. Sensors 2020, 20, 4732. [Google Scholar] [CrossRef]
- Misra, D.; Nalamada, T.; Arasanipalai, A.U.; Hou, Q. Rotate to attend: Convolutional triplet attention module. In Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2021; pp. 3139–3148. [Google Scholar] [CrossRef]
- Gulati, A.; Qin, J.; Chiu, C.C.; Parmar, N.; Zhang, Y.; Yu, J.; Han, W.; Wang, S.; Zhang, Z.; Wu, Y.; et al. Conformer: Convolution-augmented transformer for speech recognition. arXiv 2020, arXiv:2005.08100. [Google Scholar] [CrossRef]
- Ma, J.; Li, Z.; Wang, H.; Yang, X.; Xu, D.; Wang, C. On field recognition of pig cough based on multimodal audio representation and fusion with contrastive learning. Comput. Electron. Agric. 2025, 238, 110842. [Google Scholar] [CrossRef]
- Jannu, C.; Burra, M.; Vanambathina, S.D.; Parisae, V.; Krishna, C.V.M.; Madhumati, G.L. Single Channel Speech Enhancement using a Complex Dual-Path Multi Axial Transformer with Frequency Prompt. Circuits Syst. Signal Process. 2025, 44, 4224–4257. [Google Scholar] [CrossRef]
- Vandermeulen, J.; Bahr, C.; Tullo, E.; Fontana, I.; Ott, S.; Kashiha, M.; Guarino, M.; Moons, C.P.H.; Tuyttens, F.A.M.; Niewold, T.A.; et al. Discerning pig screams in production environments. PLoS ONE 2015, 10, e0123111. [Google Scholar] [CrossRef]
- Cakir, E.; Parascandolo, G.; Heittola, T.; Huttunen, H.; Virtanen, T. Convolutional recurrent neural networks for polyphonic sound event detection. IEEE/ACM Trans. Audio Speech Lang. Process. 2017, 25, 1291–1303. [Google Scholar] [CrossRef]
- Kong, Q.; Cao, Y.; Iqbal, T.; Wang, Y.; Wang, W.; Plumbley, M.D. Panns: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Trans. Audio Speech Lang. Process. 2020, 28, 2880–2894. [Google Scholar] [CrossRef]
- Chen, K.; Du, X.; Zhu, B.; Ma, Z.; Berg-Kirkpatrick, T.; Dubnov, S. HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection. In ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2022; pp. 646–650. [Google Scholar] [CrossRef]
- Zhang, J.; Wang, L.; Yu, Y.; Xu, M. ECMISM: Speech recognition via enhancing conformer models with innovative scoring matrices. In Pattern Recognition; Springer: Cham, Switzerland, 2024; pp. 335–350. [Google Scholar] [CrossRef]
- Pan, W.; Li, H.; Zhou, X.; Jiao, J.; Zhu, C.; Zhang, Q. Research on pig sound recognition based on deep neural network and hidden Markov models. Sensors 2024, 24, 1269. [Google Scholar] [CrossRef]
- Chae, H.; Lee, J.; Kim, J.; Lee, S.; Lee, J.; Chung, Y.; Park, D. Novel Method for Detecting Coughing Pigs with Audio-Visual Multimodality for Smart Agriculture Monitoring. Sensors 2024, 24, 7232. [Google Scholar] [CrossRef]
- Chen, C.; Zhu, W.; Norton, T. Behavior recognition of pigs and cattle: Journey from computer vision to deep learning. Comput. Electron. Agric. 2021, 187, 106255. [Google Scholar] [CrossRef]
- Yang, Y.; Xu, C.; Hou, W.; McElligott, A.G.; Liu, K.; Xue, Y. Transformer-based audio-visual multimodal fusion for fine-grained recognition of individual sow nursing behavior. Artif. Intell. Agric. 2025, 15, 363–376. [Google Scholar] [CrossRef]
- Pann, V.; Kwon, K.-S.; Kim, B.; Jang, D.-H.; Kim, J.; Kim, J.-B. Robustness of CNN-Based model assessment for pig vocalization classification across diverse acoustic environments. Comput. Electron. Agric. 2026, 240, 111181. [Google Scholar] [CrossRef]
- Shaikh, M.B.; Chai, D.; Islam, S.M.S.; Akhtar, N. Multimodal fusion for audio-image and video action recognition. Neural Comput. Appl. 2024, 36, 5499–5513. [Google Scholar] [CrossRef]









| Behavior Category | Sample Count | Average Duration (Seconds) | Frequency Range (kHz) |
|---|---|---|---|
| Cough | 1053 | 1.82 | 0.8–3.0 |
| Scream | 1092 | 1.75 | 3.0–8.0 |
| Estrus | 1076 | 1.94 | 1.0–3.5 |
| Feeding | 1029 | 1.88 | 0.2–2.5 |
| Normal | 1018 | 1.80 | 0–1.5 |
| Vocalization Class | Routine Source | Supplementary Source | Close-Range Sampling | Routine Split Unit | Additional Split Rule |
|---|---|---|---|---|---|
| Feeding | Finishing pens | – | – | Pen × date × session | – |
| Normal | Finishing pens | – | – | Pen × date × session | – |
| Scream | Finishing pens | – | – | Pen × date × session | – |
| Estrus | Breeding stalls | – | – | Session-level mutual exclusion | – |
| Cough | Finishing pens | Symptomatic pigs | Yes | Pen × date × session | Pig-wise exclusion * |
| Vocalization Class | Train | Validation | Test | Total |
|---|---|---|---|---|
| Cough | 839 | 108 | 106 | 1053 |
| Scream | 872 | 110 | 110 | 1092 |
| Estrus | 858 | 110 | 108 | 1076 |
| Feeding | 821 | 105 | 103 | 1029 |
| Normal | 812 | 104 | 102 | 1018 |
| Total | 4202 | 537 | 529 | 5268 |
| Network | Component | Configuration/Value |
|---|---|---|
| ATF-Conformer | Front-end encoding | 3 × 3 2D Conv + BN + ReLU; 2 × 2 AvgPool |
| Attention module | Triplet Attention | |
| Core encoder | TFD-Conformer × 2 | |
| Aggregation/classifier | Masked temporal pooling + AAM-Softmax | |
| ECMISM | Convolutional subsampling | Two-layer downsampling convolution |
| Core encoder | Conformer × 2 | |
| Auxiliary design | Skip fusion + InLoss | |
| Output strategy | Global pooling + 5-class output | |
| Conformer | Convolutional subsampling | Two-layer 2D Conv |
| Projection | Linear projection | |
| Core encoder | Conformer × 2 | |
| Aggregation | Global average pooling | |
| HTS-AT | Patch embedding | 4 × 4 Conv, stride 4 × 4 |
| Hierarchical encoder | Swin Transformer (2/2/6/2) | |
| Token processing | Patch merging + token-semantic CNN | |
| Aggregation | Global average pooling | |
| PANNs | Convolutional extraction | Stacked 3 × 3 Conv (64 → 128 → 256 → 512) |
| Pooling | Progressive 2 × 2 average pooling | |
| Aggregation | GAP + GMP | |
| Classifier | FC 2048 | |
| CRNN | Convolutional extraction | 3 × 3 Conv (64 → 128) |
| Pooling | MaxPool (2 × 2, 3 × 3) | |
| Temporal modeling | LSTM (64) | |
| Classifier | FC 128 |
| Model | Accuracy (%) | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|
| Full Model | 97.38 | 97.86 | 97.45 | 97.65 |
| w/o Spectral Gating | 94.53 | 94.81 | 94.18 | 94.49 |
| w/o TA | 94.76 | 95.09 | 94.32 | 94.70 |
| w/o TFD-Conformer | 92.08 | 92.67 | 91.83 | 92.24 |
| Conv-Front Only | 88.57 | 89.29 | 88.04 | 88.62 |
| Model | Accuracy ± SD (%) | Macro-Precision ± SD (%) | Macro-Recall ± SD (%) | Macro-F1 ± SD (%) | Macro-AUROC ± SD | p-Value vs. ATF |
|---|---|---|---|---|---|---|
| ATF-Conformer | 97.34 ± 0.42 | 97.82 ± 0.23 | 97.41 ± 0.39 | 97.61 ± 0.28 | 0.97 ± 0.01 | – |
| ECMISM | 96.38 ± 0.53 | 97.01 ± 0.38 | 96.12 ± 0.58 | 96.52 ± 0.47 | 0.96 ± 0.01 | 0.008 |
| Conformer | 95.07 ± 0.72 | 95.42 ± 0.68 | 95.31 ± 0.71 | 95.39 ± 0.69 | 0.95 ± 0.01 | <0.001 |
| HTS-AT | 95.66 ± 0.61 | 97.21 ± 0.47 | 94.23 ± 0.82 | 95.68 ± 0.59 | 0.95 ± 0.01 | <0.001 |
| PANNs | 93.95 ± 0.79 | 93.34 ± 0.88 | 94.62 ± 0.81 | 94.02 ± 0.77 | 0.95 ± 0.01 | <0.001 |
| CRNN | 92.61 ± 0.98 | 93.83 ± 0.92 | 91.94 ± 1.08 | 92.79 ± 0.96 | 0.93 ± 0.01 | <0.001 |
| Vocalization Class | Model | Precision ± SD (%) | Recall ± SD (%) | F1-Score ± SD (%) | AUROC ± SD |
|---|---|---|---|---|---|
| Cough | ATF-Conformer | 97.52 ± 0.83 | 98.21 ± 0.74 | 97.86 ± 0.72 | 0.98 ± 0.01 |
| ECMISM | 97.08 ± 0.54 | 96.48 ± 0.62 | 96.78 ± 0.56 | 0.97 ± 0.01 | |
| Conformer | 95.94 ± 0.72 | 95.36 ± 0.81 | 95.65 ± 0.74 | 0.96 ± 0.02 | |
| HTS-AT | 96.41 ± 0.69 | 94.32 ± 0.88 | 95.35 ± 0.80 | 0.96 ± 0.02 | |
| PANNs | 93.26 ± 0.91 | 94.74 ± 0.96 | 93.99 ± 0.90 | 0.95 ± 0.02 | |
| CRNN | 92.96 ± 0.98 | 91.42 ± 1.07 | 92.18 ± 0.99 | 0.94 ± 0.03 | |
| Scream | ATF-Conformer | 98.34 ± 0.67 | 97.48 ± 0.82 | 97.91 ± 0.69 | 0.97 ± 0.02 |
| ECMISM | 97.62 ± 0.61 | 96.21 ± 0.71 | 96.91 ± 0.63 | 0.96 ± 0.02 | |
| Conformer | 96.26 ± 0.74 | 95.18 ± 0.83 | 95.72 ± 0.75 | 0.95 ± 0.02 | |
| HTS-AT | 99.12 ± 0.53 | 96.87 ± 0.79 | 97.97 ± 0.62 | 0.96 ± 0.02 | |
| PANNs | 93.47 ± 0.89 | 95.63 ± 0.91 | 94.54 ± 0.83 | 0.94 ± 0.03 | |
| CRNN | 93.52 ± 0.93 | 92.34 ± 1.02 | 92.92 ± 0.95 | 0.93 ± 0.03 | |
| Estrus | ATF-Conformer | 96.78 ± 1.12 | 94.47 ± 1.06 | 95.63 ± 0.94 | 0.96 ± 0.02 |
| ECMISM | 95.92 ± 0.76 | 93.62 ± 0.84 | 94.76 ± 0.75 | 0.95 ± 0.02 | |
| Conformer | 93.86 ± 0.83 | 94.08 ± 0.92 | 93.97 ± 0.84 | 0.94 ± 0.02 | |
| HTS-AT | 94.18 ± 0.92 | 92.51 ± 1.01 | 93.33 ± 0.91 | 0.94 ± 0.02 | |
| PANNs | 90.74 ± 1.09 | 93.36 ± 1.02 | 91.92 ± 1.01 | 0.93 ± 0.03 | |
| CRNN | 93.41 ± 1.02 | 91.96 ± 1.09 | 92.64 ± 1.01 | 0.93 ± 0.03 | |
| Feeding | ATF-Conformer | 96.53 ± 0.88 | 95.81 ± 0.97 | 96.16 ± 0.91 | 0.96 ± 0.01 |
| ECMISM | 96.18 ± 0.74 | 95.46 ± 0.83 | 95.82 ± 0.75 | 0.95 ± 0.02 | |
| Conformer | 94.64 ± 0.81 | 94.92 ± 0.88 | 94.78 ± 0.82 | 0.94 ± 0.02 | |
| HTS-AT | 94.82 ± 0.93 | 92.94 ± 1.01 | 93.87 ± 0.92 | 0.93 ± 0.02 | |
| PANNs | 94.06 ± 0.92 | 96.02 ± 0.91 | 95.03 ± 0.84 | 0.95 ± 0.02 | |
| CRNN | 93.26 ± 1.01 | 92.18 ± 1.08 | 92.71 ± 1.02 | 0.92 ± 0.03 | |
| Normal | ATF-Conformer | 97.69 ± 0.54 | 98.97 ± 0.43 | 98.33 ± 0.36 | 0.99 ± 0.01 |
| ECMISM | 97.84 ± 0.52 | 98.12 ± 0.51 | 97.98 ± 0.45 | 0.98 ± 0.01 | |
| Conformer | 95.62 ± 0.73 | 96.18 ± 0.71 | 95.90 ± 0.64 | 0.96 ± 0.02 | |
| HTS-AT | 95.46 ± 0.84 | 94.78 ± 0.92 | 95.12 ± 0.76 | 0.96 ± 0.02 | |
| PANNs | 96.68 ± 0.74 | 95.92 ± 0.83 | 96.30 ± 0.72 | 0.96 ± 0.01 | |
| CRNN | 94.52 ± 0.94 | 93.68 ± 1.01 | 94.09 ± 0.93 | 0.95 ± 0.03 |
| Model | Test Morning | Test Noon | Test Evening | Mean ± SD | ΔNoon |
|---|---|---|---|---|---|
| ATF-Conformer | 96.52 | 96.03 | 96.79 | 96.45 ± 0.39 | 0.62 |
| ECMISM | 95.63 | 94.58 | 95.87 | 95.36 ± 0.69 | 1.17 |
| HTS-AT | 95.02 | 94.23 | 95.47 | 94.91 ± 0.63 | 1.02 |
| Conformer | 94.18 | 93.41 | 94.57 | 94.05 ± 0.59 | 0.97 |
| PANNs | 92.83 | 91.79 | 93.07 | 92.56 ± 0.68 | 1.16 |
| CRNN | 91.37 | 90.63 | 91.88 | 91.29 ± 0.63 | 1.00 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, J.; Liu, Y.; Geng, S.; Wei, F.; Wu, H.; Song, Y.; Lv, Y.; Li, S.; Li, Q. A Conformer-Based Time–Frequency Decoupling Network for Pig Vocalization Behavior Classification. Animals 2026, 16, 1337. https://doi.org/10.3390/ani16091337
Wang J, Liu Y, Geng S, Wei F, Wu H, Song Y, Lv Y, Li S, Li Q. A Conformer-Based Time–Frequency Decoupling Network for Pig Vocalization Behavior Classification. Animals. 2026; 16(9):1337. https://doi.org/10.3390/ani16091337
Chicago/Turabian StyleWang, Jianping, Yuqing Liu, Siao Geng, Feng Wei, Haoyu Wu, Yuzhen Song, Yingying Lv, Shugang Li, and Qian Li. 2026. "A Conformer-Based Time–Frequency Decoupling Network for Pig Vocalization Behavior Classification" Animals 16, no. 9: 1337. https://doi.org/10.3390/ani16091337
APA StyleWang, J., Liu, Y., Geng, S., Wei, F., Wu, H., Song, Y., Lv, Y., Li, S., & Li, Q. (2026). A Conformer-Based Time–Frequency Decoupling Network for Pig Vocalization Behavior Classification. Animals, 16(9), 1337. https://doi.org/10.3390/ani16091337

