A Lightweight 1D-CNN for Bark and Howl Classification from Raw Audio Waveforms Under Controlled Additive Noise
Featured Application
Abstract
1. Introduction
- A 7130-parameter 1D-CNN is evaluated for Bark and Howl classification directly from 2.0 s raw waveform segments.
- The full five-fold cross-validation procedure is repeated ten times to quantify variability across dataset partitions and stochastic training runs.
- Controlled Gaussian and uniform additive-noise performance is evaluated from 30 to 0 dB SNR, including an intermediate 15 dB condition that was deliberately excluded from training.
- A seed-42 reference ablation study evaluates noise augmentation, batch normalization, dropout, and 1 s, 2 s, and 3 s input durations.
- Parameter count, model size, FLOPs, inference time, and contextual comparisons with feature-based and compact audio architectures are reported.
2. Materials and Methods
2.1. Dataset
2.2. Raw Waveform Standardization
2.3. Controlled Additive-Noise Augmentation
2.4. Proposed 1D-CNN Architecture
2.5. Training Configuration
2.6. Repeated Five-Fold Cross-Validation
2.7. Ablation Analysis
2.8. Evaluation Metrics
2.9. Model Size, FLOPs, and Runtime Analysis
3. Results
3.1. Repeated No-Added-Noise Performance
3.2. Controlled Additive-Noise Performance
3.3. Seed-42 Ablation Results
3.4. Computational Results
3.5. Contextual Comparison with Published Architectures
4. Discussion
4.1. Classification Performance and Statistical Stability
4.2. Controlled Additive Noise and Ablation Findings
4.3. Computational Efficiency and Relation to Compact Audio Models
4.4. Interpretation of Literature Comparisons
4.5. Limitations and Future Work
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| 1D-CNN | One-dimensional convolutional neural network |
| ACC | Accuracy |
| BN | Batch normalization |
| CI | Confidence interval |
| CNN | Convolutional neural network |
| CPU | Central processing unit |
| CV | Cross-validation |
| FC | Fully connected |
| FLOPs | Floating-point operations |
| GPU | Graphics processing unit |
| LFCC | Linear-frequency cepstral coefficient |
| MFCC | Mel-frequency cepstral coefficient |
| PRE | Precision |
| ReLU | Rectified linear unit |
| SD | Standard deviation |
| SEN | Sensitivity |
| SNR | Signal-to-noise ratio |
| STFT | Short-Time Fourier Transform |
References
- Karaaslan, M.; Turkoglu, B.; Kaya, E.; Asuroglu, T. Voice analysis in dogs with deep learning: Development of a fully automatic voice analysis system for bioacoustics studies. Sensors 2024, 24, 7978. [Google Scholar] [CrossRef] [PubMed]
- Ovaskainen, O.; de Camargo, U.M.; Somervuo, P. Animal Sound Identifier (ASI): Software for automated identification of vocal animals. Ecol. Lett. 2018, 21, 1244–1254. [Google Scholar] [CrossRef] [PubMed]
- Penar, W.; Magiera, A.; Klocek, C. Applications of bioacoustics in animal ecology. Ecol. Complex. 2020, 43, 100847. [Google Scholar] [CrossRef]
- Siegford, J.M.; Steibel, J.P.; Han, J.; Benjamin, M.; Brown-Brandl, T.; Dórea, J.R.; Morris, D.; Norton, T.; Psota, E.; Rosa, G.J. The quest to develop automated systems for monitoring animal behavior. Appl. Anim. Behav. Sci. 2023, 265, 106000. [Google Scholar] [CrossRef]
- Teixeira, D.; Maron, M.; van Rensburg, B.J. Bioacoustic monitoring of animal vocal behavior for conservation. Conserv. Sci. Pract. 2019, 1, e72. [Google Scholar] [CrossRef]
- Taylor, A.M.; Reby, D.; McComb, K. Context-related variation in the vocal growling behaviour of the domestic dog (Canis familiaris). Ethology 2009, 115, 905–915. [Google Scholar] [CrossRef]
- Yin, S.; McCowan, B. Barking in domestic dogs: Context specificity and individual identification. Anim. Behav. 2004, 68, 343–355. [Google Scholar] [CrossRef]
- Yeo, C.Y.; Al-Haddad, S.; Ng, C.K. Dog voice identification (ID) for detection system. In 2012 Second International Conference on Digital Information Processing and Communications (ICDIPC); IEEE: New York, NY, USA, 2012. [Google Scholar] [CrossRef]
- Bishop, J.C.; Falzon, G.; Trotter, M.; Kwan, P.; Meek, P.D. Livestock vocalisation classification in farm soundscapes. Comput. Electron. Agric. 2019, 162, 531–542. [Google Scholar] [CrossRef]
- Nunes, L.; Ampatzidis, Y.; Costa, L.; Wallau, M. Horse foraging behavior detection using sound recognition techniques and artificial intelligence. Comput. Electron. Agric. 2021, 183, 106080. [Google Scholar] [CrossRef]
- Tsai, M.-F.; Huang, J.-Y. Sentiment analysis of pets using deep learning technologies in artificial intelligence of things system. Soft Comput. 2021, 25, 13741–13752. [Google Scholar] [CrossRef]
- Palanisamy, K.; Singhania, D.; Yao, A. Rethinking CNN models for audio classification. arXiv 2020, arXiv:2007.11154. [Google Scholar] [CrossRef]
- Abdoli, S.; Cardinal, P.; Koerich, A.L. End-to-end environmental sound classification using a 1D convolutional neural network. Expert Syst. Appl. 2019, 136, 252–263. [Google Scholar] [CrossRef]
- Dai, W.; Dai, C.; Qu, S.; Li, J.; Das, S. Very deep convolutional neural networks for raw waveforms. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2017. [Google Scholar] [CrossRef]
- Choi, S.; Seo, S.; Shin, B.; Byun, H.; Kersner, M.; Kim, B.; Kim, D.; Ha, S. Temporal convolution for real-time keyword spotting on mobile devices. arXiv 2019, arXiv:1904.03814. [Google Scholar] [CrossRef]
- Kim, B.; Chang, S.; Lee, J.; Sung, D. Broadcasted residual learning for efficient keyword spotting. arXiv 2021, arXiv:2106.04140. [Google Scholar] [CrossRef]
- Majumdar, S.; Ginsburg, B. Matchboxnet: 1d time-channel separable convolutional neural network architecture for speech commands recognition. arXiv 2020, arXiv:2004.08531. [Google Scholar] [CrossRef]
- Mohaimenuzzaman, M.; Bergmeir, C.; West, I.; Meyer, B. Environmental Sound Classification on the Edge: A Pipeline for Deep Acoustic Networks on Extremely Resource-Constrained Devices. Pattern Recognit. 2023, 133, 109025. [Google Scholar] [CrossRef]
- Nanni, L.; Maguolo, G.; Paci, M. Data augmentation approaches for improving animal audio classification. Ecol. Inform. 2020, 57, 101084. [Google Scholar] [CrossRef]
- Salamon, J.; Bello, J.P. Deep convolutional neural networks and data augmentation for environmental sound classification. IEEE Signal Process. Lett. 2017, 24, 279–283. [Google Scholar] [CrossRef]
- Loizou, P.C. Speech Enhancement: Theory and Practice; CRC Press: Boca Raton, FL, USA, 2007. [Google Scholar] [CrossRef]
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef]
- Ioffe, S.; Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning, Lille, France, 6–11 July 2015. Proceedings of Machine Learning Research (PMLR). [Google Scholar]
- Glorot, X.; Bordes, A.; Bengio, Y. Deep sparse rectifier neural networks. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics; JMLR Workshop and Conference Proceedings; JMLR.org: Norfolk, MA, USA, 2011. [Google Scholar]
- Nair, V.; Hinton, G.E. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (ICML-10), Haifa, Israel, 21–24 June 2010. [Google Scholar]
- Lin, M.; Chen, Q.; Yan, S. Network in network. arXiv 2013, arXiv:1312.4400. [Google Scholar] [CrossRef]
- Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; Salakhutdinov, R. Dropout: A simple way to prevent neural networks from overfitting. J. Mach. Learn. Res. 2014, 15, 1929–1958. [Google Scholar]
- Dumoulin, V.; Visin, F. A guide to convolution arithmetic for deep learning. arXiv 2016, arXiv:1603.07285. [Google Scholar] [CrossRef]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
- Kinga, D.; Adam, J.B. A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Kohavi, R. A study of cross-validation and bootstrap for accuracy estimation and model selection. In Proceedings of the IJCAI, Montreal, QC, Canada, 20–25 August 1995. [Google Scholar]
- Bengio, Y.; Grandvalet, Y. No unbiased estimator of the variance of k-fold cross-validation. J. Mach. Learn. Res. 2004, 5, 1089–1105. [Google Scholar]
- Sokolova, M.; Lapalme, G. A systematic analysis of performance measures for classification tasks. Inf. Process. Manag. 2009, 45, 427–437. [Google Scholar] [CrossRef]
- Sze, V.; Chen, Y.-H.; Yang, T.-J.; Emer, J.S. Efficient processing of deep neural networks: A tutorial and survey. Proc. IEEE 2017, 105, 2295–2329. [Google Scholar] [CrossRef]



| Class | Number of Samples | Percentage (%) |
|---|---|---|
| Bark | 46 | 44.66 |
| Howl | 57 | 55.34 |
| Total | 103 | 100 |
| Layer | Input Shape | Output Shape | Kernel/Setting | Trainable Parameters |
|---|---|---|---|---|
| Conv1D 1 | 1 × 32,000 | 8 × 15,985 | k = 32, s = 2, p = 0 | 264 |
| BatchNorm1D 1 | 8 × 15,985 | 8 × 15,985 | - | 16 |
| MaxPool1D 1 | 8 × 15,985 | 8 × 3996 | k = 4, s = 4 | 0 |
| Conv1D 2 | 8 × 3996 | 16 × 1991 | k = 16, s = 2, p = 0 | 2064 |
| BatchNorm1D 2 | 16 × 1991 | 16 × 1991 | - | 32 |
| MaxPool1D 2 | 16 × 1991 | 16 × 497 | k = 4, s = 4 | 0 |
| Conv1D 3 | 16 × 497 | 32 × 245 | k = 8, s = 2, p = 0 | 4128 |
| BatchNorm1D 3 | 32 × 245 | 32 × 245 | - | 64 |
| MaxPool1D 3 | 32 × 245 | 32 × 61 | k = 4, s = 4 | 0 |
| AdaptiveAvgPool1D | 32 × 61 | 32 × 1 | Output length = 1 | 0 |
| Linear 1 | 32 | 16 | - | 528 |
| Dropout | 16 | 16 | p = 0.3 | 0 |
| Linear 2 | 16 | 2 | - | 34 |
| Total | - | - | - | 7130 |
| Metric | Mean ± SD (%) | 95% CI (%) |
|---|---|---|
| Accuracy | 93.40 ± 2.62 | 91.52–95.27 |
| Macro precision | 94.43 ± 2.22 | 92.84–96.02 |
| Macro sensitivity | 92.71 ± 2.86 | 90.67–94.76 |
| Macro F1-score | 93.20 ± 2.75 | 91.23–95.16 |
| Condition | Accuracy (%) | Macro Precision (%) | Macro Sensitivity (%) | Macro F1 (%) |
|---|---|---|---|---|
| No added noise | 93.40 ± 2.62 [91.52–95.27] | 94.43 ± 2.22 [92.84–96.02] | 92.71 ± 2.86 [90.67–94.76] | 93.20 ± 2.75 [91.23–95.16] |
| 30 dB | 93.40 ± 2.62 [91.52–95.27] | 94.43 ± 2.22 [92.84–96.02] | 92.71 ± 2.86 [90.67–94.76] | 93.20 ± 2.75 [91.23–95.16] |
| 20 dB | 93.59 ± 2.30 [91.95–95.24] | 94.57 ± 1.99 [93.14–95.99] | 92.93 ± 2.50 [91.14–94.72] | 93.41 ± 2.40 [91.69–95.12] |
| 15 dB * | 94.08 ± 2.35 [92.39–95.76] | 94.89 ± 2.17 [93.34–96.44] | 93.50 ± 2.52 [91.69–95.30] | 93.92 ± 2.44 [92.18–95.66] |
| 10 dB | 94.76 ± 2.11 [93.25–96.26] | 95.28 ± 1.99 [93.85–96.71] | 94.34 ± 2.24 [92.73–95.95] | 94.64 ± 2.16 [93.10–96.19] |
| 5 dB | 94.56 ± 3.43 [92.11–97.02] | 94.73 ± 3.32 [92.35–97.10] | 94.63 ± 3.28 [92.28–96.97] | 94.51 ± 3.43 [92.06–96.97] |
| 0 dB | 79.42 ± 6.95 [74.44–84.39] | 83.30 ± 4.01 [80.43–86.16] | 81.13 ± 6.36 [76.58–85.68] | 79.06 ± 7.79 [73.49–84.64] |
| Condition | Accuracy (%) | Macro Precision (%) | Macro Sensitivity (%) | Macro F1 (%) |
|---|---|---|---|---|
| 30 dB | 93.40 ± 2.62 [91.52–95.27] | 94.43 ± 2.22 [92.84–96.02] | 92.71 ± 2.86 [90.67–94.76] | 93.20 ± 2.75 [91.23–95.16] |
| 20 dB | 93.59 ± 2.30 [91.95–95.24] | 94.57 ± 1.99 [93.14–95.99] | 92.93 ± 2.50 [91.14–94.72] | 93.41 ± 2.40 [91.69–95.12] |
| 15 dB * | 94.08 ± 2.44 [92.33–95.82] | 94.90 ± 2.22 [93.31–96.48] | 93.50 ± 2.63 [91.62–95.37] | 93.92 ± 2.53 [92.11–95.73] |
| 10 dB | 94.66 ± 2.34 [92.98–96.34] | 95.15 ± 2.28 [93.52–96.79] | 94.27 ± 2.45 [92.52–96.03] | 94.55 ± 2.40 [92.84–96.26] |
| 5 dB | 94.56 ± 3.75 [91.88–97.25] | 94.66 ± 3.66 [92.05–97.28] | 94.63 ± 3.59 [92.06–97.19] | 94.52 ± 3.76 [91.83–97.20] |
| 0 dB | 79.42 ± 7.28 [74.21–84.62] | 83.50 ± 4.06 [80.60–86.41] | 81.17 ± 6.65 [76.42–85.93] | 79.02 ± 8.22 [73.14–84.90] |
| Configuration | No-Added-Noise Accuracy (%) | Gaussian 5 dB (%) | Uniform 5 dB (%) |
|---|---|---|---|
| Full model, 2 s | 93.20 | 89.32 | 88.35 |
| Without noise augmentation | 94.17 | 68.93 | 68.93 |
| Without batch normalization | 90.29 | 88.35 | 87.38 |
| Without dropout | 91.26 | 87.38 | 86.41 |
| 1 s input duration | 82.52 | 81.55 | 84.47 |
| 3 s input duration | 92.23 | 89.32 | 89.32 |
| Parameters | Model Size (KB) | Conv1D FLOPs | FC FLOPs | Total FLOPs | CPU (ms/Sample) | GPU (ms/Sample) |
|---|---|---|---|---|---|---|
| 7130 | 27.85 | 18.35 M | 0.001 M | 18.35 M | 0.62 | 0.37 |
| Model | Input | Protocol | ACC (%) | SEN (%) | PRE (%) | F1 (%) | Params | FLOPS |
|---|---|---|---|---|---|---|---|---|
| AlexNet | MFCC | 80/20 split | 90.00 | 90.00 | 90.00 | 90.00 | 61.1 M | 714 M |
| DenseNet | Mel spectrogram | 80/20 split | 90.00 | 90.00 | 90.00 | 90.00 | 8.0 M | 2.9 G |
| EfficientNet | LFCC | 80/20 split | 90.00 | 90.00 | 90.00 | 90.00 | 5.3 M | 390 M |
| ResNet50 | MFCC/LFCC | 80/20 split | 86.00 | 86.00 | 85.00 | 86.00 | 25.6 M | 4.1 G |
| ResNet152 | MFCC/LFCC | 80/20 split | 86.00 | 86.00 | 85.00 | 86.00 | 60.2 M | 11.3 G |
| Proposed | Raw waveform | 10 × five-fold CV | 93.40 ± 2.62 | 92.71 ± 2.86 | 94.43 ± 2.22 | 93.20 ± 2.75 | 7.13 K | 18.35 M |
| Architecture | Input | Task | Parameters | Reported Operation Count |
|---|---|---|---|---|
| BC-ResNet-1 [16] | Log-Mel spectrogram | Keyword spotting | 9.2 K | 3.1 M multiplications |
| Micro-ACDNet [18] | Raw waveform | Environmental sound classification | 131 K | 14.82 M FLOPs |
| Proposed model | Raw waveform | Bark/Howl classification | 7.13 K | 18.35 M FLOPs |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dinsel, E.A.; Kodaz, H. A Lightweight 1D-CNN for Bark and Howl Classification from Raw Audio Waveforms Under Controlled Additive Noise. Appl. Sci. 2026, 16, 6819. https://doi.org/10.3390/app16136819
Dinsel EA, Kodaz H. A Lightweight 1D-CNN for Bark and Howl Classification from Raw Audio Waveforms Under Controlled Additive Noise. Applied Sciences. 2026; 16(13):6819. https://doi.org/10.3390/app16136819
Chicago/Turabian StyleDinsel, Emir Ali, and Halife Kodaz. 2026. "A Lightweight 1D-CNN for Bark and Howl Classification from Raw Audio Waveforms Under Controlled Additive Noise" Applied Sciences 16, no. 13: 6819. https://doi.org/10.3390/app16136819
APA StyleDinsel, E. A., & Kodaz, H. (2026). A Lightweight 1D-CNN for Bark and Howl Classification from Raw Audio Waveforms Under Controlled Additive Noise. Applied Sciences, 16(13), 6819. https://doi.org/10.3390/app16136819
