Next Article in Journal
Hamiltonian Monte Carlo with Random Effect for Analyzing Cyclist Crash Severity
Next Article in Special Issue
Blind Source Separation in Polyphonic Music Recordings Using Deep Neural Networks Trained via Policy Gradients
Previous Article in Journal
Omnidirectional Haptic Guidance for the Hearing Impaired to Track Sound Sources
Previous Article in Special Issue
Efficient Retrieval of Music Recordings Using Graph-Based Index Structures
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms

1
Graduate School of Informatics, Kyoto University, Kyoto 606-8501, Japan
2
PRESTO, Japan Science and Technology Agency (JST), Saitama 332-0012, Japan
*
Author to whom correspondence should be addressed.
Signals 2021, 2(3), 508-526; https://doi.org/10.3390/signals2030031
Submission received: 8 January 2021 / Revised: 19 July 2021 / Accepted: 21 July 2021 / Published: 13 August 2021
(This article belongs to the Special Issue Advances in Processing and Understanding of Music Signals)

Abstract

This paper describes an automatic drum transcription (ADT) method that directly estimates a tatum-level drum score from a music signal in contrast to most conventional ADT methods that estimate the frame-level onset probabilities of drums. To estimate a tatum-level score, we propose a deep transcription model that consists of a frame-level encoder for extracting the latent features from a music signal and a tatum-level decoder for estimating a drum score from the latent features pooled at the tatum level. To capture the global repetitive structure of drum scores, which is difficult to learn with a recurrent neural network (RNN), we introduce a self-attention mechanism with tatum-synchronous positional encoding into the decoder. To mitigate the difficulty of training the self-attention-based model from an insufficient amount of paired data and to improve the musical naturalness of the estimated scores, we propose a regularized training method that uses a global structure-aware masked language (score) model with a self-attention mechanism pretrained from an extensive collection of drum scores. The experimental results showed that the proposed regularized model outperformed the conventional RNN-based model in terms of the tatum-level error rate and the frame-level F-measure, even when only a limited amount of paired data was available so that the non-regularized model underperformed the RNN-based model.
Keywords: automatic drum transcription; self-attention mechanism; transformer; positional encoding; masked language model automatic drum transcription; self-attention mechanism; transformer; positional encoding; masked language model

Share and Cite

MDPI and ACS Style

Ishizuka, R.; Nishikimi, R.; Yoshii, K. Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms. Signals 2021, 2, 508-526. https://doi.org/10.3390/signals2030031

AMA Style

Ishizuka R, Nishikimi R, Yoshii K. Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms. Signals. 2021; 2(3):508-526. https://doi.org/10.3390/signals2030031

Chicago/Turabian Style

Ishizuka, Ryoto, Ryo Nishikimi, and Kazuyoshi Yoshii. 2021. "Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms" Signals 2, no. 3: 508-526. https://doi.org/10.3390/signals2030031

APA Style

Ishizuka, R., Nishikimi, R., & Yoshii, K. (2021). Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms. Signals, 2(3), 508-526. https://doi.org/10.3390/signals2030031

Article Metrics

Back to TopTop