Reducing latency of neural automatic piano transcription models 인공신경망 기반 저지연 피아노 채보 모델

Citations

SCOPUS

4

초록

Automatic Music Transcription (AMT) is a task that detects and recognizes musical note events from a given audio recording. In this paper, we focus on reducing the latency of real-time AMT systems on piano music. Although neural AMT models have been adapted for real-time piano transcription, they suffer from high latency, which hinders their usefulness in interactive scenarios. To tackle this issue, we explore several techniques for reducing the intrinsic latency of a neural network for piano transcription, including reducing window and hop sizes of Fast Fourier Transformation (FFT), modifying convolutional layer’s kernel size, and shifting the label in the time-axis to train the model to predict onset earlier. Our experiments demonstrate that combining these approaches can lower latency while maintaining high transcription accuracy. Specifically, our modified model achieved note F1 scores of 92.67 % and 90.51 % with latencies of 96 ms and 64 ms, respectively, compared to the baseline model’s note F1 score of 93.43 % with a latency of 160 ms. This methodology has potential for training AMT models for various interactive scenarios, including providing real-time feedback for piano education. ©©2023 The Acoustical Society of Korea.

키워드

Automatic music transcriptionPiano transcriptionLow-latencyConvolutional neural network자동 채보피아노 채보저지연합성곱 신경망
제목
Reducing latency of neural automatic piano transcription models 인공신경망 기반 저지연 피아노 채보 모델
저자
Lee, Dasol정다샘
DOI
10.7776/ASK.2023.42.2.102
발행일
2023
유형
Article
저널명
한국음향학회지
42
2
페이지
102 ~ 111