Can Audio Reveal Music Performance Difficulty? Insights From the Piano Syllabus Dataset

  • Ramoneda, Pedro
  • Lee, Minhee
  • Jeong, Dasaem
  • Valero-Mas, Jose J.
  • Serra, Xavier
Citations

WEB OF SCIENCE

3

초록

Automatically estimating the performance difficulty of a music piece represents a key process in music education to create tailored curricula according to the individual needs of the students. Given its relevance, the Music Information Retrieval (MIR) field comprises some proof-of-concept works addressing this task that mainly focus on high-level music abstractions such as machine-readable scores or music sheet images. In this regard, the potential of directly analyzing audio recordings has generally been neglected. This work addresses this gap in the field with two contributions: (i) PSyllabus, the first audio-based difficulty estimation dataset-collected from Piano Syllabus community-featuring 7,901 piano pieces across 11 difficulty levels from 1,233 composers as well as two additional benchmark datasets particularly compiled for evaluation purposes; and (ii) a recognition framework capable of managing different input representations-both in unimodal and multimodal manners-derived from audio to perform the difficulty estimation task. The comprehensive experimentation comprising different pre-training schemes, input modalities, and multi-task scenarios proves the validity of the hypothesis and establishes PSyllabus as a reference dataset for audio-based difficulty estimation in the MIR field. The dataset, developed code, and trained models are publicly shared to promote further research in the field.

키워드

EstimationTrainingAcousticsBenchmark testingMultitaskingSpeech processingPrompt engineeringHidden Markov modelsAudio recordingAttention mechanismsMusic difficultymusic information retrievalmusic technology educationperformance analysisplayabilityCURRICULUM DESIGN
제목
Can Audio Reveal Music Performance Difficulty? Insights From the Piano Syllabus Dataset
저자
Ramoneda, PedroLee, MinheeJeong, DasaemValero-Mas, Jose J.Serra, Xavier
DOI
10.1109/TASLPRO.2025.3539018
발행일
2025
유형
Article
저널명
IEEE Transactions on Audio, Speech and Language Processing
33
페이지
1129 ~ 1141