Research Team Led by Jeong Da-saem, Professor in the Department of Art & Technology, Publishes Paper in IEEE TASLPRO, a Prestigious International Journal in Signal Processing

작성일: 2026-05-15
Research Team Led by Jeong Da-saem, Professor in the Department of Art & Technology,  Publishes Paper in IEEE TASLPRO, a Prestigious International Journal in Signal Processing
A paper co-authored by Professor Jeong Da-saem’s research team in the Department of Art & Technology (including Jeong Jong-min (Master’s), Kim Dong-min (Master’s), Ph.D. student Lee Si-hoon, and Master’s student Cho Seol-a), in collaboration with Postdoctoral Researcher So Hyung-jun from Seoul National University and Professor Chris Donahue’s research team at Carnegie Mellon University, has been published in IEEE Transactions on Audio, Speech and Language Processing (TASLPRO), a prestigious international journal in signal processing.

In the paper, titled "U-MusT: A Unified Framework for Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio," the research team proposes a universal model designed to simultaneously handle translation tasks across various musical modalities.

Music exists in various modalities, including score images, symbolic notation, MIDI, and audio. Translation tasks between these modalities—such as automatic music transcription and optical music recognition—are core challenges in music information retrieval (MIR). While previous research has focused on models specialized for individual translation tasks, Professor Jeong Da-saem’s research team has developed a universal model capable of simultaneously learning translation tasks across multiple modalities.

The model proposed in this study achieved the current state-of-the-art (SOTA) symbol error rate in piano sheet music recognition. Furthermore, it is the world’s first model capable of generating expressive performance audio directly from sheet music images without any intermediate steps. The research team also significantly contributed to the music information retrieval (MIR) community by releasing a large-scale dataset of over 1,300 hours of sheet music image-performance audio pairs, which was constructed specifically for training the proposed model.

The paper is also scheduled to be presented at ICASSP 2026, the world’s largest conference in signal processing, which will be held in Barcelona, Spain, starting May 4, 2026.

[Keyword]

Music Information Retrieval, AI Music Generation, Cross-modal Learning

[Summary]

Through music information retrieval research, Professor Jeong Da-saem's team has developed a universal framework, "U-MusT," for cross-modal learning between score images and audio. This AI music generation model, published in IEEE TASLPRO, is the first to directly produce expressive performance audio from sheet music, achieving state-of-the-art accuracy in music recognition.