ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Modeling the natural contour of fundamental frequency (F0) plays a critical role in music audio synthesis. However, transcribing and managing multiple F0 contours in polyphonic music is challenging, and explicit F0 contour modeling has not yet been explored for polyphonic instrumental synthesis. In this paper, we present ViolinDiff, a two-stage diffusion-based synthesis framework. For a given violin MIDI file, the first stage estimates the F0 contour as pitch bend information, and the second stage generates mel spectrogram incorporating these expressive details. The quantitative metrics and listening test results show that the proposed model generates more realistic violin sounds than the model without explicit pitch bend modeling. Audio samples are available online: daewoung.github.io/ ViolinDiff-Demo. © 2025 Institute of Electrical and Electronics Engineers Inc.. All rights reserved.

키워드

Diffusion ModelsExpressive PerformanceNeural Audio SynthesisPitch Bend ModelingViolin Synthesis
제목
ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
저자
Kim, DaewoongDong, Hao WenJeong, Da saem
DOI
10.1109/ICASSP49660.2025.10890613
발행일
2025
유형
Proceedings Paper
저널명
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings