Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection

Wei, Weixing; Zhao, Jiahao; Wu, Yulun; Yoshii, Kazuyoshi

Computer Science > Sound

arXiv:2503.01362 (cs)

[Submitted on 3 Mar 2025]

Title:Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection

Authors:Weixing Wei, Jiahao Zhao, Yulun Wu, Kazuyoshi Yoshii

View PDF HTML (experimental)

Abstract:This paper describes a streaming audio-to-MIDI piano transcription approach that aims to sequentially translate a music signal into a sequence of note onset and offset events. The sequence-to-sequence nature of this task may call for the computationally-intensive transformer model for better performance, which has recently been used for offline transcription benchmarks and could be extended for streaming transcription with causal attention mechanisms. We assume that the performance limitation of this naive approach lies in the decoder. Although time-frequency features useful for onset detection are considerably different from those for offset detection, the single decoder is trained to output a mixed sequence of onset and offset events without guarantee of the correspondence between the onset and offset events of the same note. To overcome this limitation, we propose a streaming encoder-decoder model that uses a convolutional encoder aggregating local acoustic features, followed by an autoregressive Transformer decoder detecting a variable number of onset events and another decoder detecting the offset events for the active pitches with validation of the sustain pedal at each time frame. Experiments using the MAESTRO dataset showed that the proposed streaming method performed comparably with or even better than the state-of-the-art offline methods while significantly reducing the computational cost.

Comments:	Accepted to ISMIR 2024
Subjects:	Sound (cs.SD); Information Retrieval (cs.IR); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2503.01362 [cs.SD]
	(or arXiv:2503.01362v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2503.01362

Submission history

From: Weixing Wei [view email]
[v1] Mon, 3 Mar 2025 09:55:54 UTC (757 KB)

Computer Science > Sound

Title:Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators