arXiv · 2411.09037
Pay Attention to the Keys: Visual Piano Transcription Using Transformers
Abstract
Visual piano transcription (VPT) is the task of obtaining a symbolic representation of a piano performance from visual information alone (e.g., from a top-down video of the piano keyboard). In this work we propose a VPT system based on the vision transformer (ViT), which surpasses previous methods based on convolutional neural networks (CNNs). Our system is trained on the newly introduced R3 dataset, consisting of ca.~31 hours of synchronized video and MIDI recordings of piano performances. We additionally introduce an approach to predict note offsets, which has not been previously explored in this context. We show that our system outperforms the state-of-the-art on the PianoYT dataset for onset prediction and on the R3 dataset for both onsets and offsets.
Explore related subjects
Keep this discovery
Uros Zivanovic, Ivan Pilkov, Carlos Eduardo Cancino-Chacón. 2024-11-13. Pay Attention to the Keys: Visual Piano Transcription Using Transformers. https://doi.org/10.24963/ijcai.2025%2F1138
Cite the original work for its findings. Save a collection to share your selection of sources.