SKY-Piano: A Multimodal Piano Performance Dataset
Source: arXiv:2607.27296 · Published 2026-07-29 · By Joonhyung Bae, Dawon Park, Taegyun Kwon, Yoon-Seok Choi, Hyeon Hur, Satoshi Obata et al.
TL;DR
SKY-Piano addresses a critical gap in music information retrieval (MIR) research by providing a multimodal piano performance dataset with real measured hand and body motion, synchronized with audio, MIDI, multi-view video, and symbolic scores. Prior datasets either lacked real hand motion capture or did not combine motion with detailed annotations across performer expertise, technique, and difficulty. SKY-Piano contains 11 hours of recordings from 19 pianists (7 professional, 12 amateur) performing a carefully selected repertoire that enables controlled studies across skills and musical pieces. The dataset includes both flagged and imputed motion data, Visual3D body kinematics, and an interactive web browser for easy exploration. They also contribute a fingering annotation pipeline that derives pseudo fingering labels from motion and MIDI, achieving high precision on expert-audited samples.
As a practical demonstration, the authors fine-tune a leading MIDI-to-motion generation model (Tipiano) on SKY-Piano, verifying the dataset supports state-of-the-art kinematic synthesis and adapts well across domains despite subtle differences in fingertip definitions. Overall, SKY-Piano brings precise multimodal measurements, structured repertoire, and advanced annotation tools in one publicly released corpus, enabling fine-grained expressive performance modeling and piano-specific MIR beyond audio and MIDI alone.
Key findings
- Dataset contains 11 hours of multimodal synchronized data from 19 pianists (7 professional, 12 amateur) performing 35 pieces including 26 structured exercises/pieces.
- Motion data includes 46 hand markers (23 per hand) and 17 Visual3D body joints, recorded at 120 Hz with real optical motion capture.
- Marked missing motion data averages 14% corpus-wide, up to 44% on thumb-carpal markers, provided in both flagged and SAITS-imputed forms.
- Fingering annotation pipeline achieves 94.5% strict precision on 2,269 expert-annotated professional notes with 94.0% coverage; imputation tier extends coverage to 100%.
- MIDI-to-motion finetuning of Tipiano model reduces MPJPE from 65.9 mm (zero-shot) to 48.8 mm (fine-tuned), a 25.9% reduction, approaching original in-domain error of 32.8 mm.
- The dataset supports cross-comparison of playing technique, difficulty, and performer expertise on shared musical material, enabling controlled skill and expression analyses.
- Multi-view 4-camera video plus synchronized multi-modal streams are precisely aligned using hardware SMPTE timecode and manual verification, achieving ∼3 ms audio-MIDI alignment accuracy.
- Pseudo fingering uses measured fingertip depth (z) and key contact geometry, improving on prior video-only estimation by exploiting real 3D motion capture.
Methodology — deep read
Threat Model & Assumptions: Not a security paper; dataset targets MIR and piano performance modeling research. No adversarial threat model.
Data: The corpus includes 11 hours of performances by 19 pianists (7 professionals, 12 amateurs) covering 35 pieces selected based on playing technique, difficulty, and expertise. The repertoire includes two technique exercise categories, graded pieces spanning beginner to advanced, and free professional repertoire. Audio is recorded at 48kHz/24-bit, MIDI from a Yamaha Disklavier DC7X, hand and body motion via optical mocap (OptiTrack for hands with 46 markers total, Qualisys for full-body with 17 Visual3D joints) at 120 Hz, and multi-view 4K/60fps video. Motion data is released as both raw with validity flags and imputed-filling via SAITS neural model. Pieces are segmented via MIDI silence and aligned to trimmed MusicXML scores verified by a musicologist.
Architecture/Algorithms: A geometry-based fingering annotation pipeline adapts PianoVAM's video approach but uses direct fingertip z-depth readings from mocap, with three tiers—geometric scoring, z-onset refinement (to resolve multi/no candidate ambiguity), and BiLSTM imputation for residual cases—enabling 100% fingering coverage with explicit ambiguity tags. For generation, Tipiano’s cascaded hand motion synthesis model (21 MANO keypoints predicted from MIDI+fingering) is fine-tuned on SKY-Piano.
Training Regime: For fingering imputation, the BiLSTM uses a 4.2M parameter model trained with masked classification on professional trials augmented with symbolic prior datasets; details on epochs, batch size, or optimizer are not specified. Tipiano is fine-tuned under 5-fold leave-one-pianist-out cross-validation on the professional subset, with MPJPE and key-contact F1 as metrics.
Evaluation Protocol: Fingering precision is audited on 2,269 hand-labeled notes from three graded pieces covering beginner to advanced levels. Fingering coverage and precision are reported. For MIDI-to-motion generation, MPJPE and key-contact F1 are compared across zero-shot, fine-tuned, and original in-domain conditions. The multi-modal data alignment is verified manually and with cross-correlation metrics, achieving ∼3 ms audio-MIDI sync error. No adversarial or distribution-shift evaluations reported.
Reproducibility: The dataset, pipeline code, interactive browser, and MIDI-to-motion baseline Fine-tuning experiments are publicly released via the project webpage with a Creative Commons CC-BY 4.0 license. Recorded sessions comply with IRB and participant consent with modality-specific release permissions. However, some proprietary capture hardware and software (Visual3D kinematics) may limit full replication.
Example: In one professional trial, a MIDI note’s onset is matched with measured fingertip z-depth crossing the keyboard surface within ±50 ms at −10.1 mm median depth confirming actual key press, enabling candidate finger assignment. Ambiguous notes are resolved with BiLSTM imputation achieving 94.5% precision on this subset.
The multi-view videos are face-anonymized and timecode-synchronized with motion and audio streams, facilitating simultaneous browsing across modalities in a web interface.
Technical innovations
- Release of a large-scale multimodal piano corpus combining real optical hand and body motion capture with synchronized multi-view video, audio, MIDI, and MusicXML aligned score annotations.
- A novel pseudo fingering annotation pipeline leveraging precise mocap fingertip z-depth measurements combined with geometric scoring and a BiLSTM imputation tier to achieve full coverage and high precision.
- Implementation of frame-level synchronized multi-modal data streams aligned via SMPTE timecode and refined by cross-correlation and spatial Procrustes alignment to ensure precise inter-modality timing.
- Fine-tuning and cross-domain adaptation of the Tipiano MIDI-to-motion generation model to a real measured-motion dataset, demonstrating improved MPJPE despite different fingertip conventions.
Datasets
- SKY-Piano — 11 hours — Publicly released under CC-BY 4.0, with data from 19 (7 professional / 12 amateur) pianists including audio, MIDI, real hand/body mocap, multi-view video, and scores
Baselines vs proposed
- Tipiano (trained on Für Elise pseudo-motion corpus): MPJPE = 32.8 mm, Key-contact F1 = 0.910
- Tipiano zero-shot on SKY-Piano professional test set: MPJPE = 65.9 mm, Key-contact F1 = 0.93
- Tipiano fine-tuned on SKY-Piano (5-fold leave-one-pianist-out): MPJPE = 48.8 mm, Key-contact F1 = 0.66
Figures from the paper
Figures are reproduced from the source paper for academic discussion. Original copyright: the paper authors. See arXiv:2607.27296.

Fig 1: SKY-Piano at a glance. Hardware-synchronized modalities for one released trial in the interactive viewer.
Limitations
- Relatively small participant pool (19 total pianists) limits generalizability and modeling power.
- Repertoire is focused on Western classical technique exercises and graded pieces; lacks extended or contemporary piano techniques.
- Fingering annotations are pseudo labels with ambiguity flags rather than fully hand-verified ground truth across entire dataset.
- Marker occlusion and motion dropout are significant (up to 44% on some finger markers), requiring imputation which may introduce artifacts.
- No adversarial or out-of-domain robustness evaluations for fingering or motion models provided.
- The fine-tuning experiment reveals motion conventions differ (marker vs MANO joints), complicating metric transfer and interpretation.
Open questions / follow-ons
- How does performer expertise influence biomechanical and expressive features across extended repertoires beyond the dataset’s limited pieces?
- Can richer hand and finger kinematics from measured mocap improve automatic fingering estimation models beyond geometry-based heuristics?
- What are the best modeling approaches to address missing data imputation uncertainties and their impact on downstream tasks like motion synthesis?
- How generalizable are the motion-to-MIDI and MIDI-to-motion models across different keyboard instruments or playing styles beyond classical piano?
Why it matters for bot defense
For bot-defense and CAPTCHA practitioners focused on bot detection through behavioral biometrics or multimodal activity recognition, SKY-Piano illustrates how combining precise, synchronized multimodal data streams (video, motion capture, symbolic events) can enable detailed attribution of subtle, skilled human motor behaviors. Although specific to piano performance, the dataset’s approach to label ambiguity handling, multi-tier annotation pipelines, and fusion of diverse sensor modalities provides methodological insights. Applying similar multimodal synchronization and imputation strategies may improve robustness in behavior-based human verification systems, especially where fine-grained motion or pose tracking data is noisy or partially observed. Moreover, their interactive browser tool models best practices in visualizing complex synchronized modalities, useful when designing analytics workflows for security tasks. However, domain differences mean direct application to CAPTCHA or bot defense requires adaptation beyond musical contexts.
Cite
@article{arxiv2607_27296,
title={ SKY-Piano: A Multimodal Piano Performance Dataset },
author={ Joonhyung Bae and Dawon Park and Taegyun Kwon and Yoon-Seok Choi and Hyeon Hur and Satoshi Obata and Shigeru Kai and Yohei Wada and Yu Takahashi and Akira Maezawa and Jaebum Park and Jonghwa Park and Juhan Nam },
journal={arXiv preprint arXiv:2607.27296},
year={ 2026 },
url={https://arxiv.org/abs/2607.27296}
}