Source-linked AI summary

Expressive Robotic Pianist: Mastering Complex Piano Repertoire with Graph-Mimic and Musical Dynamics

Yanhong Liang, Xianwei Liu, Chaojie Fu, Shaowen Cheng, Yanyan Yuan, Chengwei Zhuo, Xi Chen, Yongbin Jin, Wei Yang, Hongtao Wang

arXiv:2609.10844v1cs.RO

TL;DR

Robotic pianists struggle to combine human-like fingering with nuanced dynamic control. This work uses Graph-Mimic and a physics-inspired acoustic model within reinforcement learning, achieving expressive performance comparable to humans while retaining accurate note execution.

  • Problem

    Robotic piano systems struggle to reproduce fluid fingering, contact behavior, and nuanced dynamics required for expressive performance.

  • Method

    The framework combines graph-based motion imitation with a physics-inspired acoustic model that maps key velocity to musical loudness.

  • Results

    The system achieves expressive performance comparable to human performance across varied repertoire, while maintaining accurate note-level execution.

  • Takeaways & Limitations

    Graph-based hand-motion modeling and acoustic dynamics control provide a pathway from mechanically accurate robotic piano playing toward human-like expressivity.

  • Takeaways & Limitations

    The evaluation relies on F1 for pitch–onset accuracy, while four-level loudness classification cannot fully represent expressive timing and the continuous dynamic range.

Abstract

from arXiv · show

Enabling robots to perform musical instruments with human-level expressivity represents a frontier in bridging the gap between mechanical precision and artistic interpretation. Despite advances in robotic dexterity, replicating the fluid finger transitions and nuanced dynamic control characteristic of human pianists remains a significant challenge. Through a reinforcement learning-based control framework, we demonstrate that a dexterous robotic hand can achieve high-fidelity performance across a diverse piano repertoire. Central to our approach is a graph-based optimization strategy that guides the robot to generate natural pre-press and key-press fingering strategies that closely resemble human movement patterns. To achieve expressive sound production, the control system is coupled with a physics-inspired acoustic model that modulates keypress velocity to accurately reproduce the dynamic variations specified in musical scores. Quantitative evaluations demonstrate that our expressive control model significantly outperforms baseline methods in both finger morphology similarity and dynamic velocity accuracy. In a perceptual test involving participants from diverse listener groups, performances generated by our system are significantly preferred over baseline robotic performances and are indistinguishable from human performances for non-professional audiences. Furthermore, extensive experiments across multiple musical styles confirm that our method maintains high note-level accuracy while achieving expressive performance. Our approach provides a robust pathway for robotic systems to move beyond mere mechanical accuracy, elevating robotic musicianship to a level of expressive performance comparable to human pianists.

Introduction

Robotic piano performance remains difficult because it requires coordinated fingering, articulation, timing, and expressive dynamics beyond mechanical note execution. The proposed framework combines Graph-Mimic motion imitation with acoustic modeling to address these challenges across demanding repertoire.

  • Introduction: Robotic piano performance requires fluid fingering, articulation-specific touch, precise timing, and dynamic control, making it a demanding test of dexterity.These requirements span pre-press coordination, key-press contact behavior, and expressive sound production.
  • Introduction: Existing robotic piano systems and reinforcement-learning approaches generally demonstrate simpler note execution, timing, or duration control rather than full expressive performance.Prior systems include pre-programmed hands, force-modulated mechanisms, and RL agents trained for simple sequences or multiple pieces.
  • Introduction: The framework combines motion imitation with a simplified acoustic model that relates key angular velocity to note volume for score-consistent dynamic control.The acoustic model focuses on energy transfer during key presses instead of directly simulating string vibration.
  • Introduction: Graph-Mimic represents human hand motion through normalized spatial relationship graphs, enabling morphology-agnostic retargeting based on shared feature points.The representation targets spatial joint relationships rather than requiring strict joint-to-joint correspondence.
  • Introduction: Across varied piano pieces, the system generates musically appropriate fingering and nuanced loudness modulation with expressivity comparable to human performance.The evaluation includes objective metrics and a human perceptual study in addition to repertoire spanning difficulty, duration, and style.

Results

The integrated controller produces more natural fingering, score-aligned dynamics, and expressive performances while retaining high note-level accuracy across varied repertoire and techniques.

  • Articulation-specific touch techniques during key-press phase: The robot differentiated staccato through vertical fingertip contact and legato through overlapping depressions with stable pad contact.These contrasting contact strategies correspond to crisp versus continuous articulation styles.
  • Articulation-specific touch techniques during key-press phase: Graph-Mimic produced natural extension–flexion–extension trajectories and coordinated postures, unlike the baseline’s abnormal PIP extension and inefficient curled motions.The thumb-under ablation further showed that removing Graph-Mimic leads to suboptimal wrist compensation and severely curled postures.
  • Piano acoustics model for tuning musical dynamics: 85.7% velocity accuracy on challenging wide-range pieces shows that dynamic guidance preserves reference loudness categories as MIDI velocity variance increases.Baseline accuracy drops sharply and reaches zero for wide-range pieces.
  • The piano perceptual test and comparative analysis: Listener rankings significantly favored the expressive robot over the baseline, while untrained listeners found no significant preference difference between it and either human performance.Formally trained pianists still ranked both human performances above the robots.
  • The piano perceptual test and comparative analysis: Expressive control reduced reference MIDI-velocity error to 10.72 versus 15.50 for the baseline, while velocity accuracy reached 50% versus 37.5%.The expressive robot largely remained within the human dynamic range and reproduced a clear opening crescendo absent from the baseline.
  • Experimental validation across diverse musical styles: 0.96 average simulated F1 and 0.80–0.95 real-world scores show high note-level accuracy across evaluated pieces.The system also reproduced sophisticated chords, overlapping presses, rapid passages, and large octave spans.

Discussion

The study combines Graph-Mimic, acoustic modeling, and reinforcement learning to produce human-like fingering and expressive dynamics, while identifying limits in dynamic resolution, arm control, evaluation, sensing, and policy generality.

  • Discussion: Graph-Mimic and acoustic modeling enable technically demanding repertoire, duet coordination, and expressive robotic piano performance.The system performs Grade 7 works, including Croatian Rhapsody, and models musical dynamics through loudness control.
  • Discussion: F1 reaches 0.92 on Twinkle, Twinkle, Little Star versus approximately 0.60 for previous approaches, while non-experts statistically distinguish the system from neither human performances nor baseline expressivity.
  • Discussion: The graph representation encodes phalangeal and inter-fingertip relationships and may extend to more natural, efficient, and compliant manipulation strategies.
  • Discussion: Four-level loudness classes cannot capture the full ppp–fff dynamic continuum, motivating pressure-sensitive tactile control and continuous dynamics modeling.
  • Discussion: The system lacks proximal arm kinematics, relies on simulation piano-state variables, and evaluates expressivity only partially through F1, dynamics metrics, and perception.
  • Discussion: Song-specific policies constrain generalization toward generalist, real-time, and improvisational piano controllers.

Study Design

The study benchmarks expressive robotic piano performance against professional human pianists and a baseline controller using perceptual judgments and objective pitch–onset accuracy.

  • Study Design: The evaluation compares an expressive robotic controller with professional human pianists and a baseline robotic controller across two experimental dimensions.
  • Study Design: A perceptual test used 125 participants grouped as trained pianists, generally educated listeners, and untrained listeners to rank randomized audio excerpts.
  • Study Design: Objective validation used F1 for mechanical precision, aggregating 10 physical trials per piece and selecting top-scoring human and simulation references.

Reinforcement Learning Setup

The reinforcement-learning setup augments standard piano rewards with graph-based hand-motion similarity and physics-inspired musical-dynamics objectives.

  • Reinforcement Learning Setup: The setup combines Key Press, Energy Penalty, and Fingering rewards with Graph Distance and Musical Dynamics rewards.These objectives respectively encourage correct notes, lower energy use, target-directed finger motion, human-like morphology, and expressive loudness.
  • Reinforcement Learning Setup: Graph Distance compares normalized human and robotic Action Frame Graphs built from corresponding hand feature points and vectors.
  • Reinforcement Learning Setup: The graph design preserves finger kinematics and hand contour while preventing short, functionally critical segments from being overshadowed by longer ones.
  • Reinforcement Learning Setup: A tolerance-based Graph Distance reward provides smooth, differentiable decay as morphology mismatch exceeds the allowed range.The implementation uses bounds [0, 5], margin 10, and reward value 0.1 at the margin.
  • Reinforcement Learning Setup: The acoustic model links key angular velocity to sound loudness through energy transfer from key motion to hammer and string vibration.The controller augments observations with angular velocities for all 88 keys and rewards agreement between computed and reference MIDI velocity.
  • Reinforcement Learning Setup: MIDI Velocity represents loudness on a 0–127 scale, with a scaling constant mapping key angular velocity to that value.

Physical World Setup

The physical platform pairs a Yamaha digital piano and UR5 arm with the InReal dexterous hand, whose speed, force, precision, and span support expressive piano performance.

  • Physical World Setup: The platform comprises a digital piano, UR5 robotic arm, and dexterous hand.
  • Physical World Setup: The Yamaha P-48B provides 88 weighted hammer-action keys, stereo-sampled tones, expressive dynamics, and USB MIDI acquisition for quantitative analysis.
  • Physical World Setup: The UR5 offers six degrees of freedom, an 850 mm working envelope, 5 kg payload, and ±0.1 mm repeatability, but limits rapid octave transitions in virtuoso pieces.
  • Physical World Setup: The InReal hand is benchmarked across degrees of freedom, actuation speed, fingertip force, fingertip width, and maximum span.
  • Physical World Setup: The hand provides 18 active DoFs, speeds above 1000°/s, 10 Hz key strikes, 18.5 N fingertip force, 14 mm fingertips, and a 265 mm thumb-to-pinky span.
  • Physical World Setup: Together, these capabilities establish the InReal hand as a highly versatile platform for robotic piano performance.

Statistical Analysis

Listener preferences were analyzed using rank-based nonparametric tests across four audio excerpts and three listener cohorts.

  • Data and cohorts: The analysis used participant rank data from trained pianists, general-education participants, and untrained listeners.The cohort sizes were n=24, 64, and 37, respectively.
  • Overall tests: Friedman tests evaluated overall ranking differences among four excerpts within each listener cohort.The four conditions were human1, human2, expressive robot, and baseline robot.
  • Pairwise tests: Pairwise Wilcoxon signed-rank tests were applied when Friedman tests indicated a significant overall effect.These tests evaluated within-subject differences between two excerpts.
  • Multiple comparisons: Holm–Bonferroni adjustment controlled family-wise error across the multiple pairwise comparisons.Statistical significance was defined as p < 0.05 after correction.

Supplementary materials

The supplementary materials provide figures covering the robotic pianist’s models, fingering, dynamics, performance, hardware, and reinforcement-learning setup.

  • Supplementary package: The supplementary package includes supplementary text, figures, tables, references, movies with legends, and participant preference data.The listed materials span S1–S6, Figs. S1–S5, Tables S1–S6, Movies S1–S13, and Data S1.
  • Learning setup: The reinforcement-learning setup table summarizes observation, action, and reward components together with their meanings and dimensions.It accompanies the supplementary documentation of the learning configuration.

Repertoire with Graph-Mimic and Musical Dynamics

The framework models configuration-dependent finger mechanics and combines graph-based coordination with acoustic dynamics for expressive piano performance. Across repertoire and evaluations, it improves coordination, dynamic control, and execution of technically demanding passages.

  • Finger mechanics: Configuration-dependent fingertip stiffness varies with finger flexion because changing joint geometry alters the Jacobian and effective stiffness.Simulation results further show that total torque decreases with increasing flexion, reaching a minimum near q1 = q2 = q3 ≈ 0.67 rad.
  • Graph-Mimic coordination: Graph-Mimic prioritizes finger articulation over wrist compensation, avoiding the missing-key failure observed with fingering-reward-only control.The baseline seeks suboptimal wrist movements to satisfy key-finger constraints, disrupting temporal continuity at a critical transition.
  • Expert evaluation: Expert assessments identify finger-balance and accentuation issues in human reference performances, providing detailed interpretive context for evaluating robotic phrasing.The evaluations discuss uneven fourth- and fifth-finger execution, motivic emphasis, and the naturalness of phrase resolution.
  • Performance evaluation: The expressive robotic system shows more consistent timing, articulation, coordination, and dynamic shaping than the baseline, while subtler phrasing remains imperfect.The system produces gradual softening across measures 9–12, but some accents remain abrupt or overemphasized.
  • Musical dynamics: Velocity modulation reproduces dynamic contrasts in Für Elise by aligning robotic keypresses with reference MIDI dynamics.
  • Repertoire evaluation: Experiments cover mixed-key chords, overlapping presses, rapid sixteenth-note passages, multi-finger chords, wide octave spans, and human–robot ensemble timing.These demonstrations test adaptive posture, high-speed precision, temporal control, extended reach, and stability across varied repertoire.
Loading 2609.10844v1…