SearcharxivSearch

arXiv subjects

Jaebum Park

Publications and source records attributed to Jaebum Park.

8 recordsLinked to original sources

SKY-Piano: A Multimodal Piano Performance Dataset

Music information retrieval research on piano performance increasingly involves diverse modalities of data and annotations beyond audio and MIDI. We present SKY-Piano, a multimodal piano performance dataset that includes 11 hours of performance recordings of motion, multi-view video, audio, MIDI from 7 professional and 12 amateur pianists along with MusicXML scores. The performance pieces were selected considering playing technique, difficulty, and performer expertise on a shared core repertoire. The motion data include both hand and body motion, released in both flagged form, where samples lost to marker occlusion are marked as unreliable, and imputed form, where those gaps are reconstructed, together with Visual3D body-segment kinematics and other time-synchronized modalities. To easily browse different modalities of data at a glance, we provide an interactive web browser. In addition, we developed a fingering annotation model and tool for deriving pseudo fingering annotations from the MIDI and motion data. Lastly, we present MIDI-to-motion generation through a fine-tuning experiment as a use case of the dataset.

cs.SD

Energy Efficiency Optimization in Distributed MIMO vRAN via Cross-Layer Link Abstraction

Virtualized radio access networks (vRAN) run the compute-intensive multiple-input multiple-output (MIMO) baseband as software on shared servers, which makes energy efficiency (EE) a primary design objective. Distributed MIMO vRAN consumes power across virtualized distributed unit (vDU) baseband, fronthaul transport, and per-radio-unit operation. We build a power model that resolves these three components. We then develop a framework that jointly selects modulation, transmission rank, and per-subcarrier power to maximize system EE. Exponential effective SNR mapping induces a convex per-subcarrier power constraint, which yields a convex power minimization problem with a closed-form waterfilling-like solution. We show that radio frequency-only models underestimate the spectral efficiency range where single-input multiple-output (SIMO) transmission saves power, and our power model extends this range by 24%. We further extend the framework to a traffic-aware setting with realistic user trajectories from the multi-agent transport simulator. We propose a traffic-aware strategy that switches each radio unit among MIMO, SIMO, and sleep modes based on demand. Simulation results over 3GPP NR compliant fading channels show that, after a one-time offline calibration, the framework predicts link performance without further link-level simulation. The proposed framework achieves higher average EE than a traffic-agnostic always-on MIMO baseline, while maintaining comparable throughput at peak hours.

eess.SP

Switch-DFT: Adaptive Waveform and MIMO Switching for Energy-Efficient Base Stations

Energy efficiency has emerged as a critical challenge in modern base stations (BSs), as the power amplifier (PA) consumes a substantial portion of the total power due to its limited efficiency. We investigate waveform and mode adaptation to enhance the energy efficiency of BSs. We propose Switch-DFT, an adaptive switching framework that selects between cyclic prefix orthogonal frequency division multiplexing (CP-OFDM) and discrete Fourier transform-spread-OFDM (DFT-s-OFDM) waveforms, as well as between single-input multiple-output (SIMO) and multiple-input multiple-output (MIMO) modes. Switch-DFT improves efficiency by reducing PA backoff with DFT-s-OFDM and achieves the target rate at lower power by leveraging higher MIMO throughput. This results in superior energy efficiency over a wide range of the spectral efficiencies compared with static configurations.

eess.SP

Accelerating vRAN and O-RAN with SIMD: Architectural Perspectives and Performance Evaluation

The evolution of radio access networks (RANs) toward virtualization and openness creates new opportunities for flexible, cost-effective, and high-performance deployments. Achieving real-time and energy-efficient baseband processing on commercial off-the-shelf platforms, however, remains a critical challenge. This article explores how single instruction multiple data (SIMD) architectures can accelerate RAN workloads. We first outline why key physical-layer functions, such as channel estimation, multiple-input multiple-output (MIMO) detection, and forward error correction, are well aligned with SIMD's data-level parallelism. We then present practical design guidelines and prototype results, showing significant improvements in throughput and energy efficiency compared to conventional CPU-only processing, while retaining programmability and ease of integration. Finally, we discuss open challenges in workload balancing and hardware heterogeneity, and highlight the role of SIMD as an enabling technology for flexible, efficient, and sustainable 6G-ready RANs.

eess.SP

PiAnnotate: A Web Annotation Tool for Piano Fingering, with a Diagnostic Probe

Piano fingering shapes how a passage can be played, yet it is difficult to label after a performance. An annotator must decide which finger produced each note while reconciling the score, timing, video, and hand motion. We present PiAnnotate, a web-based pipeline for adding expert fingering annotations to the FurElise performance dataset. The tool brings together a piano-roll view, performance video, and a 3D MANO hand mesh so that reviewers can inspect each assignment in musical and physical context. Rather than storing only the final answer, PiAnnotate keeps paired rule-based and human-edited fingering tracks. These paired tracks make the annotation history auditable by showing where a geometric rule was sufficient, where experts intervened, and how labels changed across review passes. As a final diagnostic, we train a small Transformer probe on the paired tracks. The probe improves on the rule baseline on held-out pieces while remaining conservative about changing labels that were already correct, suggesting that the edited labels contain learnable structure rather than only isolated fixes.

cs.SD

Tipiano: Cascaded Piano Hand Motion Synthesis via Fingertip Priors

Synthesizing realistic piano hand motions requires both precision and naturalness. Physics-based methods achieve precision but produce stiff motions; data-driven models learn natural dynamics but struggle with positional accuracy. Piano motion exhibits a natural hierarchy: fingertip positions are nearly deterministic given piano geometry and fingering, while wrist and intermediate joints offer stylistic freedom. We present [OURS], a four-stage framework exploiting this hierarchy: (1) statistics-based fingertip positioning, (2) FiLM-conditioned trajectory refinement, (3) wrist estimation, and (4) STGCN-based pose synthesis. We contribute expert-annotated fingerings for the FürElise dataset (153 pieces, ~10 hours). Experiments demonstrate F1 = 0.910, substantially outperforming diffusion baselines (F1 = 0.121), with user study (N=41) confirming quality approaching motion capture. Expert evaluation by professional pianists (N=5) identified anticipatory motion as the key remaining gap, providing concrete directions for future improvement.

cs.AI

Precoding Matrix Indicator in the 5G NR Protocol: A Tutorial on 3GPP Beamforming Codebooks

This paper bridges this critical gap by providing a systematic examination of the beamforming codebook technology, i.e., precoding matrix indicator (PMI), in the 5G NR from theoretical, standardization, and implementation perspectives. We begin by introducing the background of beamforming in multiple-input multiple-output (MIMO) systems and the signaling procedures for codebook-based beamforming in practical 5G systems. Then, we establish the fundamentals of regular codebooks and port-selection codebooks in 3GPP standards. Next, we provide rigorous technical analysis of 3GPP codebook evolution spanning Releases 15-18, with particular focus on: 1) We elucidate the core principles underlying codebook design, 2) provide clear physical interpretations for each symbolic variable in the codebook formulas, summarized in tabular form, and 3) offer intuitive visual illustrations to explain how codebook parameters convey information. These essential pedagogical elements are almost entirely absent in the often-obscure standardization documents. Through mathematical modeling, performance benchmarking, feedback comparisons, and scenario-dependent applicability analysis, we provide researchers and engineers with a unified understanding of beamforming codebooks in real-world systems. Furthermore, we identify future directions and other beamforming scenarios for ongoing research and development efforts. This work serves as both an informative tutorial and a guidance for future research, facilitating more effective collaboration between academia and industry in advancing wireless communication technologies.

cs.IT

Designing a Multimodal Viewer for Piano Performance Analysis -- a Pedagogy-First Approach

Abstract instructions in piano education, such as "raise your wrist" and "relax your tension," lead to varying interpretations among learners, preventing instructors from effectively conveying their intended pedagogical guidance. To address this problem, this study conducted systematic interviews with a piano professor with 18 years teaching experience, and two researchers derived seven core need groups through cross-validation. Based on these findings, we developed a web-based dashboard prototype integrating video, motion capture, and musical scores, enabling instructors to provide concrete, visual feedback instead of relying solely on abstract verbal instructions. Technical feasibility was validated through 109 performance datasets.

cs.MM