SearcharxivSearch

arXiv subjects

Rajeev Nongpiur

Publications and source records attributed to Rajeev Nongpiur.

3 recordsLinked to original sources

MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

Recent advances in multimodal audio generation have enabled music synthesis from text, visual cues, and other high-level conditions. However, most systems are designed for a single operating mode: either generating music without a reference mixture or extracting a target source from an existing mixture. This fixed-task design limits their use when different combinations of text, visual, and mixture inputs are available. To address this gap, we propose MAGE, a modality-agnostic framework for conditional music generation and mixture-grounded target-source extraction within a shared continuous latent space. Our approach introduces three key components. First, a Controlled Multimodal FluxFormer models the conditional flow from noise to a target audio latent, enabling the same backbone to operate with or without a mixture condition. Second, Audio-Visual Nexus Alignment maps frame-level visual features onto the audio latent sequence, allowing visual evidence to condition the generation process at the audio-token level. Third, a cross-gated modulation mechanism uses the aligned visual representation to regulate intermediate audio features, while text provides separate semantic guidance. We further train MAGE with dynamic modality masking, exposing the same model to text-only, visual-only, joint text-visual, mixture-conditioned, and unconditional configurations. Experiments on the MUSIC benchmark evaluate MAGE under separate protocols for mixture-free generation and mixture-grounded target-source extraction. The results show that MAGE provides a shared conditioning interface across both settings, and that the proposed alignment and gating components improve interference suppression in the extraction task.

cs.SD

Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering

In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time adaptation to diverse physical scenes, causing a sensory mismatch between visual and auditory cues that disrupts user immersion. To address this, we introduce SAMOSA, a novel on-device system that renders spatially accurate sound by dynamically adapting to its physical environment. SAMOSA leverages a synergistic multimodal scene representation by fusing real-time estimations of room geometry, surface materials, and semantic-driven acoustic context. This rich representation then enables efficient acoustic calibration via scene priors, allowing the system to synthesize a highly realistic Room Impulse Response (RIR). We validate our system through technical evaluation using acoustic metrics for RIR synthesis across various room configurations and sound types, alongside an expert evaluation (N=12). Evaluation results demonstrate SAMOSA's feasibility and efficacy in enhancing XR auditory realism.

cs.HC

Soli-enabled Noncontact Heart Rate Detection for Sleep and Meditation Tracking

Heart rate (HR) is a crucial physiological signal that can be used to monitor health and fitness. Traditional methods for measuring HR require wearable devices, which can be inconvenient or uncomfortable, especially during sleep and meditation. Noncontact HR detection methods employing microwave radar can be a promising alternative. However, the existing approaches in the literature usually use high-gain antennas and require the sensor to face the user's chest or back, making them difficult to integrate into a portable device and unsuitable for sleep and meditation tracking applications. This study presents a novel approach for noncontact HR detection using a miniaturized Soli radar chip embedded in a portable device (Google Nest Hub). The chip has a $6.5 \mbox{ mm} \times 5 \mbox{ mm} \times 0.9 \mbox{ mm}$ dimension and can be easily integrated into various devices. The proposed approach utilizes advanced signal processing and machine learning techniques to extract HRs from radar signals. The approach is validated on a sleep dataset (62 users, 498 hours) and a meditation dataset (114 users, 1131 minutes). The approach achieves a mean absolute error (MAE) of $1.69$ bpm and a mean absolute percentage error (MAPE) of $2.67\%$ on the sleep dataset. On the meditation dataset, the approach achieves an MAE of $1.05$ bpm and a MAPE of $1.56\%$. The recall rates for the two datasets are $88.53\%$ and $98.16\%$, respectively. This study represents the first application of the noncontact HR detection technology to sleep and meditation tracking, offering a promising alternative to wearable devices for HR monitoring during sleep and meditation.

eess.SP