SearcharxivSearch

arXiv subjects

Gary Scavone

Publications and source records attributed to Gary Scavone.

5 recordsLinked to original sources

Physics-Informed Neural Operator for Speech Production Analysis

Physics-informed neural operators (PINOs) have recently gained attention as fast numerical simulators with potential for solving inverse problems. This study proposes the first PINO-based method for speech production analysis. The model learns the governing one-dimensional wave equations directly without requiring pre-computed supervised training data. Using vocal tract shape data as input features, we compare the proposed model's predicted f0, glottal volume velocity and sound pressure at the lip for five static vowels to a conventional Runge Kutta/Finite difference approach. With errors of 0.8% for glottal volume flow and 3.2% for speech waveforms, the proposed model enables efficient GPU-parallelized simulation without iterative calculations. We conclude that PINO is a promising approach for fast analysis of speech.

cs.SD

Masked Wavelet Scattering Transform Neural Field for Sound Field Reconstruction

In this paper, we propose a reconstruction framework that leverages the Wavelet Scattering Transform (WST) as a multi-scale feature extractor to impose statistical priors under sparse observation conditions. The reconstruction problem is formulated as an optimization task and solved using a neural field, with the WST incorporated into the training loss function. As a proof of concept, we validate the proposed method on HRTF upsampling. A masking strategy is applied to the WST coefficients, resulting in a two-phase procedure. The first phase learns a binary mask from a small multi-subject dataset, while the second phase applies the learned mask to the WST coefficients of an individual HRTF to preserve informative statistical structures during reconstruction. Validation against baseline methods, which also serve as an ablation study of the different components of the framework, demonstrates the effectiveness of the proposed approach.

eess.AS

Physics-Informed Deep Learning for Nonlinear Friction Model of Bow-string Interaction

This study investigates the use of an unsupervised, physics-informed deep learning framework to model a one-degree-of-freedom mass-spring system subjected to a nonlinear friction bow force and governed by a set of ordinary differential equations. Specifically, it examines the application of Physics-Informed Neural Networks (PINNs) and Physics-Informed Deep Operator Networks (PI-DeepONets). Our findings demonstrate that PINNs successfully address the problem across different bow force scenarios, while PI-DeepONets perform well under low bow forces but encounter difficulties at higher forces. Additionally, we analyze the Hessian eigenvalue density and visualize the loss landscape. Overall, the presence of large Hessian eigenvalues and sharp minima indicates highly ill-conditioned optimization. These results underscore the promise of physics-informed deep learning for nonlinear modelling in musical acoustics, while also revealing the limitations of relying solely on physics-based approaches to capture complex nonlinearities. We demonstrate that PI-DeepONets, with their ability to generalize across varying parameters, are well-suited for sound synthesis. Furthermore, we demonstrate that the limitations of PI-DeepONets under higher forces can be mitigated by integrating observation data within a hybrid supervised-unsupervised framework. This suggests that a hybrid supervised-unsupervised DeepONets framework could be a promising direction for future practical applications.

eess.AS

Acoustic Field Reconstruction in Tubes via Physics-Informed Neural Networks

This study investigates the application of Physics-Informed Neural Networks (PINNs) to inverse problems in acoustic tube analysis, focusing on reconstructing acoustic fields from noisy and limited observation data. Specifically, we address scenarios where the radiation model is unknown, and pressure data is only available at the tube's radiation end. A PINNs framework is proposed to reconstruct the acoustic field, along with the PINN Fine-Tuning Method (PINN-FTM) and a traditional optimization method (TOM) for predicting radiation model coefficients. The results demonstrate that PINNs can effectively reconstruct the tube's acoustic field under noisy conditions, even with unknown radiation parameters. PINN-FTM outperforms TOM by delivering balanced and reliable predictions and exhibiting robust noise-tolerance capabilities.

eess.AS

Acoustic Characterization of the Resonator in the Chinese Transverse Flute (dizi)

The dizi is a traditional Chinese transverse flute and is most distinguished from the western flute by the presence of a hole covered by a wrinkled membrane. In this study, we analyze the linear acoustical behavior of the dizi resonator through a detailed acoustical model that incorporates drilled toneholes, back end-holes, membrane hole, and upstream embouchure hole. The input admittance of the dizi is measured and modeled using the Transfer Matrix Method (TMM) and Transfer Matrix Method with external Interactions (TMMI). In comparison to measurements, the TMMI is shown to more accurately model the dizi than the TMM when compared to measurements. Our analysis reveals that attaching the membrane shifts admittance peaks to lower frequencies, reduces their magnitude, and influences tuning and harmonicity for different peaks and fingerings. The study further shows that the upstream branch, which includes the embouchure hole, complicates the evaluation of the tonehole lattice cutoff frequency, suggesting that it may not need to be considered for flute instruments. Cutoff frequencies exhibit distinct groupings across fingerings, influenced by the different tonehole lattices in the dizi: finger-hole lattice and end-hole lattice.

physics.app-ph