SearcharxivSearch

arXiv subjects

Brad Story

Publications and source records attributed to Brad Story.

3 recordsLinked to original sources

Towards detecting the pathological subharmonic voicing with fully convolutional neural networks

Many voice disorders induce subharmonic phonation, but voice signal analysis is currently lacking a technique to detect the presence of subharmonics reliably. Distinguishing subharmonic phonation from normal phonation is a challenging task as both are nearly periodic phenomena. Subharmonic phonation adds cyclical variations to the normal glottal cycles. Hence, the estimation of subharmonic period requires a wholistic analysis of the signals. Deep learning is an effective solution to this type of complex problem. This paper describes fully convolutional neural networks which are trained with synthesized subharmonic voice signals to classify the subharmonic periods. Synthetic evaluation shows over 98% classification accuracy, and assessment of sustained vowel recordings demonstrates encouraging outcomes as well as the areas for future improvements.

eess.AS

Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals

In this paper, we propose a new method for the accurate estimation and tracking of formants in speech signals using time-varying quasi-closed-phase (TVQCP) analysis. Conventional formant tracking methods typically adopt a two-stage estimate-and-track strategy wherein an initial set of formant candidates are estimated using short-time analysis (e.g., 10--50 ms), followed by a tracking stage based on dynamic programming or a linear state-space model. One of the main disadvantages of these approaches is that the tracking stage, however good it may be, cannot improve upon the formant estimation accuracy of the first stage. The proposed TVQCP method provides a single-stage formant tracking that combines the estimation and tracking stages into one. TVQCP analysis combines three approaches to improve formant estimation and tracking: (1) it uses temporally weighted quasi-closed-phase analysis to derive closed-phase estimates of the vocal tract with reduced interference from the excitation source, (2) it increases the residual sparsity by using the $L_1$ optimization and (3) it uses time-varying linear prediction analysis over long time windows (e.g., 100--200 ms) to impose a continuity constraint on the vocal tract model and hence on the formant trajectories. Formant tracking experiments with a wide variety of synthetic and natural speech signals show that the proposed TVQCP method performs better than conventional and popular formant tracking tools, such as Wavesurfer and Praat (based on dynamic programming), the KARMA algorithm (based on Kalman filtering), and DeepFormants (based on deep neural networks trained in a supervised manner). Matlab scripts for the proposed method can be found at: https://github.com/njaygowda/ftrack

eess.AS

Prospectively accelerated dynamic speech MRI at 3 Tesla using a self-navigated spiral based manifold regularized scheme

This work proposes a self-navigated variable density spiral(VDS) based manifold regularization scheme to prospectively improve dynamic speech MRI at 3T. Short readout 1.3ms spirals were used to minimize off-resonance. A custom 16-channel speech coil was used for improved parallel imaging of vocal tract. The manifold model leveraged similarities between frames sharing similar speech postures without explicit motion binning. The self-navigating capability of VDS was leveraged to learn the Laplacian matrix of the manifold. Reconstruction was posed as a SENSE-based non-local soft weighted temporal regularization scheme. Our approach was compared against view-sharing, low-rank, finite difference, extra-dimension-based sparsity reconstruction constraints. Under-sampling experiments were conducted on five volunteers performing repetitive and arbitrary speaking tasks at different speaking rates. Quantitative evaluation in terms of mean square error over moving edges were performed in a retrospectively under-sampled data. For prospective under-sampling, blinded image quality evaluation in the categories of alias artifacts, spatial blurring, and temporal blurring were performed by three voice research experts. Region of interest analysis at articulator boundaries were performed to assess articulatory motion. Our scheme provided improved reconstruction over the others. With prospective under-sampling, a spatial resolution of 2.4mm2/pixel and a temporal resolution of 17.4 ms/frame for single slice imaging, and 52.2 ms/frame for 3-slice imaging were achieved. We demonstrated implicit motion binning by analyzing the mechanics of the Laplacian matrix. Our method demonstrated superior image quality scores in reducing spatial and temporal blurring. While it exhibited faint alias artifacts similar to temporal finite-difference, it provided statistically significant improvements over remaining constraints.

eess.IV