SearcharxivSearch

arXiv subjects

Zachary Miller

Publications and source records attributed to Zachary Miller.

13 recordsLinked to original sources

HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection

Building a diagnosis model for primary progressive aphasia (PPA) has been challenging due to the data scarcity. Collecting clinical data at scale is limited by the high vulnerability of clinical population and the high cost of expert labeling. To circumvent this, previous studies simulate dysfluent speech to generate training data. However, those approaches are not comprehensive enough to simulate PPA as holistic, multi-level phenotypes, instead relying on isolated dysfluencies. To address this, we propose a novel, clinically grounded simulation framework, Hierarchical Aphasic Speech Simulation (HASS). HASS aims to simulate behaviors of logopenic variant of PPA (lvPPA) with varying degrees of severity. To this end, semantic, phonological, and temporal deficits of lvPPA are systematically identified by clinical experts, and simulated. We demonstrate that our framework enables more accurate and generalizable detection models.

eess.AS

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling

Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current models failing to capture these discrepancies effectively. We propose a verbatim phoneme recognition framework using multi-task training with novel phoneme similarity modeling that transcribes what speakers actually say rather than what they're supposed to say. We develop and open-source \textit{VCTK-accent}, a simulated dataset containing phonetic errors, and propose two novel metrics for assessing pronunciation differences. Our work establishes a new benchmark for phonetic error detection.

eess.AS

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis

Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting their performance. In this work, we propose Neural LCS, a novel approach for dysfluent text-text and speech-text alignment. Neural LCS addresses key challenges, including partial alignment and context-aware similarity mapping, by leveraging robust phoneme-level modeling. We evaluate our method on a large-scale simulated dataset, generated using advanced data simulation techniques, and real PPA data. Neural LCS significantly outperforms state-of-the-art models in both alignment accuracy and dysfluent speech segmentation. Our results demonstrate the potential of Neural LCS to enhance automated systems for diagnosing and analyzing speech disorders, offering a more accurate and linguistically grounded solution for dysfluent speech alignment.

eess.AS

Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection

Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency generation, existing synthetic datasets suffer from unnatural prosody and limited contextual diversity. To address these limitations, we propose LLM-Dys -- the most comprehensive dysfluent speech corpus with LLM-enhanced dysfluency simulation. This dataset captures 11 dysfluency categories spanning both word and phoneme levels. Building upon this resource, we improve an end-to-end dysfluency detection framework. Experimental validation demonstrates state-of-the-art performance. All data, models, and code are open-sourced at https://github.com/Berkeley-Speech-Group/LLM-Dys.

eess.AS

Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection

Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems.

eess.AS

Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection

Speech dysfluency modeling is a task to detect dysfluencies in speech, such as repetition, block, insertion, replacement, and deletion. Most recent advancements treat this problem as a time-based object detection problem. In this work, we revisit this problem from a new perspective: tokenizing dysfluencies and modeling the detection problem as a token-based automatic speech recognition (ASR) problem. We propose rule-based speech and text dysfluency simulators and develop VCTK-token, and then develop a Whisper-like seq2seq architecture to build a new benchmark with decent performance. We also systematically compare our proposed token-based methods with time-based methods, and propose a unified benchmark to facilitate future research endeavors. We open-source these resources for the broader scientific community. The project page is available at https://rorizzz.github.io/

eess.AS

YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection

Dysfluent speech detection is the bottleneck for disordered speech analysis and spoken language learning. Current state-of-the-art models are governed by rule-based systems which lack efficiency and robustness, and are sensitive to template design. In this paper, we propose YOLO-Stutter: a first end-to-end method that detects dysfluencies in a time-accurate manner. YOLO-Stutter takes imperfect speech-text alignment as input, followed by a spatial feature aggregator, and a temporal dependency extractor to perform region-wise boundary and class predictions. We also introduce two dysfluency corpus, VCTK-Stutter and VCTK-TTS, that simulate natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation. Our end-to-end method achieves state-of-the-art performance with a minimum number of trainable parameters for on both simulated data and real aphasia speech. Code and datasets are open-sourced at https://github.com/rorizzz/YOLO-Stutter

eess.AS

Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection

Current de-facto dysfluency modeling methods utilize template matching algorithms which are not generalizable to out-of-domain real-world dysfluencies across languages, and are not scalable with increasing amounts of training data. To handle these problems, we propose Stutter-Solver: an end-to-end framework that detects dysfluency with accurate type and time transcription, inspired by the YOLO object detection algorithm. Stutter-Solver can handle co-dysfluencies and is a natural multi-lingual dysfluency detector. To leverage scalability and boost performance, we also introduce three novel dysfluency corpora: VCTK-Pro, VCTK-Art, and AISHELL3-Pro, simulating natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation through articulatory-encodec and TTS-based methods. Our approach achieves state-of-the-art performance on all available dysfluency corpora. Code and datasets are open-sourced at https://github.com/eureka235/Stutter-Solver

eess.AS

Lower Bound for Independence Covering in $C_4$-Free Graphs

An independent set in a graph $G$ is a set $S$ of pairwise non-adjacent vertices in $G$. A family $\mathcal{F}$ of independent sets in $G$ is called a $k$-independence covering family if for every independent set $I$ in $G$ of size at most $k$, there exists an $S \in \mathcal{F}$ such that $I \subseteq S$. Lokshtanov et al. [ACM Transactions on Algorithms, 2018] showed that graphs of degeneracy $d$ admit $k$-independence covering families of size $\binom{k(d+1)}{k} \cdot 2^{o(kd)} \cdot \log n$, and used this result to design efficient parameterized algorithms for a number of problems, including STABLE ODD CYCLE TRANSVERSAL and STABLE MULTICUT. In light of the results of Lokshtanov et al. it is quite natural to ask whether even more general families of graphs admit $k$-independence covering families of size $f(k)n^{O(1)}$. Graphs that exclude a complete bipartite graph $K_{d+1,d+1}$ with $d+1$ vertices on both sides as a subgraph, called $K_{d+1,d+1}$-free graphs, are a frequently considered generalization of $d$-degenerate graphs. This motivates the question whether $K_{d,d}$-free graphs admit $k$-independence covering families of size $f(k,d)n^{O(1)}$. Our main result is a resounding "no" to this question -- specifically we prove that even $K_{2,2}$-free graphs (or equivalently $C_4$-free graphs) do not admit $k$-independence covering families of size $f(k)n^{\frac{k}{4}-ε}$.

cs.DM

Motion Compensated Self Supervised Deep Learning for Highly Accelerated 3D Ultrashort Echo Time Pulmonary MRI

Purpose: To investigate motion compensated, self-supervised, model based deep learning (MBDL) as a method to reconstruct free breathing, 3D Pulmonary ultrashort echo time (UTE) acquisitions. Theory and Methods: A self-supervised eXtra Dimension MBDL architecture (XD-MBDL) was developed that combined respiratory states to reconstruct a single high-quality 3D image. Non-rigid, GPU based motion fields were incorporated into this architecture by estimating motion fields from a low resolution motion resolved (XD-GRASP) iterative reconstruction. Motion Compensated XD-MBDL was evaluated on lung UTE datasets with and without contrast and was compared to constrained reconstructions and variants of self-supervised MBDL that do not consider respiratory motion. Results: Images reconstructed using XD-MBDL demonstrate improved image quality as measured by apparent SNR, CNR and visual assessment relative to self-supervised MBDL approaches that do not account for dynamic respiratory states, XD-GRASP and a recently proposed motion compensated iterative reconstruction strategy (iMoCo). Additionally, XD-MBDL reduced reconstruction time relative to both XD-GRASP and iMoCo. Conclusion: A method was developed to allow self-supervised MBDL to combine multiple respiratory states to reconstruct a single image. This method was combined with GPU-based image registration to further improve reconstruction quality. This approach showed promising results reconstructing a user-selected respiratory phase from free breathing 3D pulmonary UTE acquisitions.

physics.med-ph

Memory Efficient Model Based Deep Learning Reconstructions for High Spatial Resolution 3D Non-Cartesian Acquisitions

Objective: Model based deep learning (MBDL) has been challenging to apply to the reconstruction of 3D non-Cartesian MRI acquisitions due to extreme GPU memory demand (>250 GB using traditional backpropagation) primarily because the entire volume is needed for data-consistency steps embedded in the model. The goal of this work is to develop and apply a memory efficient method called block-wise learning that combines gradient checkpointing with patch-wise training to allow for fast and high-quality 3D non-Cartesian reconstructions using MBDL. Approach: Block-wise learning applied to a single unroll decomposes the input volume into smaller patches, gradient checkpoints each patch, passes each patch iteratively through a neural network regularizer, and then rebuilds the full volume from these output patches for data-consistency. This method is applied across unrolls during training. Block-wise learning significantly reduces memory requirements by tying GPU memory to user selected patch size instead of the full volume. This algorithm was used to train a MBDL architecture to reconstruct highly undersampled, 1.25mm isotropic, pulmonary magnetic resonance angiography volumes with matrix sizes varying from 300-450 x 200-300 x 300-450 on a single GPU. We compared block-wise learning reconstructions against L1 wavelet compressed reconstructions and proxy ground truth images. Main results: MBDL with block-wise learning significantly improved image quality relative to L1 wavelet compressed sensing while simultaneously reducing average reconstruction time 38x. Significance: Block-wise learning allows for MBDL to be applied to high spatial resolution, 3D non-Cartesian datasets with improved image quality and significant reductions in reconstruction time relative to traditional iterative methods

physics.med-ph

Motion Compensated Extreme MRI: Multi-Scale Low Rank Reconstructions for Highly Accelerated 3D Dynamic Acquisitions (MoCo-MSLR)

Purpose: To improve upon Extreme MRI, a recently proposed method by Ong Et al. for reconstructing high spatiotemporal resolution, 3D non-Cartesian acquisitions by incorporating motion compensation into these reconstructions using an approach termed MoCo-MSLR. Methods: Motion compensation is challenging to incorporate into high spatiotemporal resolution reconstruction due to the memory footprint of the motion fields and the potential to lose dynamics by relying on an initial high temporal resolution, low spatial resolution reconstruction. Motivated by the work of Ong Et al. and Huttinga Et al., we estimate low spatial resolution motion fields through a loss enforced in k-space and represent these motion fields in a memory efficient manner using multi-scale low rank components. We interpolate these motion fields to the desired spatial resolution, and then incorporate these fields into Extreme MRI. Results: MoCo-MSLR was able to improve image quality for reconstructions around 500ms temporal resolution and capture bulk motion not seen in Extreme MRI. Further, MoCo-MSLR was able to resolve realistic cardiac dynamics at near 100ms temporal resolution while Extreme MRI struggled to resolve these dynamics. Conclusion: MoCo-MSLR improved image quality over Extreme MRI and was able to resolve both respiratory and cardiac motion in 3D.

physics.med-ph

Tempered fractional Brownian motion on finite intervals

Diffusive transport in many complex systems features a crossover between anomalous diffusion at short times and normal diffusion at long times. This behavior can be mathematically modeled by cutting off (tempering) beyond a mesoscopic correlation time the power-law correlations between the increments of fractional Brownian motion. Here, we investigate such tempered fractional Brownian motion confined to a finite interval by reflecting walls. Specifically, we explore how the tempering of the long-time correlations affects the strong accumulation and depletion of particles near reflecting boundaries recently discovered for untempered fractional Brownian motion. We find that exponential tempering introduces a characteristic size for the accumulation and depletion zones but does not affect the functional form of the probability density close to the wall. In contrast, power-law tempering leads to more complex behavior that differs between the superdiffusive and subdiffusive cases.

cond-mat.stat-mech