SearcharxivSearch

arXiv subjects

Stephanie Dodson

Publications and source records attributed to Stephanie Dodson.

9 recordsLinked to original sources

No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

Supplying context at inference time to a large multimodal model is an inexpensive lever for adapting speech transcription to a domain, and earlier results on smaller models reported large gains. This work tested that mechanism where it ships, in the prompt-conditioning layer of a production oral-history transcription tool, on a sample from its own production corpus. Full prompt-level context did not detectably change side-level word error rate (WER), and none of the four preregistered hypotheses was supported. The design was a within-item paired ablation, preregistered with the analysis code frozen by hash before the confirmatory batch was scored; two disclosed gpt-4o pilot sides had been scored earlier, during scorer development. Nineteen cassette sides, about 10.6 hours of degraded 1970s-80s interview audio, were reprocessed through the production code path under three prompt arms, crossed with two deployed commercial configurations, gpt-4o-transcribe and gemini-2.5-flash, and scored against operator-corrected verbatim references. For gpt-4o-transcribe the median paired difference between the full-context and no-context arms was +0.6 WER points, with a side-resampled interval of [-1.1, +1.0]; the Gemini estimates were too unstable to support a comparable negative inference. A post-hoc rerun found run-to-run pipeline variability larger than the confirmatory differences, so effects of that size cannot be resolved from one transcription per cell. An implementation audit verified the manipulation was live, and sequence-alignment analysis found a small improvement on complete context-listed phrases, too small to materially change side-level WER, and for Gemini coexisting with worsened unlisted-token error. Evaluating context mechanisms therefore requires sequence-aligned term-level, insertion, and speaker-label measures alongside aggregate accuracy.

cs.CL

Kernel-Dependent Pattern Formation in a Population Model with Nonlocal Facilitation and Competition

Spatial patterns, such as those in dryland vegetation models, have historically been studied in systems of reaction-diffusion systems with pattern onset via a Turing bifurcation from a spatially uniform state. More recently, spatial patterns have been considered in models that incorporate spatially extended interactions via nonlocal interaction kernels. It remains largely underexplored if and how the choice of nonlocal interaction kernel contributes to differences in pattern formation and persistence, particularly in models that contain competition and facilitation. Here, we investigate spatial patterns in a reaction-diffusion model for a single species that includes nonlocal competition and facilitation processes; Gaussian, exponential, algebraic, hat, and smooth hat kernels are considered as specific examples. Via a center manifold analysis, and using the relative spatial scale of competition to facilitation and the death rate as bifurcation parameters, we identify that the choice of kernel has impacts on the pattern forming bifurcation. Bifurcations using the Gaussian, exponential, and algebraic kernels largely follow expectations of Turing patterns, but patterns in the hat and smooth hat kernels can form even when the scale of competition is less than that of facilitation. The dynamics of patterns far from onset are investigated via numerical continuation methods. The model produces the so-called "Turing-before-Tipping" phenomenon demonstrating that the arrangement into spatial patterns is an effective resilience mechanism against harsh conditions. Again, there is a kernel-dependent dichotomy in pattern behavior. Early warning signs for population extinction are observed with the Gaussian, exponential, and algebraic kernels, but not under the hat or smooth hat cases.

math.DS

Efficient numerical computation of spiral spectra with exponentially-weighted preconditioners

The stability of nonlinear waves on spatially extended domains is commonly probed by computing the spectrum of the linearization of the underlying PDE about the wave profile. It is known that convective transport, whether driven by the nonlinear pattern itself or an underlying fluid flow, can cause exponential growth of the resolvent of the linearization as a function of the domain length. In particular, sparse eigenvalue algorithms may result in inaccurate and spurious spectra in the convective regime. In this work, we focus on spiral waves, which arise in many natural processes and which exhibit convective transport. We prove that exponential weights can serve as effective, inexpensive preconditioners that result in resolvents that are uniformly bounded in the domain size and that stabilize numerical spectral computations. We also show that the optimal exponential rates can be computed reliably from a simpler asymptotic problem posed in one space dimension.

math.NA

Behavior of Spiral Wave Spectra with a Rank-Deficient Diffusion Matrix

Spiral waves emerge in numerous pattern forming systems and are commonly modeled with reaction-diffusion systems. Some systems used to model biological processes, such as ion-channel models, fall under the reaction-diffusion category and often have one or more non-diffusing species which results in a rank-deficient diffusion matrix. Previous theoretical research focused on spiral spectra for strictly positive diffusion matrices. In this paper, we use a general two-variable reaction-diffusion system to compare the essential and absolute spectra of spiral waves for strictly positive and rank-deficient diffusion matrices. We show that the essential spectrum is not continuous in the limit of vanishing diffusion in one component. Moreover, we predict locations for the absolute spectrum in the case of a non-diffusing slow variable. Predictions are confirmed numerically for the Barkley and Karma models.

math.DS

Reflections in excitable media linked to existence and stability of one-dimensional spiral waves

When propagated action potentials in cardiac tissue interact with local heterogeneities, reflected waves can sometimes be induced. These reflected waves have been associated with the onset of cardiac arrhythmias, and while their generation is not well understood, their existence is linked to that of one-dimensional (1D) spiral waves. Thus, understanding the existence and stability of 1D spirals plays a crucial role in determining the likelihood of the unwanted reflected pulses. Mathematically, we probe these issues by viewing the 1D spiral as a time-periodic antisymmetric source defect. Through a combination of direct numerical simulation and continuation methods, we investigate existence and stability of a 1D spiral wave in a qualitative ionic model to determine how the systems propensity for reflections are influenced by system parameters. Our results support and extend a previous hypothesis that the 1D spiral is an unstable periodic orbit that emerges through a global rearrangement of heteroclinic orbits and we identify key parameters and physiological processes that promote and deter reflection behavior.

math.DS

Determining the source of period-doubling instabilities in spiral waves

Spiral wave patterns observed in models of cardiac arrhythmias and chemical oscillations develop alternans and stationary line defects, which can both be thought of as period-doubling instabilities. These instabilities are observed on bounded domains, and may be caused by the spiral core, far-field asymptotics, or boundary conditions. Here, we introduce a methodology to disentangle the impacts of each region on the instabilities by analyzing spectral properties of spiral waves and boundary sinks on bounded domains with appropriate boundary conditions. We apply our techniques to spirals formed in reaction-diffusion systems to investigate how and why alternans and line defects develop. Our results indicate that the mechanisms driving these instabilities are quite different; alternans are driven by the spiral core, whereas line defects appear from boundary effects. Moreover, we find that the shape of the alternans eigenfunction is due to the interaction of a point eigenvalue with curves of continuous spectra.

math.DS

Using a Big Data Database to Identify Pathogens in Protein Data Space

Current metagenomic analysis algorithms require significant computing resources, can report excessive false positives (type I errors), may miss organisms (type II errors / false negatives), or scale poorly on large datasets. This paper explores using big data database technologies to characterize very large metagenomic DNA sequences in protein space, with the ultimate goal of rapid pathogen identification in patient samples. Our approach uses the abilities of a big data databases to hold large sparse associative array representations of genetic data to extract statistical patterns about the data that can be used in a variety of ways to improve identification algorithms.

cs.DB

Rapid Sequence Identification of Potential Pathogens Using Techniques from Sparse Linear Algebra

The decreasing costs and increasing speed and accuracy of DNA sample collection, preparation, and sequencing has rapidly produced an enormous volume of genetic data. However, fast and accurate analysis of the samples remains a bottleneck. Here we present D$^{4}$RAGenS, a genetic sequence identification algorithm that exhibits the Big Data handling and computational power of the Dynamic Distributed Dimensional Data Model (D4M). The method leverages linear algebra and statistical properties to increase computational performance while retaining accuracy by subsampling the data. Two run modes, Fast and Wise, yield speed and precision tradeoffs, with applications in biodefense and medical diagnostics. The D$^{4}$RAGenS analysis algorithm is tested over several datasets, including three utilized for the Defense Threat Reduction Agency (DTRA) metagenomic algorithm contest.

q-bio.QM

Genetic Sequence Matching Using D4M Big Data Approaches

Recent technological advances in Next Generation Sequencing tools have led to increasing speeds of DNA sample collection, preparation, and sequencing. One instrument can produce over 600 Gb of genetic sequence data in a single run. This creates new opportunities to efficiently handle the increasing workload. We propose a new method of fast genetic sequence analysis using the Dynamic Distributed Dimensional Data Model (D4M) - an associative array environment for MATLAB developed at MIT Lincoln Laboratory. Based on mathematical and statistical properties, the method leverages big data techniques and the implementation of an Apache Acculumo database to accelerate computations one-hundred fold over other methods. Comparisons of the D4M method with the current gold-standard for sequence analysis, BLAST, show the two are comparable in the alignments they find. This paper will present an overview of the D4M genetic sequence algorithm and statistical comparisons with BLAST.

q-bio.QM