SearcharxivSearch

arXiv subjects

Michael Wilson

Publications and source records attributed to Michael Wilson.

At least 19 recordsLinked to original sources

A face-isolation proof of the two-variable Gaussian Moments Conjecture

Let $(X,Y)$ be independent standard real Gaussian random variables, and let $(P\in\mathbb C[X,Y])$ satisfy $\mathbb E!\left(P(X,Y)^m\right)=0$ for every $(m\ge1)$. Using the complex coordinates $Z=\frac{X+iY}{\sqrt2}, \qquad W=\frac{X-iY}{\sqrt2},$ we prove that the monomial support of $(P)$ is strictly one-sided with respect to the weight $\operatorname{wt}(Z^aW^b)=a-b$. Thus either every monomial occurring in $(P)$ has positive weight, or every monomial has negative weight. It follows that, for every $(Q\in\mathbb C[X,Y])$, $\mathbb E!\left(Q(X,Y)P(X,Y)^m\right)=0$ whenever $(m>\deg Q)$. This proves the two-variable Gaussian Moments Conjecture with the explicit threshold $m\ge\deg Q+1$. The main ingredient is a prime-isolation theorem for exposed faces of the Newton polygon of $(P)$. A $(p)$-adic valuation argument at moment indices of the form $(m=qp)$ separates the contribution of a chosen face from all remaining multinomial strata. Frobenius reduction then forces the constant terms of all positive powers of the associated one-variable Laurent polynomial to vanish. The theorem of Duistermaat and van der Kallen implies that the face has weights of one strict sign, while a planar convex-geometric argument rules out support containing weights of both signs. Combined with the known one-variable case and counterexamples in dimensions $(n\ge3)$, this determines the dimensions in which the Gaussian Moments Conjecture holds.

math.PR

Power-Law Approach of the Stress-Energy Tensor to the Unruh State after Gravitational Collapse

We establish the rate at which the renormalized stress--energy tensor of a massless minimally coupled scalar field in the in-vacuum state of a collapsing null-shell spacetime approaches the corresponding Unruh-state value. At finite exterior radius, we establish the upper bound \[ |\Delta\langle T_{\mu\nu}\rangle|\leq C(r)\,t_s^{-3} \] from the Cauchy-surface decomposition of the Hadamard difference and the branch-cut structure of the retarded Green function. At future null infinity, we show that the leading coefficient in the late-time expansion \[ \Delta\langle T_{uu}\rangle\sim C_{uu}\,u_s^{-3} \] is nonzero, by computing the branch-cut residue explicitly at small frequency and using the Planck suppression of the thermal spectrum at large frequency to show that the dominant contribution to $C_{uu}$ has a definite sign. The result gives \[ \Delta\langle T_{uu}\rangle\big|_{\Iscr^+}(u_s) \sim C_{uu}\,u_s^{-3}, \qquad u_s\to\infty, \] with $C_{uu}\neq 0$. The exponent is determined by the $\omega^2\ln\omega$ branch-point singularity in the Wronskian of the $\ell=0$ radial wave equation, the same structure responsible for Price's law. The sign $C_{uu}<0$ is supported by a physical argument and by the numerical mode data of Gholizadeh Siahmazgi, Anderson, and Fabbri. The result confirms their conjecture that the approach is a power law. We conjecture that the same mechanism gives an analogous $t_s^{-7}$ bound for gravitational perturbations ($\ell_{\min}=2$), though the extension to the spin-2 case involves gauge issues not addressed here.

gr-qc

Olmo 3

We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets long-context reasoning, function calling, coding, instruction following, general chat, and knowledge recall. This release includes the entire model flow, i.e., the full lifecycle of the family of models, including every stage, checkpoint, data point, and dependency used to build it. Our flagship model, Olmo 3 Think 32B, is the strongest fully-open thinking model released to-date.

cs.CL

Infrared Universality: The $r^{-3}$ Spectral Threshold for Coupled Gravitational and Electromagnetic Fields

We identify the $r^{-3}$ curvature-decay rate as a universal geometric threshold separating compact from non-compact perturbations of Laplace-type operators on asymptotically flat manifolds. For the coupled Einstein--Maxwell system, we prove that the linearized operator $\mathcal{L}$ is essentially self-adjoint and that curvature and field strengths decaying faster than $r^{-3}$ act as relatively compact perturbations, while decay exactly at $r^{-3}$ places $0\in\sigma_{\mathrm{ess}}(\mathcal{L})$ through delocalized zero modes. This threshold mechanism unifies the infrared behavior of spin-1, spin-2, and mixed spin-$(1\oplus2)$ fields, linking the onset of spectral delocalization with the appearance of gravitational and electromagnetic memory. Finite-difference simulations corroborate the analytic scaling and reproduce the characteristic quadrupolar and dipolar sky maps predicted for the coupled memory fields. These results demonstrate that curvature decay at $r^{-3}$ constitutes a fundamental geometric boundary underlying infrared universality in gauge and gravitational theories, providing a spectral counterpart to the asymptotic-symmetry and soft-theorem formulations of memory.

gr-qc

Threshold Resolvent Singularities and the Infrared Structure of Linearized Gravity

We identify a sharp geometric threshold governing the infrared spectral behavior of the spatial Lichnerowicz operator on asymptotically flat three-dimensional manifolds. Let $(M,g)$ be asymptotically flat and let $L=\Delta_L$ denote the spatial Lichnerowicz operator acting on symmetric $2$-tensors. Assume \[ |{\rm Riem}(x)| \lesssim r(x)^{-p} \quad \text{as } r(x)\to\infty. \] If $p>3$, curvature is spectrally short-range: $L$ exhibits regular low-energy scattering and zero energy is not singular. At the critical decay \[ |{\rm Riem}(x)| \sim r^{-3}, \] dispersion and curvature balance. Zero enters the essential spectrum, and the weighted resolvent develops a threshold singularity. For $s\in(1/2,1)$, \[ \|\langle r\rangle^{-s}(L-i\varepsilon)^{-1}\langle r\rangle^{-s}\| \gtrsim \varepsilon^{-(1-s)} \quad \text{as } \varepsilon \downarrow 0 . \] Thus, the limiting absorption principle fails at zero energy. This singularity provides a spatial spectral mechanism for the infrared sector of linearized gravity. The same inverse-cube scaling governs long-range correlations, irregular low-frequency scattering, and soft gravitational modes. Numerical simulations of a radial model and the full tensor operator confirm that $p=3$ marks a sharp transition between negligible and marginal curvature. The associated branch point at zero energy determines late-time relaxation, yielding the universal tail exponent \[ t^{-(2\ell+3)}, \] a spectral consequence of nonzero ADM mass. More generally, in $d$ spatial dimensions, the critical decay \[ |{\rm Riem}(x)| \sim r^{-d} \] forms a universal boundary for curvature-coupled Laplace-type operators, encoding the infrared structure of gravity in the spectral geometry of a Cauchy slice.

gr-qc

Curvature Decay and the Spectrum of the Non-Abelian Laplacian on $\mathbb{R}^3$

I study the spectral behavior of the covariant Laplacian $\Delta_A = d_A^* d_A$ associated with smooth $\mathrm{SU}(2)$ connections on $\mathbb{R}^3$. The main result establishes a sharp threshold for the pointwise decay of curvature governing the essential spectrum of $\Delta_A$. Specifically, if the curvature satisfies the bound $|F_A(x)| \le C(1 + |x|)^{-3-\varepsilon}$ for some $\varepsilon > 0$, then $\Delta_A$ is a relatively compact perturbation of the flat Laplacian and hence $\sigma_{\mathrm{ess}}(\Delta_A) = [0,\infty)$. At the critical decay rate $|F_A(x)| \sim |x|^{-3}$, I construct a smooth connection for which $0 \in \sigma_{\mathrm{ess}}(\Delta_A)$, showing that the threshold is sharp. Moreover, a genuinely non-Abelian example based on the hedgehog ansatz is given to demonstrate that the commutator term $A \wedge A$ contributes at the same order. This work identifies the exact decay rate separating stable preservation of the essential spectrum from the onset of delocalized modes in the non-Abelian setting, providing a counterpart to classical results on magnetic Schr\"odinger operators.

math-ph

DESI DR2 Galaxy Luminosity Functions

We present galaxy luminosity functions (LFs) for the Dark Energy Spectroscopic Instrument (DESI) DR2 Bright Galaxy Survey (BGS) in the g,r,z, and w1 bands over 0.002 -15, which is stronger for red galaxies than for blue. We show that our LFs are largely complete for galaxies with surface brightness mu_50<25, and that an apparent steepening fainter than -13 is driven primarily by local overdensity and fragmentation of large galaxies. A systematic North-South offset at the brightest magnitudes is traced to red galaxies and may reflect shallower North photometry underestimating extended early-type profiles, although this remains inconclusive. We therefore also provide LFs based on model-Petrosian magnitudes. Redshift-splitting reveals small but significant residuals, indicating limitations of a simple global evolutionary model. Using the redshift limits of Loveday (2011) we find excellent agreement with GAMA, with substantially reduced statistical errors. These measurements provide a precise reference for studies of environmental and population-dependent LFs and for testing galaxy formation models.

astro-ph.GA

DESI Spectroscopy of HETDEX Emission-line Candidates I: Line Discrimination Validation

The Hobby-Eberly Dark Energy Experiment (HETDEX) is an untargeted spectroscopic galaxy survey that uses Ly$\alpha$ emitting galaxies (LAEs) as tracers of 1.9 < z < 3.5 large scale structure. Most detections consist of a single emission line, whose identity is inferred via a Bayesian analysis of ancillary data. To determine the accuracy of these line identifications, HETDEX detections were observed with the Dark Energy Spectroscopic Instrument (DESI). In two DESI pointings, high confidence spectroscopic redshifts are obtained for 1157 sources, including 982 LAEs. The DESI spectra are used to evaluate the accuracy of the HETDEX object classifications, and tune the methodology to achieve the HETDEX science requirement of $\lesssim 2\%$ contamination of the LAE sample by low-redshift emission-line galaxies, while still assigning $96\%$ of the true Ly$\alpha$ emission sample with the correct spectroscopic redshift. We compare emission line measurements between the two experiments assuming a simple Gaussian line fitting model. Fitted values for the central wavelength of the emission line, the measured line flux and line widths are consistent between the surveys within uncertainties. Derived spectroscopic redshifts, from the two classification pipelines, when both agree as an LAE classification, are consistent to within $\langle \Delta z / (1 + z) \rangle = 6.9\times 10^{-5}$ with an rms scatter of $3.3\times 10^{-4}$.

astro-ph.CO

2 OLMo 2 Furious

We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model weights, full training data, training code and recipes, training logs and thousands of intermediate checkpoints. In this work, we describe our modified model architecture and training recipe, focusing on techniques for achieving better training stability and improved per-token efficiency. Our updated pretraining data mixture introduces a new, specialized data mix called Dolmino Mix 1124, which significantly improves model capabilities across many downstream task benchmarks when introduced via late-stage curriculum training (i.e. specialized data during the annealing phase of pretraining). Finally, we incorporate best practices from T\"ulu 3 to develop OLMo 2-Instruct, focusing on permissive data and extending our final-stage reinforcement learning with verifiable rewards (RLVR). Our OLMo 2 base models sit at the Pareto frontier of performance to training compute, often matching or outperforming open-weight only models like Llama 3.1, Qwen 2.5, and Gemma 2 while using fewer FLOPs and with fully transparent training data, code, and recipe. Our fully open OLMo 2-Instruct models are competitive with open-weight only models of comparable size and even some proprietary models like GPT-3.5 Turbo and GPT 4o Mini.

cs.CL

Fused Gromov-Wasserstein Variance Decomposition with Linear Optimal Transport

Wasserstein distances form a family of metrics on spaces of probability measures that have recently seen many applications. However, statistical analysis in these spaces is complex due to the nonlinearity of Wasserstein spaces. One potential solution to this problem is Linear Optimal Transport (LOT). This method allows one to find a Euclidean embedding, called LOT embedding, of measures in some Wasserstein spaces, but some information is lost in this embedding. So, to understand whether statistical analysis relying on LOT embeddings can make valid inferences about original data, it is helpful to quantify how well these embeddings describe that data. To answer this question, we present a decomposition of the Fr\'echet variance of a set of measures in the 2-Wasserstein space, which allows one to compute the percentage of variance explained by LOT embeddings of those measures. We then extend this decomposition to the Fused Gromov-Wasserstein setting. We also present several experiments that explore the relationship between the dimension of the LOT embedding, the percentage of variance explained by the embedding, and the classification accuracy of machine learning classifiers built on the embedded data. We use the MNIST handwritten digits dataset, IMDB-50000 dataset, and Diffusion Tensor MRI images for these experiments. Our results illustrate the effectiveness of low dimensional LOT embeddings in terms of the percentage of variance explained and the classification accuracy of models built on the embedded data.

stat.ME

FairLENS: Assessing Fairness in Law Enforcement Speech Recognition

Automatic speech recognition (ASR) techniques have become powerful tools, enhancing efficiency in law enforcement scenarios. To ensure fairness for demographic groups in different acoustic environments, ASR engines must be tested across a variety of speakers in realistic settings. However, describing the fairness discrepancies between models with confidence remains a challenge. Meanwhile, most public ASR datasets are insufficient to perform a satisfying fairness evaluation. To address the limitations, we built FairLENS - a systematic fairness evaluation framework. We propose a novel and adaptable evaluation method to examine the fairness disparity between different models. We also collected a fairness evaluation dataset covering multiple scenarios and demographic dimensions. Leveraging this framework, we conducted fairness assessments on 1 open-source and 11 commercially available state-of-the-art ASR models. Our results reveal that certain models exhibit more biases than others, serving as a fairness guideline for users to make informed choices when selecting ASR models for a given real-world scenario. We further explored model biases towards specific demographic groups and observed that shifts in the acoustic domain can lead to the emergence of new biases.

eess.AS

A Wasserstein-type Distance for Gaussian Mixtures on Vector Bundles with Applications to Shape Analysis

This paper uses sample data to study the problem of comparing populations on finite-dimensional parallelizable Riemannian manifolds and more general trivial vector bundles. Utilizing triviality, our framework represents populations as mixtures of Gaussians on vector bundles and estimates the population parameters using a mode-based clustering algorithm. We derive a Wasserstein-type metric between Gaussian mixtures, adapted to the manifold geometry, in order to compare estimated distributions. Our contributions include an identifiability result for Gaussian mixtures on manifold domains and a convenient characterization of optimal couplings of Gaussian mixtures under the derived metric. We demonstrate these tools on some example domains, including the pre-shape space of planar closed curves, with applications to the shape space of triangles and populations of nanoparticles. In the nanoparticle application, we consider a sequence of populations of particle shapes arising from a manufacturing process, and utilize the Wasserstein-type distance to perform change-point detection.

stat.ME

How Abstract Is Linguistic Generalization in Large Language Models? Experiments with Argument Structure

Language models are typically evaluated on their success at predicting the distribution of specific words in specific contexts. Yet linguistic knowledge also encodes relationships between contexts, allowing inferences between word distributions. We investigate the degree to which pre-trained Transformer-based large language models (LLMs) represent such relationships, focusing on the domain of argument structure. We find that LLMs perform well in generalizing the distribution of a novel noun argument between related contexts that were seen during pre-training (e.g., the active object and passive subject of the verb spray), succeeding by making use of the semantically-organized structure of the embedding space for word embeddings. However, LLMs fail at generalizations between related contexts that have not been observed during pre-training, but which instantiate more abstract, but well-attested structural generalizations (e.g., between the active object and passive subject of an arbitrary verb). Instead, in this case, LLMs show a bias to generalize based on linear order. This finding points to a limitation with current models and points to a reason for which their training is data-intensive.s reported here are available at https://github.com/clay-lab/structural-alternations.

cs.CL

Regularising experimental correlations in LHC data: theory and application to a global analysis of parton distributions

We show how an inaccurate determination of experimental uncertainty correlations in high-precision LHC measurements may undermine the reliability of the associated $\chi^2$. We formulate the problem rigorously, and devise a regularisation procedure that increases the stability of the $\chi^2$ by altering the covariance matrix of the measurement as little as possible. We apply the procedure to the NNPDF4.0 global analysis of parton distribution functions that utilises a large amount of LHC measurements. We find that the regularised $\chi^2$ of the NNPDF4.0 determination is lowered by about $3\sigma$, without significantly altering the resulting PDFs upon refitting.

hep-ph

Static Analysis for AWS Best Practices in Python Code

Amazon Web Services (AWS) is a comprehensive and broadly adopted cloud provider, offering over 200 fully featured services, including compute, database, storage, networking and content delivery, machine learning, Internet of Things and many others. AWS SDKs provide access to AWS services through API endpoints. However, incorrect use of these APIs can lead to code defects, crashes, performance issues, and other problems. This paper presents automated static analysis rules, developed in the context of a commercial service for detection of code defects and security vulnerabilities, to identify deviations from AWS best practices in Python applications that use the AWS SDK. Such applications use the AWS SDK for Python, called "Boto3", to access AWS cloud services. However, precise static analysis of Python applications that use cloud SDKs requires robust type inference for inferring the types of cloud service clients. The dynamic style of Boto3 APIs poses unique challenges for type resolution, as does the interprocedural style in which service clients are used in practice. In support of our best-practices goal, we present a layered strategy for type inference that combines multiple type-resolution and tracking strategies in a staged manner. From our experiments across >3,000 popular Python GitHub repos that make use of the AWS SDK, our layered type inference system achieves 85% precision and 100% recall in inferring Boto3 clients in Python client code. Additionally, we present a representative sample of eight AWS best-practice rules that detect a wide range of issues including pagination, polling, and batch operations. We have assessed the efficacy of these rules based on real-world developer feedback. Developers have accepted more than 85% of the recommendations made by five out of eight Python rules, and almost 83% of all recommendations.

cs.PL

BioSimulators: a central registry of simulation engines and services for recommending specific tools

Computational models have great potential to accelerate bioscience, bioengineering, and medicine. However, it remains challenging to reproduce and reuse simulations, in part, because the numerous formats and methods for simulating various subsystems and scales remain siloed by different software tools. For example, each tool must be executed through a distinct interface. To help investigators find and use simulation tools, we developed BioSimulators (https://biosimulators.org), a central registry of the capabilities of simulation tools and consistent Python, command-line, and containerized interfaces to each version of each tool. The foundation of BioSimulators is standards, such as CellML, SBML, SED-ML, and the COMBINE archive format, and validation tools for simulation projects and simulation tools that ensure these standards are used consistently. To help modelers find tools for particular projects, we have also used the registry to develop recommendation services. We anticipate that BioSimulators will help modelers exchange, reproduce, and combine simulations.

q-bio.QM

Do Language Models Learn Position-Role Mappings?

How is knowledge of position-role mappings in natural language learned? We explore this question in a computational setting, testing whether a variety of well-performing pertained language models (BERT, RoBERTa, and DistilBERT) exhibit knowledge of these mappings, and whether this knowledge persists across alternations in syntactic, structural, and lexical alternations. In Experiment 1, we show that these neural models do indeed recognize distinctions between theme and recipient roles in ditransitive constructions, and that these distinct patterns are shared across construction type. We strengthen this finding in Experiment 2 by showing that fine-tuning these language models on novel theme- and recipient-like tokens in one paradigm allows the models to make correct predictions about their placement in other paradigms, suggesting that the knowledge of these mappings is shared rather than independently learned. We do, however, observe some limitations of this generalization when tasks involve constructions with novel ditransitive verbs, hinting at a degree of lexical specificity which underlies model performance.

cs.CL

Machine Learning Trivializing Maps: A First Step Towards Understanding How Flow-Based Samplers Scale Up

A trivializing map is a field transformation whose Jacobian determinant exactly cancels the interaction terms in the action, providing a representation of the theory in terms of a deterministic transformation of a distribution from which sampling is trivial. Recently, a proof-of-principle study by Albergo, Kanwar and Shanahan [arXiv:1904.12072] demonstrated that approximations of trivializing maps can be `machine-learned' by a class of invertible, differentiable neural models called \textit{normalizing flows}. By ensuring that the Jacobian determinant can be computed efficiently, asymptotically exact sampling from the theory of interest can be performed by drawing samples from a simple distribution and passing them through the network. From a theoretical perspective, this approach has the potential to become more efficient than traditional Markov Chain Monte Carlo sampling techniques, where autocorrelations severely diminish the sampling efficiency as one approaches the continuum limit. A major caveat is that it is not yet understood how the size of models and the cost of training them is expected to scale. As a first step, we have conducted an exploratory scaling study using two-dimensional $\phi^4$ with up to $20^2$ lattice sites. Although the scope of our study is limited to a particular model architecture and training algorithm, initial results paint an interesting picture in which training costs grow very quickly indeed. We describe a candidate explanation for the poor scaling, and outline our intentions to clarify the situation in future work.

hep-lat