SearcharxivSearch

arXiv subjects

Martin T. Wells

Publications and source records attributed to Martin T. Wells.

At least 19 recordsLinked to original sources

Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization

We provide a theoretical analysis of Adam under non-stationary stochastic objectives, separating two regimes: Euclidean tracking under adaptive strong monotonicity of the Adam-preconditioned mean-gradient operator, and high-probability projected stationarity guarantees under general $L$-smooth objectives. In the tracking regime, we derive finite-time expected and high-probability bounds that decompose sharply into four components: initialization, objective drift, a first-moment tracking error governed by $β_1$, and a preconditioner perturbation governed by $β_2$. We characterize the burn-in time required for the transient terms to decay to the asymptotic tracking bound under constant and step-decay schedules. We also prove a high-probability bound on the average projected stationarity gap for Adam under distribution shift. Across both analyses, our bounds reveal a noise--drift tradeoff: in noise-dominated regimes, first-moment averaging and adaptive preconditioning can yield favorable upper guarantees, whereas in drift-dominated regimes, stale first-moment information and preconditioner perturbations can enlarge Adam's tracking guarantee, potentially allowing vanilla SGD to attain a smaller tracking error. Our explicit $(β_1,β_2,ε)$-dependent bounds identify mechanisms through which adaptive step-sizing can help or hurt under nonstationarity and provide theoretical explanations consistent with Adam's empirical instability and stabilization under distribution shift.

stat.ML

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization

In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness. Our finite-time bounds reveal a sharp decomposition of tracking error into transient, noise-induced, and drift-induced components. This decomposition exposes a fundamental trade-off: while momentum is often used as a gradient-smoothing heuristic, under distribution shift it incurs an explicit drift-amplification penalty that diverges as the momentum parameter $β$ approaches 1, yielding systematic tracking lag. We complement these upper bounds with minimax lower bounds under gradient-variation constraints, proving this momentum-induced tracking penalty is not an analytical artifact but an information-theoretic barrier: in drift-dominated regimes, momentum is unavoidably worse because stale-gradient averaging forces systematic lag. Our results provide theoretical grounding for the empirical instability of momentum in dynamic settings and precisely delineate regime boundaries where vanilla SGD provably outperforms its accelerated counterparts.

stat.ML

Nonparametric Regression via Tree-Guided Feature Aggregation

In regression problems where covariates are naturally organized in a hierarchical tree structure, a central challenge is to select the resolution at which covariates enter the model. Determining this level of feature aggregation is of intrinsic scientific interest and can improve statistical efficiency by inducing sparsity. While a rich literature addresses this problem in the linear setting, extending feature aggregation to the nonlinear setting remains an open challenge. In this work, we propose to simultaneously perform model selection and feature aggregation through a penalized Nadaraya-Watson-type estimator. Our proposed estimator, Kernel Regression with Tree-EXploring AggregationS (KR-TEXAS), constructs adaptive penalty weights for the features based on pilot estimators of the regression function's partial derivatives. Under mild conditions, we establish model selection consistency for a well-defined target aggregation set, and our simulations show strong performance in both model selection and prediction. Finally, we demonstrate the utility of our procedure by applying it to a microbiome data set to predict short chain fatty acids. A user-friendly implementation of our procedure is available in the R package krtexas.

stat.ME

Double Local-to-Unity: Inference under Nearly Nonstationary Volatility

This article develops a moderate-deviation limit theory for autoregressive models with jointly persistent mean and volatility dynamics. The autoregressive coefficient is allowed to drift toward unity slower than the classical 1/n rate, while the volatility persistence parameter also converges to one at an even slower, logarithmic order, so that the conditional variance process is itself nearly nonstationary and its unconditional moments may diverge. This double localization allows the variance process to be nearly nonstationary and to evolve slowly, as observed in financial data and during asset price bubble episodes. Under standard regularity conditions, we establish consistency and distributional limits for the OLS estimator of the autoregressive coefficient that remains valid in the presence of highly persistent stochastic volatility. We show that the effective normalization for least squares inference is governed by an average volatility scale, and we derive martingale limit theorems for the OLS estimator under joint drift and volatility dynamics. In a mildly stationary regime (where the autoregressive root approaches one from below), the OLS estimator is asymptotically normal. In a mildly explosive regime (where the root approaches one from above), an OLS based self normalized statistic converges to a Cauchy limit. Strikingly, in both regimes, the limiting laws of our statistics are invariant to the detailed specification of the volatility process, even though the conditional variance is itself nearly nonstationary. Overall, the results extend moderate-deviation asymptotics to settings with drifting volatility persistence, unify local to unity inference with nearly nonstationary stochastic volatility, and deliver practically usable volatility robust statistics for empirical work in settings approaching instability and exhibiting bubbles.

math.ST

Is There an AI Bubble? Robust Date-Stamping for Periods of Exuberance

The recent surge in valuations among AI related firms has renewed concerns that markets may be entering a new phase of speculative exuberance, especially in the technology and semiconductor sectors at the center of the AI investment wave. This paper develops a practical econometric framework for detecting, date-stamping, and drawing inference on the origination and collapse of bubble episodes when prices evolve under persistent, time-varying volatility. Standard bubble tests are typically derived under homoskedasticity or weak heteroskedasticity and may therefore yield misleading inference in more general settings. We extend right-tailed Dickey-Fuller unit root tests to autoregressive models with highly persistent mean and volatility dynamics, delivering a stochastic-volatility-robust ADF (SV-ADF) test that accommodates persistent variance without imposing strict parametric structure. Building on a moderate-deviation asymptotic theory, the SV-ADF yields nuisance-parameter-free procedures with distinct critical values for origination and collapse, producing more stable alarms and fewer transient false positives around volatility spikes. We establish consistency of the date-stamping estimator and show that it remains asymptotically tractable. Monte Carlo simulations document strong power and substantial gains over homoskedastic (PWY) procedures when volatility dynamics are pronounced. An empirical analysis of AI-exposed equities, including the "Magnificent Seven" and leading semiconductor firms, finds pervasive exuberance with substantial heterogeneity in timing, intensity, and duration. The evidence points to especially strong bubble dynamics for Alphabet and TSMC in the current cycle, while Tesla and Nvidia exhibited pronounced explosive episodes in earlier phases of the AI-driven market cycle.

stat.ME

Modeling Dynamic Correlation Matrices with Shrinkage Priors

Estimating time-varying correlation matrices is challenging because existing methods may adapt slowly to structural changes, impose insufficient regularization, or produce diffuse posterior uncertainty. In moderate dimensions, an additional difficulty is summarizing the estimated evolving dependence structure for downstream decision-making tasks. We propose a Bayesian approach based on a low-rank factor representation, with latent states evolving under a dynamic shrinkage prior and observation errors following a multivariate factor stochastic volatility model. This specification allows locally adaptive regularization of the estimated correlation structure over time and informative uncertainty quantification. We establish, to our knowledge, a first-of-its-kind posterior contraction result for dynamically regularized Bayesian models, showing contraction around the true model parameters at an explicit rate under averaged Hellinger distance. To summarize the estimated correlation matrices, we build on the information-theoretic concept of total correlation to obtain a scalar measure of cross-sectional dependence. Simulation studies show improved accuracy and responsiveness relative to competing methods in a range of challenging scenarios. We then apply our method to monitoring the correlation evolution of equity portfolios during periods of financial market stress, providing an ex post framework for assessing the changing benefits of diversification in backtesting analyses.

stat.ME

Foreclassing: A new machine learning perspective on human decision making with temporal data

Time series forecasts are widely used to inform decisions. Human decision-makers interpret these forecasts, incorporate prior experience and uncertainty about future outcomes, and then make a decision. In this paper, we propose a new machine learning problem, which we call Foreclassing, which addresses settings in which the aim is to automate human involvement in such decision-making processes. Our aim is to develop a unified end-to-end model that takes a time series as input, produces a forecast, accounts for its predictive uncertainty, and makes a downstream classification decision, enabling models to support or automate such temporal decision-making tasks. Related problems arise across a range of applications, yet the literature lacks both a unified methodology and a formal problem statement. By formalizing the task, we aim to stimulate research on such models and encourage cross-domain collaboration. To solve the Foreclassing problem, we propose a deep Bayesian neural network, ForeClassNet. As part of this framework, we introduce a new type of neural network layer, Boltzmann convolutions, which enable probabilistic learning of kernel sizes in convolutional layers. We evaluate the Foreclassing framework against standard time series classification methods and demonstrate the efficacy of ForeClassNet on real-world Foreclassing datasets from the weather, energy, and finance domains, achieving superior performance relative to state-of-the-art time series classifiers.

stat.ML

A Milestone-Based Framework for Characterizing Time-Varying Treatment Effects in Immunotherapy Trials

Immune checkpoint inhibitor--based therapies often produce heterogeneous survival responses, including early risk, delayed treatment benefit, and durable long-term survival in a subset of patients. In these settings, conventional summary measures such as the hazard ratio may not adequately describe how treatment effects evolve over follow-up. We propose a milestone-based framework that separates long-term survival beyond a clinically meaningful time point from earlier outcomes and provides a practical way to characterize patient heterogeneity in treatment response. The framework summarizes treatment differences through milestone survival probabilities and, among patients who do not reach the milestone, characterizes short-term treatment ordering over time using a tau-based summary that helps identify hazard reversal. We illustrate the approach using reconstructed individual-level data from three landmark phase III trials: CheckMate~067, CheckMate~227, and CLEAR. Across these examples, the framework captures patterns that are difficult to summarize with conventional measures, including settings in which early disadvantage coexists with later durable benefit. It also helps clarify when treatment benefit begins to emerge and how short-term and long-term effects differ within the same trial. This approach provides a clinically interpretable and statistically principled way to evaluate heterogeneous and time-varying treatment effects in oncology trials with nonproportional hazards.

stat.ME

Online Distributionally Robust LLM Alignment via Regression to Relative Reward

Reinforcement Learning with Human Feedback (RLHF) has become crucial for aligning Large Language Models (LLMs) with human intent. However, existing offline RLHF approaches suffer from overoptimization, where language models degrade by overfitting inaccuracies and drifting from preferred behaviors observed during training. Distributionally robust optimization (DRO) is a natural solution, but existing DRO-DPO methods are sample-inefficient, ignore heterogeneous preferences, and lean on brittle heuristics. We introduce \emph{DRO-REBEL}, a family of robust online REBEL updates built on type-$p$ Wasserstein, Kullback-Leibler (KL), and $χ^2$ ambiguity sets. Strong duality reduces each update to a relative-reward regression, retaining REBEL's scalability without PPO-style clipping or value networks. Under linear rewards, log-linear policies, and a standard coverage condition, we prove $\widetilde{O}(\sqrt{d/n})$ bounds on squared parameter error, with sharper constants than prior DRO-DPO analyses, and give the first parametric $\widetilde{O}(d/n)$ rate for DRO-based alignment under preference shift, matching non-robust RLHF in benign regimes. Each divergence yields a tractable SGD-based algorithm: gradient regularization for Wasserstein, importance weighting for KL, and a 1-D dual solve for $χ^2$. On Emotion Alignment, the ArmoRM multi-objective benchmark, and HH-Alignment, DRO-REBEL outperforms prior robust and non-robust baselines across unseen preference mixtures, model sizes, and dataset scales.

cs.LG

Minimaxity and Admissibility of Bayesian Neural Networks

Bayesian neural networks (BNNs) offer a natural probabilistic formulation for inference in deep learning models. Despite their popularity, their optimality has received limited attention through the lens of statistical decision theory. In this paper, we study decision rules induced by deep, fully connected feedforward ReLU BNNs in the normal location model under quadratic loss. We show that, for fixed prior scales, the induced Bayes decision rule is not minimax. We then propose a hyperprior on the effective output variance of the BNN prior that yields a superharmonic square-root marginal density, establishing that the resulting decision rule is simultaneously admissible and minimax. We further extend these results from the quadratic loss setting to the predictive density estimation problem with Kullback--Leibler loss. Finally, we validate our theoretical findings numerically through simulation.

math.ST

Fragility Measures For Typical Cases

The fragility index is a clinically motivated metric designed to supplement the $p$ value during hypothesis testing. The measure relies on two pillars: selecting cases to have their outcome modified and modifying the outcomes. The measure is interesting but the case selection suffers from a drawback which can hamper its interpretation. This work presents the drawback and a method, the stochastic generalized fragility indices, designed to remedy it. Two examples concerning electoral outcomes and the causal effect of smoking cessation illustrate the method.

stat.ME

SEMMS with Random Effects: A Mixed-Model Extension for Variable Selection in Clustered and Longitudinal Data

SEMMS (Scalable Empirical-Bayes Model for Marker Selection) is a variable-selection procedure for generalized linear models that uses a three-component normal mixture prior on regression coefficients. In its original form, SEMMS assumes that all observations are independent. Many real-world datasets, however, arise from repeated-measures or clustered designs in which observations within the same subject are correlated. Ignoring this correlation inflates the apparent residual variance and can severely degrade variable-selection performance. We extend SEMMS to accommodate random intercepts, random slopes, or both, via an alternating coordinate-ascent algorithm. After each round of fixed-effect variable selection, the subject-level best linear unbiased predictors (BLUPs) are updated with \texttt{lmer} (Gaussian) or \texttt{glmer} (non-Gaussian); the fixed-effect step then operates on the random-effect-adjusted response. We describe the algorithm, evaluate its performance in three Gaussian simulation studies spanning a range of signal strengths, random-effect magnitudes, and sample/predictor-space regimes, and present a semi-synthetic real-data example. We further extend the framework to non-Gaussian families (Poisson, binomial) via an IRLS working-response adaptation: at each outer iteration the fixed-effects step uses the RE-adjusted working response computed from the current \texttt{glmer} fitted values rather than the raw response. When the fixed-effect signal is strong relative to the random-effect variance, both the original and extended procedures perform comparably. When the random-effect variance dominates -- the scenario most likely to cause plain SEMMS to fail -- the mixed-model extension recovers the exact true predictor set in 93\% of simulated datasets (Gaussian), 61\% (Poisson), and 65\% (binomial), compared with 1\%, 45\%, and 39\% for plain SEMMS respectively.

stat.CO

Democratic Preference Alignment via Sortition-Weighted RLHF

Whose values should AI systems learn? Preference based alignment methods like RLHF derive their training signal from human raters, yet these rater pools are typically convenience samples that systematically over represent some demographics and under represent others. We introduce Democratic Preference Optimization, or DemPO, a framework that applies algorithmic sortition, the same mechanism used to construct citizen assemblies, to preference based fine tuning. DemPO offers two training schemes. Hard Panel trains exclusively on preferences from a quota satisfying mini public sampled via sortition. Soft Panel retains all data but reweights each rater by their inclusion probability under the sortition lottery. We prove that Soft Panel weighting recovers the expected Hard Panel objective in closed form. Using a public preference dataset that pairs human judgments with rater demographics and a seventy five clause constitution independently elicited from a representative United States panel, we evaluate Llama models from one billion to eight billion parameters fine tuned under each scheme. Across six aggregation methods, the Hard Panel consistently ranks first and the Soft Panel consistently outperforms the unweighted baseline, with effect sizes growing as model capacity increases. These results demonstrate that enforcing demographic representativeness at the preference collection stage, rather than post hoc correction, yields models whose behavior better reflects values elicited from representative publics.

cs.AI

Quantum Cognition Machine Learning for Forecasting Chromosomal Instability

The accurate prediction of chromosomal instability from the morphology of circulating tumor cells (CTCs) enables real-time detection of CTCs with high metastatic potential in the context of liquid biopsy diagnostics. However, it presents a significant challenge due to the high dimensionality and complexity of single-cell digital pathology data. Here, we introduce the application of Quantum Cognition Machine Learning (QCML), a quantum-inspired computational framework, to estimate morphology-predicted chromosomal instability in CTCs from patients with metastatic breast cancer. QCML leverages quantum mechanical principles to represent data as state vectors in a Hilbert space, enabling context-aware feature modeling, dimensionality reduction, and enhanced generalization without requiring curated feature selection. QCML outperforms conventional machine learning methods when tested on out of sample verification CTCs, achieving higher accuracy in identifying predicted large-scale state transitions (pLST) status from CTC-derived morphology features. These preliminary findings support the application of QCML as a novel machine learning tool with superior performance in high-dimensional, low-sample-size biomedical contexts. QCML enables the simulation of cognition-like learning for the identification of biologically meaningful prediction of chromosomal instability from CTC morphology, offering a novel tool for CTC classification in liquid biopsy.

q-bio.QM

Quantum Geometry of Data

We demonstrate how Quantum Cognition Machine Learning (QCML) encodes data as quantum geometry. In QCML, features of the data are represented by learned Hermitian matrices, and data points are mapped to states in Hilbert space. The quantum geometry description endows the dataset with rich geometric and topological structure - including intrinsic dimension, quantum metric, and Berry curvature - derived directly from the data. QCML captures global properties of data, while avoiding the curse of dimensionality inherent in local methods. We illustrate this on a number of synthetic and real-world examples. Quantum geometric representation of QCML could advance our understanding of cognitive phenomena within the framework of quantum cognition.

cs.LG

Quantitative Relaxations of Arrow's Axioms

In this paper we develop a novel approach to relaxing Arrow's axioms for voting rules, addressing a long-standing critique in social choice theory. Classical axioms (often styled as fairness axioms or fairness criteria) are assessed in a binary manner, so that a voting rule fails the axiom if it fails in even one corner case. Many authors have proposed a probabilistic framework to soften the axiomatic approach. Instead of immediately passing to random preference profiles, we begin by measuring the degree to which an axiom is upheld or violated on a given profile. We focus on two foundational axioms-Independence of Irrelevant Alternatives (IIA) and Unanimity (U)-and extend them to take values in $[0,1]$. Our $σ_{IIA}$ measures the stability of a voting rule when candidates are removed from consideration, while $σ_{U}$ captures the degree to which the outcome respects majority preferences. Together, these metrics quantify how a voting rule navigates the fundamental trade-off highlighted by Arrow's Theorem. We show that $σ_{IIA}\equiv 1$ recovers classical IIA, and $σ_{U}>0$ recovers classical Unanimity, allowing a quantitative restatement of Arrow's Theorem. In the empirical part of the paper, we test these metrics on two kinds of data: a set of over 1000 ranked choice preference profiles from Scottish local elections, and a batch of synthetic preference profiles generated with a Bradley-Terry-type model. We use those to investigate four positional voting rules-Plurality, 2-Approval, 3-Approval, and the Borda rule-as well as the iterative rule known as Single Transferable Vote (STV). The Borda rule consistently receives the highest $σ_{IIA}$ and $σ_{U}$ scores across observed and synthetic elections. This compares interestingly with a recent result of Maskin showing that weakening IIA to include voter preference intensity uniquely selects Borda.

cs.GT

BLOG: Bayesian Longitudinal Omics with Group Constraints

Clinical investigators are increasingly interested in discovering computational biomarkers from short-term longitudinal omics data sets. This work focuses on Bayesian regression and variable selection for longitudinal omics datasets, which can quantify uncertainty and control false discovery. In our univariate approach, Zellner's $g$ prior is used with two different options of the tuning parameter $g$: $g=\sqrt{n}$ and a $g$ that minimizes Stein's unbiased risk estimate (SURE). Bayes Factors were used to quantify uncertainty and control for false discovery. In the multivariate approach, we use Bayesian Group LASSO with a spike and slab prior for group variable selection. In both approaches, we use the first difference ($Δ$) scale of longitudinal predictor and the response. These methods work together to enhance our understanding of biomarker identification, improving inference and prediction. We compare our method against commonly used linear mixed effect models on simulated data and real data from a Tuberculosis (TB) study on metabolite biomarker selection. With an automated selection of hyperparameters, the Zellner's $g$ prior approach correctly identifies target metabolites with high specificity and sensitivity across various simulation and real data scenarios. The Multivariate Bayesian Group Lasso spike and slab approach also correctly selects target metabolites across various simulation scenarios.

stat.ME

Constructing Bayes Minimax Estimators through Integral Transformations

The problem of Bayes minimax estimation for the mean of a multivariate normal distribution under quadratic loss has attracted significant attention recently. These estimators have the advantageous property of being admissible, similar to Bayes procedures, while also providing the conservative risk guarantees typical of frequentist methods. This paper demonstrates that Bayes minimax estimators can be derived using integral transformation techniques, specifically through the \( I \)-transform and the Laplace transform, as long as appropriate spherical priors are selected. Several illustrative examples are included to highlight the effectiveness of the proposed approach.

math.ST