SearcharxivSearch

arXiv subjects

Stephen Becker

Publications and source records attributed to Stephen Becker.

At least 19 recordsLinked to original sources

Gradient-Free Optimization for Matrix functions

We consider the task of optimizing smooth, possibly non-convex functions of a matrix variable given access only to directional derivatives rather than full gradients. This setting arises when fine-tuning large neural networks on consumer-grade hardware: the network's weights are matrices, memory constraints rule out backward-mode automatic differentiation, but directional derivatives remain available through forward mode. We frame gradient estimation in this setting as a structured recovery problem, in the spirit of signal processing. From this perspective we provide three contributions. First, we introduce an alternative to the standard random gradient estimator; the difference corresponds to replacing the adjoint of the sampling operator with its pseudoinverse. Second, when the gradient satisfies an approximate low-rank condition, techniques from matrix sensing yield a family of highly accurate gradient estimators that drop into any first-order method. Third, we note that while the computational cost of such estimators is high, this can be amortized by combining them with a matrix-aware optimizer such as spectral descent. Specifically, the gradient estimator computes a factorization of the gradient, allowing for the projection step of spectral descent to be done at no extra cost. We demonstrate our findings with two careful numerical experiments on synthetic functions with approximately low-rank gradients. We show that by exploiting this low-rank property one obtains much faster convergence to good approximate solutions.

math.OC

Nonnegative Low-Rank Matrix Correction under an Orthogonality Constraint in Conservative Vlasov Simulations

In low-rank numerical methods for Vlasov dynamics, the SVD-type truncation procedure may introduce negative entries into the numerical solution. Such negative values are unphysical because the solution is a probability distribution function. We design optimization-based post-processing algorithms to recover nonnegativity while preserving the macroscopic quantities (density, momentum, and energy) pointwise. The preservation of the macroscopic quantities is written as an orthogonality constraint on the correction term. For a convex formulation based on squared nuclear norm minimization, we show that the proximal operator with the orthogonality constraint is characterized by an implicit singular value thresholding equation, and the threshold can be computed efficiently by bisection. Based on this result, we develop five algorithms for the convex formulation: Douglas--Rachford splitting, restarted dual FISTA, restarted dual accelerated gradient descent, dual PR+ conjugate gradient, and dual L-BFGS. We also consider a non-convex formulation with an explicit rank constraint and develop a tangent-space accelerated alternating projection algorithm that only requires a \(2r \times 2r\) SVD per iteration. Numerical results for a Landau damping test case show that the proposed algorithms give comparable correction quality. Among them, the tangent-space accelerated alternating projection is the most cost-efficient, increasingly so as the problem size grows. We further demonstrate the correction as a positivity limiter inside a time-dependent conservative low-rank Vlasov solver, where it removes the negativity introduced by the SVD-type truncation while preserving the conserved mass, momentum, and energy.

math.NA

Singular value soft-thresholding via the polar decomposition

Singular value soft-thresholding can be computed via a reduction to the matrix polar decomposition, which allows one to exploit GPU-friendly algorithms for computing the polar decomposition. Empirically, there is a significant speed-up on GPUs compared to the standard approach using the SVD. We leave the investigation of robustness to future work, but note that due to the discontinuous nature of the sign function, the reduction to the polar decomposition is likely only suitable for low-accuracy applications.

math.NA

Functional Liu Regression for Scalar-on-Functional Models in High-Dimensional Settings

This study develops a functional Liu-type shrinkage estimator (fLiu) for scalar-on-function regression in the presence of strong multicollinearity and high-dimensional functional predictors. The approach extends the classical Liu estimator to the functional setting by combining directional shrinkage with smoothness regularization, providing flexible control over the bias-variance trade-off. Theoretical analysis is used to examine the behavior of the estimator and the associated parameter selection problem. In particular, an explicit mean squared error (MSE) decomposition is derived, characterizing the risk of the estimator in terms of variance reduction and shrinkage bias. This further yields an explicit optimal choice of the shrinkage parameter of the fLiu estimator through a one-dimensional convex risk minimization problem, leading to a practical plug-in tuning rule. Moreover, it is shown that in high-dimensional (underdetermined) settings, commonly used criterion such as GCV (and equivalently PRESS/LOO-CV) become constant with respect to the parameter d, thus uninformative for tuning. This provides a theoretical explanation for the predominant focus on the overdetermined regime in existing Liu-type methods. Numerical results demonstrate that the estimator achieves competitive predictive accuracy relative to existing methods. Implementation is carried out in R using the fda package, and in Python via the fLiu.py package developed for this study.

stat.OT

WSINDy for Model Predictive Control with Applications to Fusion, Drones, and Chaos

The control of complex dynamical systems remains a fundamental challenge in science and engineering, where strong nonlinearities, the presence of noise, and computational constraints often pose significant obstacles in traditional control approaches. Recent advances in data-driven methods, particularly system identification techniques, have shown a powerful alternative by providing fast, parsimonious, interpretable models that are well-suited for model predictive control (MPC). Building on these developments, the present article embeds WSINDy with actuation inputs (WSINDYc) within a MPC framework. Compared to benchmark data-driven methods, WSINDYc enables a more robust identification of the governing dynamics, particularly in the presence of high noise levels, resulting in more accurate and efficient control. The capabilities of the proposed WSINDY-MPC framework are demonstrated on a range of problems, including a tokamak plasma boundary model that includes main ion gas puff actuation, drone tracking and collision avoidance, the chaotic Lorenz system, and a simplified flight control model for an F-8 aircraft. The proposed framework achieves superior performance in the presence of noise, enabling longer prediction horizons, lower trajectory tracking error, and a more reliable obstacle clearance, while simultaneously achieving lower MPC cost values compared to the baseline methods.

math.DS

A Bifidelity Proximal Quasi-Newton Method for Dense Rigid Body Suspension Collision Resolution

Direct numerical simulation of dense rigid body suspensions poses significant computational challenges. A popular approach to resolve collisions necessitates solving a linear complementary problem (LCP) per time step. Each matrix vector product (MVP) inside the LCP requires solving an expensive partial differential equation. In this work, we show the LCP can be solved efficiently, often in only three to four MVPs. Specifically, we develop a custom monofidelity proximal quasi-Newton (Mono-PQN) method and a bi-fidelity variant (Bi-PQN). Our approach is validated through an application to representative systems of dense Stokesian Janus particles. Notably, in contact resolution our Mono-PQN and Bi-PQN achieve $\approx 1.5 \times$ and $> 2 \times$ speed up respectively against a competitive baseline, with the latter method displaying robust, problem-size-independent convergence. For our largest simulation involving $216$ particles, our Bi-PQN cut total simulation runtime to five days, as compared to the eight days required by the prior state-of-the-art method.

math.OC

In Situ Training of Implicit Neural Compressors for Scientific Simulations via Sketch-Based Regularization

Focusing on implicit neural representations, we present a novel in situ training protocol that employs limited memory buffers of full and sketched data samples, where the sketched data are leveraged to prevent catastrophic forgetting. The theoretical motivation for our use of sketching as a regularizer is presented via a simple Johnson-Lindenstrauss-informed result. While our methods may be of wider interest in the field of continual learning, we specifically target in situ neural compression using implicit neural representation-based hypernetworks. We evaluate our method on a variety of complex simulation data in two and three dimensions, over long time horizons, and across unstructured grids and non-Cartesian geometries. On these tasks, we show strong reconstruction performance at high compression rates. Most importantly, we demonstrate that sketching enables the presented in situ scheme to approximately match the performance of the equivalent offline method.

cs.LG

Generating Entangled Steady States in Multistable Open Quantum Systems via Initial State Control

Entanglement underpins the power of quantum technologies, yet it is fragile and typically destroyed by dissipation. Paradoxically, the same dissipation, when carefully engineered, can drive a system toward robust entangled steady states. However, this engineering task is nontrivial, as dissipative many-body systems are complex, particularly when they support multiple steady states. Here, we derive analytic expressions that predict how the steady state of a system evolving under a Lindblad equation depends on the initial state, without requiring integration of the dynamics. These results extend the frameworks developed in Refs. [Phys. Rev. A 89, 022118 (2014) and Phys. Rev. X 6, 041031 (2016)], showing that while the steady-state manifold is determined by the Liouvillian kernel, the weights within it depend on both the Liouvillian and the initial state. We identify a special class of Liouvillians for which the steady state depends only on the initial overlap with the kernel. Our framework provides analytical insight and a computationally efficient tool for predicting steady states in open quantum systems. As an application, we propose schemes to generate metrologically useful entangled steady states in spin ensembles via balanced collective decay.

quant-ph

SPIRiT Regularization: Parallel MRI with a Combination of Sensitivity Encoding and Linear Predictability

Accelerated Magnetic Resonance Imaging (MRI) permits high quality images from fewer samples that can be collected with a faster scan. Two established methods for accelerating MRI include parallel imaging and compressed sensing. Two types of parallel imaging include linear predictability, which assumes that the Fourier samples are linearly related, and sensitivity encoding, which incorporates a priori knowledge of the sensitivity maps. In this work, we combine compressed sensing with both types of parallel imaging using a novel regularization term: SPIRiT regularization. When combined, the reconstructed images are improved. We demonstrate results on data of a brain, a knee, and an ankle.

eess.IV

Rigorous Maximum Likelihood Estimation for Quantum States

Existing quantum state tomography methods are limited in scalability due to their high computation and memory demands, making them impractical for recovery of large quantum states. In this work, we address these limitations by reformulating the maximum likelihood estimation (MLE) problem using the Burer-Monteiro factorization, resulting in a non-convex but low-rank parameterization of the density matrix. We derive a fully unconstrained formulation by analytically eliminating the trace-one and positive semidefinite constraints, thereby avoiding the need for projection steps during optimization. Furthermore, we determine the Lagrange multiplier associated with the unit-trace constraint a priori, reducing computational overhead. The resulting formulation is amenable to scalable first-order optimization, and we demonstrate its tractability using limited-memory BFGS (L-BFGS). Importantly, we also propose a low-memory version of the above algorithm to fully recover certain large quantum states with Pauli-based POVM measurements. Our low-memory algorithm avoids explicitly forming any density matrix, and does not require the density matrix to have a matrix product state (MPS) or other tensor structure. For a fixed number of measurements and fixed rank, our algorithm requires just $\mathcal{O}(d \log d)$ complexity per iteration to recover a $d \times d$ density matrix. Additionally, we derive a useful error bound that can be used to give a rigorous termination criterion. We numerically demonstrate that our method is competitive with state-of-the-art algorithms for moderately sized problems, and then demonstrate that our method can solve a 20-qubit problem on a laptop in under 5 hours.

quant-ph

Stochastic Subspace Descent Accelerated via Bi-fidelity Line Search

Efficient optimization remains a fundamental challenge across numerous scientific and engineering domains, especially when objective function and gradient evaluations are computationally expensive. While zeroth-order optimization methods offer effective approaches when gradients are inaccessible, their practical performance can be limited by the high cost associated with function queries. This work introduces the bi-fidelity stochastic subspace descent (BF-SSD) algorithm, a novel zeroth-order optimization method designed to reduce this computational burden. BF-SSD leverages a bi-fidelity framework, constructing a surrogate model from a combination of computationally inexpensive low-fidelity (LF) and accurate high-fidelity (HF) function evaluations. This surrogate model facilitates an efficient backtracking line search for step size selection, for which we provide theoretical convergence guarantees under standard assumptions. We perform a comprehensive empirical evaluation of BF-SSD across four distinct problems: a synthetic optimization benchmark, dual-form kernel ridge regression, black-box adversarial attacks on machine learning models, and transformer-based black-box language model fine-tuning. Numerical results demonstrate that BF-SSD consistently achieves superior optimization performance while requiring significantly fewer HF function evaluations compared to relevant baseline methods. This study highlights the efficacy of integrating bi-fidelity strategies within zeroth-order optimization, positioning BF-SSD as a promising and computationally efficient approach for tackling large-scale, high-dimensional problems encountered in various real-world applications.

cs.LG

A Relaxed Primal-Dual Hybrid Gradient Method with Line Search

The primal-dual hybrid gradient method (PDHG) is useful for optimization problems that commonly appear in image reconstruction. A downside of PDHG is that there are typically three user-set parameters and performance of the algorithm is sensitive to their values. Toward a parameter-free algorithm, we combine two existing line searches. The first, by Malitsky et al., is over two of the step sizes in the PDHG iterations. We then use the connection between PDHG and the primal-dual form of Douglas-Rachford splitting to construct a line search over the relaxation parameter. We demonstrate the efficacy of the combined line search on multiple problems, including a novel inverse problem in magnetic resonance image reconstruction. The method presented in this manuscript is the first parameter-free variant of PDHG (across all numerical experiments, there were no changes to line search hyperparameters).

math.OC

Aligning to What? Limits to RLHF Based Alignment

Reinforcement Learning from Human Feedback (RLHF) is increasingly used to align large language models (LLMs) with human preferences. However, the effectiveness of RLHF in addressing underlying biases remains unclear. This study investigates the relationship between RLHF and both covert and overt biases in LLMs, particularly focusing on biases against African Americans. We applied various RLHF techniques (DPO, ORPO, and RLOO) to Llama 3 8B and evaluated the covert and overt biases of the resulting models using matched-guise probing and explicit bias testing. We performed additional tests with DPO on different base models and datasets; among several implications, we found that SFT before RLHF calcifies model biases. Additionally, we extend the tools for measuring biases to multi-modal models. Through our experiments we collect evidence that indicates that current alignment techniques are inadequate for nebulous tasks such as mitigating covert biases, highlighting the need for capable datasets, data curating techniques, or alignment tools.

cs.CL

WENDy for Nonlinear-in-Parameters ODEs

The Weak-form Estimation of Non-linear Dynamics (WENDy) framework is a recently developed approach for parameter estimation and inference of systems of ordinary differential equations (ODEs). Prior work demonstrated WENDy to be robust, computationally efficient, and accurate, but only works for ODEs which are linear-in-parameters. In this work, we derive a novel extension to accommodate systems of a more general class of ODEs that are nonlinear-in-parameters. Our new WENDy-MLE algorithm approximates a maximum likelihood estimator via local non-convex optimization methods. This is made possible by the availability of analytic expressions for the likelihood function and its first and second order derivatives. WENDy-MLE has better accuracy, a substantially larger domain of convergence, and is often faster than other weak form methods and the conventional output error least squares method. Moreover, we extend the framework to accommodate data corrupted by multiplicative log-normal noise. The WENDy.jl algorithm is efficiently implemented in Julia. In order to demonstrate the practical benefits of our approach, we present extensive numerical results comparing our method, other weak form methods, and output error least squares on a suite of benchmark systems of ODEs in terms of accuracy, precision, bias, and coverage.

cs.LG

Exploring Exploration in Bayesian Optimization

A well-balanced exploration-exploitation trade-off is crucial for successful acquisition functions in Bayesian optimization. However, there is a lack of quantitative measures for exploration, making it difficult to analyze and compare different acquisition functions. This work introduces two novel approaches - observation traveling salesman distance and observation entropy - to quantify the exploration characteristics of acquisition functions based on their selected observations. Using these measures, we examine the explorative nature of several well-known acquisition functions across a diverse set of black-box problems, uncover links between exploration and empirical performance, and reveal new relationships among existing acquisition functions. Beyond enabling a deeper understanding of acquisition functions, these measures also provide a foundation for guiding their design in a more principled and systematic manner.

cs.LG

A Unified Framework for Entropy Search and Expected Improvement in Bayesian Optimization

Bayesian optimization is a widely used method for optimizing expensive black-box functions, with Expected Improvement being one of the most commonly used acquisition functions. In contrast, information-theoretic acquisition functions aim to reduce uncertainty about the function's optimum and are often considered fundamentally distinct from EI. In this work, we challenge this prevailing perspective by introducing a unified theoretical framework, Variational Entropy Search, which reveals that EI and information-theoretic acquisition functions are more closely related than previously recognized. We demonstrate that EI can be interpreted as a variational inference approximation of the popular information-theoretic acquisition function, named Max-value Entropy Search. Building on this insight, we propose VES-Gamma, a novel acquisition function that balances the strengths of EI and MES. Extensive empirical evaluations across both low- and high-dimensional synthetic and real-world benchmarks demonstrate that VES-Gamma is competitive with state-of-the-art acquisition functions and in many cases outperforms EI and MES.

stat.ML

Evaluation of data driven low-rank matrix factorization for accelerated solutions of the Vlasov equation

Low-rank methods have shown success in accelerating simulations of a collisionless plasma described by the Vlasov equation, but still rely on computationally costly linear algebra every time step. We propose a data-driven factorization method using artificial neural networks, specifically with convolutional layer architecture, that trains on existing simulation data. At inference time, the model outputs a low-rank decomposition of the distribution field of the charged particles, and we demonstrate that this step is faster than the standard linear algebra technique. Numerical experiments show that the method effectively interpolates time-series data, generalizing to unseen test data in a manner beyond just memorizing training data; patterns in factorization also inherently followed the same numerical trend as those within algebraic methods (e.g., truncated singular-value decomposition). However, when training on the first 70% of a time-series data and testing on the remaining 30%, the method fails to meaningfully extrapolate. Despite this limiting result, the technique may have benefits for simulations in a statistical steady-state or otherwise showing temporal stability.

math.NA

Online randomized interpolative decomposition with a posteriori error estimator for temporal PDE data reduction

Traditional low-rank approximation is a powerful tool to compress the huge data matrices that arise in simulations of partial differential equations (PDE), but suffers from high computational cost and requires several passes over the PDE data. The compressed data may also lack interpretability thus making it difficult to identify feature patterns from the original data. To address these issues, we present an online randomized algorithm to compute the interpolative decomposition (ID) of large-scale data matrices {\em in situ}. Compared to previous randomized IDs that used the QR decomposition to determine the column basis, we adopt a streaming ridge leverage score-based column subset selection algorithm that dynamically selects proper basis columns from the data and thus avoids an extra pass over the data to compute the coefficient matrix of the ID. In particular, we adopt a single-pass error estimator based on the non-adaptive Hutch++ algorithm to provide real-time error approximation for determining the best coefficients. As a result, our approach only needs a single pass over the original data and thus is suitable for large and high-dimensional matrices stored outside of core memory or generated in PDE simulations. A strategy to improve the accuracy of the reconstructed data gradient, when desired, within the ID framework is also presented. We provide numerical experiments on turbulent channel flow and ignition simulations, and on the NSTX Gas Puff Image dataset, comparing our algorithm with the offline ID algorithm to demonstrate its utility in real-world applications.

math.NA