SearcharxivSearch

arXiv subjects

James McInerney

Publications and source records attributed to James McInerney.

At least 19 recordsLinked to original sources

Mult-DPO: Multinomial Direct Preference Optimization for Recommender Systems

Direct preference optimization (DPO) is a simple and effective alignment strategy for large language models (LLMs) based on pairwise preferences. In recommender systems, however, user feedback is rarely pairwise. For a given context, e.g., a user, a session, or a conversation, we typically observe set-wise preferences with multiple positive items, where every positive item should outrank every unobserved or explicitly negative item, with no prescribed order among the positives or the negatives themselves. A natural generalization is to use the Plackett-Luce (PL) reward model, which extends the Bradley-Terry reward model underlying vanilla DPO from pairwise preferences to full rankings of candidates. However, we show that adapting the PL model to set-wise preferences requires marginalizing over all positive orderings, where the resulting expression is combinatorial in complexity. To address this fundamental challenge, we propose Mult-DPO, a novel DPO objective with a tractable multinomial surrogate likelihood over set-wise preference events for the user-preference alignment of LLM-based recommender systems. The multinomial construction is not itself a ranking distribution, but it is defined on the same reward-induced weight space and admits a closed-form DPO-style objective, enabling direct alignment of LLMs with multiple candidates through a classification-style objective. In addition, we prove that the multinomial DPO loss is a tractable upper bound on the marginalized PL DPO loss when optimizing against the set-wise preference data. We further characterize the tightness of this bound in terms of the relative total weight of positives versus negatives, which provides insights into tightening the bound with richer or harder negatives. Finally, we extend Mult-DPO to the alignment of LLMs with multiple preference levels. Code is available at https://github.com/yaochenzhu/Mult_DPO

cs.IR

Entropy After for reasoning model early exiting

Reasoning LLMs show improved performance with longer chains of thought. However, recent work has highlighted their tendency to overthink, continuing to revise answers even after reaching the correct solution. We quantitatively confirm this inefficiency from the distribution dynamics perspective by tracking Pass@1 for answers averaged over a large number of rollouts and find the model often begins to always produce the correct answer early in the reasoning, making extra reasoning tokens wasteful. To detect and prevent overthinking, we propose a simple and inexpensive novel signal, Entropy After (EAT), for monitoring and deciding whether to exit reasoning early. By appending a stop thinking token ( ) and monitoring the entropy of the following token as the model reasons, we obtain a trajectory that decreases and stabilizes when Pass@1 plateaus; thresholding its variance under an exponential moving average yields a practical stopping rule. Importantly, our approach enables adaptively allocating compute based on the EAT trajectory, allowing us to spend compute in a more efficient way compared with fixing the token budget for all questions. Empirically, on MATH500 and AIME2025, EAT reduces token usage by 12 - 22% without harming accuracy. EAT also remains effective in black box settings where logits from the reasoning model are not accessible, and EAT is computed with proxy models: We verified the feasibility via early stopping Llama 70B with a 1.5B model and Claude 3.7 with a local 4B model.

cs.LG

Optimization of Epsilon-Greedy Exploration

Modern recommendation systems rely on exploration to learn user preferences for new items, typically implementing uniform exploration policies (e.g., epsilon-greedy) due to their simplicity and compatibility with machine learning (ML) personalization models. Within these systems, a crucial consideration is the rate of exploration - what fraction of user traffic should receive random item recommendations and how this should evolve over time. While various heuristics exist for navigating the resulting exploration-exploitation tradeoff, selecting optimal exploration rates is complicated by practical constraints including batched updates, time-varying user traffic, short time horizons, and minimum exploration requirements. In this work, we propose a principled framework for determining the exploration schedule based on directly minimizing Bayesian regret through stochastic gradient descent (SGD), allowing for dynamic exploration rate adjustment via Model-Predictive Control (MPC). Through extensive experiments with recommendation datasets, we demonstrate that variations in the batch size across periods significantly influence the optimal exploration strategy. Our optimization methods automatically calibrate exploration to the specific problem setting, consistently matching or outperforming the best heuristic for each setting.

cs.LG

Novel mechanical response of parallelogram-face origami governed by topological characteristics

Origami principles are used to create strong, lightweight structures with complex mechanical response. However, identifying the fundamental physical principles that determine a sheet's behavior remains a challenge. We introduce a new analytic theory in which commonly studied origami sheets fall into distinct topological classes that predict sharply varying mechanical behavior, including effective stiffness and smoothness of mechanical response under external loads. Origami sheets with negative Poisson's ratios, such as the Miura ori, have conventional, smooth mechanical response amenable to continuum-based approaches. In contrast, positive Poisson's ratio, as in the Eggbox ori, generates a topological transition to lines of doubly degenerate zero modes that lead to dramatically softer structures with uneven, complex patterns of spatial response. These patterns interact in complicated ways with origami boundary conditions and source terms, leading to rich physical phenomena in experimentally accessible systems. This approach highlights topological mechanics, with deep connections to topologically protected quantum-mechanical systems, as a design principle for controlling the mechanical response of thin, complex sheets.

cond-mat.soft

Adjusting Regression Models for Conditional Uncertainty Calibration

Conformal Prediction methods have finite-sample distribution-free marginal coverage guarantees. However, they generally do not offer conditional coverage guarantees, which can be important for high-stakes decisions. In this paper, we propose a novel algorithm to train a regression function to improve the conditional coverage after applying the split conformal prediction procedure. We establish an upper bound for the miscoverage gap between the conditional coverage and the nominal coverage rate and propose an end-to-end algorithm to control this upper bound. We demonstrate the efficacy of our method empirically on synthetic and real-world datasets.

stat.ML

Variation Due to Regularization Tractably Recovers Bayesian Deep Learning

Uncertainty quantification in deep learning is crucial for safe and reliable decision-making in downstream tasks. Existing methods quantify uncertainty at the last layer or other approximations of the network which may miss some sources of uncertainty in the model. To address this gap, we propose an uncertainty quantification method for large networks based on variation due to regularization. Essentially, predictions that are more (less) sensitive to the regularization of network parameters are less (more, respectively) certain. This principle can be implemented by deterministically tweaking the training loss during the fine-tuning phase and reflects confidence in the output as a function of all layers of the network. We show that regularization variation (RegVar) provides rigorous uncertainty estimates that, in the infinitesimal limit, exactly recover the Laplace approximation in Bayesian deep learning. We demonstrate its success in several deep learning architectures, showing it can scale tractably with the network size while maintaining or improving uncertainty quantification quality. Our experiments across multiple datasets show that RegVar not only identifies uncertain predictions effectively but also provides insights into the stability of learned representations.

stat.ML

Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning

We propose training fitted Q-iteration with log-loss (FQI-log) for batch reinforcement learning (RL). We show that the number of samples needed to learn a near-optimal policy with FQI-log scales with the accumulated cost of the optimal policy, which is zero in problems where acting optimally achieves the goal and incurs no cost. In doing so, we provide a general framework for proving small-cost bounds, i.e. bounds that scale with the optimal achievable cost, in batch RL. Moreover, we empirically verify that FQI-log uses fewer samples than FQI trained with squared loss on problems where the optimal policy reliably achieves the goal.

cs.LG

Orisometry formalism reveals duality and exotic nonuniform response in origami sheets

Origami metamaterial design enables drastic qualitative changes in the response properties of a thin sheet via the addition of a repeating pattern of folds based around a rigid folding motion. Known also as a mechanism, this folding motion will have a very small energy cost when applied uniformly; and yet uniform activation of such remains highly difficult to observe, these sheets instead generically displaying nonuniform response patterns which are not yet well understood. Here, we present a purely geometric continuum theory which captures the nonuniform, nonlinear response to generic loading as composed locally of the planar mechanism, as well as previously identified ``twist'' and ``bend'' modes which enable the patterned sheet to curve out of the plane across long distances. Our numerical analysis confirms that these three modes govern the observed nonuniform response, varying smoothly across the sheet according to three PDEs which guarantee compatibility. In analogy with the recently solved case of planar mechanism metamaterials, these ``Orisometries'' (origami + isometries, so named by us) are subextensive but infinite in number, with each mode displaying, in the linear limit, ``sheared analytic'' spatial patterns which are controlled by the Poisson's ratio of the uniform folding mechanism. Furthermore, the ``planar'' mechanism-based deformation patterns superimpose with a mathematically dual space of ``non-planar'' twist/bend deformations to span the available soft linear response. Together, our findings furnish the first quantification of the number of soft response modes available, as well as the first intuitive quantification of their spatial distribution.

cond-mat.soft

Rigidity percolation in a random tensegrity via analytic graph theory

Functional structures from across the engineered and biological world combine rigid elements such as bones and columns with flexible ones such as cables, fibers and membranes. These structures are known loosely as tensegrities, since these cable-like elements have the highly nonlinear property of supporting only extensile tension. Marginally rigid systems are of particular interest because the number of structural constraints permits both flexible deformation and the support of external loads. We present a model system in which tensegrity elements are added at random to a regular backbone. This system can be solved analytically via a directed graph theory, revealing a novel mechanical critical point generalizing that of Maxwell. We show that even the addition of a few cable-like elements fundamentally modifies the nature of this transition point, as well as the later transition to a fully rigid structure. Moreover, the tensegrity network displays a fundamentally new collective avalanche behavior, in which the addition of a single cable leads to the elimination of multiple floppy modes, a phenomenon that becomes dominant at the transition point. These phenomena have implications for systems with nonlinear mechanical constraints, from biopolymer networks to soft robots to jammed packings to origami sheets.

cond-mat.soft

The Implicit Delta Method

Epistemic uncertainty quantification is a crucial part of drawing credible conclusions from predictive models, whether concerned about the prediction at a given point or any downstream evaluation that uses the model as input. When the predictive model is simple and its evaluation differentiable, this task is solved by the delta method, where we propagate the asymptotically-normal uncertainty in the predictive model through the evaluation to compute standard errors and Wald confidence intervals. However, this becomes difficult when the model and/or evaluation becomes more complex. Remedies include the bootstrap, but it can be computationally infeasible when training the model even once is costly. In this paper, we propose an alternative, the implicit delta method, which works by infinitesimally regularizing the training loss of the predictive model to automatically assess downstream uncertainty. We show that the change in the evaluation due to regularization is consistent for the asymptotic variance of the evaluation estimator, even when the infinitesimal change is approximated by a finite difference. This provides both a reliable quantification of uncertainty in terms of standard errors as well as permits the construction of calibrated confidence intervals. We discuss connections to other approaches to uncertainty quantification, both Bayesian and frequentist, and demonstrate our approach empirically.

stat.ML

Omnimodal topological polarization of bilayer networks: analysis in the Maxwell limit and experiments on a 3D-printed prototype

Periodic networks on the verge of mechanical instability, called Maxwell lattices, are known to exhibit zero-frequency modes localized to their boundaries. Topologically polarized Maxwell lattices, in particular, focus these zero modes to one of their boundaries in a manner that is protected against disorder by the reciprocal-space topology of the lattice's band structure. Here, we introduce a class of mechanical bilayers as a model system for designing topologically protected edge modes that couple in-plane dilational and shearing modes to out-of-plane flexural modes, a paradigm that we refer to as omnimodal polarization. While these structures exhibit a high-dimensional design space that makes it difficult to predict the topological polarization of generic geometries, we are able to identify a family of mirror-symmetric bilayers that inherit the in-plane modal localization of their constitutive monolayers whose topological polarization can be determined analytically. Importantly, the coupling between the layers results in the emergence of omnimodal polarization, whereby in-plane and out-of-plane edge modes localize on the same edge. We demonstrate these theoretical results by fabricating a mirror-symmetric, topologically polarized kagome bilayer consisting of a network of elastic beams via additive manufacturing and confirm this finite-frequency polarization via finite element analysis and laser-vibrometry experiments.

cond-mat.soft

Band theory and boundary modes of high-dimensional representations of infinite hyperbolic lattices

Periodic lattices in hyperbolic space are characterized by symmetries beyond Euclidean crystallographic groups, offering a new platform for classical and quantum waves, demonstrating great potentials for a new class of topological metamaterials. One important feature of hyperbolic lattices is that their translation group is nonabelian, permitting high-dimensional irreducible representations (irreps), in contrast to abelian translation groups in Euclidean lattices. Here we introduce a general framework to construct wave eigenstates of high-dimensional irreps of infinite hyperbolic lattices, thereby generalizing Bloch's theorem, and discuss its implications on unusual mode-counting and degeneracy, as well as bulk-edge correspondence in hyperbolic lattices. We apply this method to a mechanical hyperbolic lattice, and characterize its band structure and zero modes of high-dimensional irreps.

cond-mat.mes-hall

Discrete symmetries control mechanical response in parallelogram-based origami

Geometric compatibility constraints dictate the mechanical response of soft systems that can be utilized for the design of mechanical metamaterials such as the negative Poisson ratio Miura-ori origami crease pattern. Here, we develop a formalism for linear compatibility that enables explicit investigation of the interplay between geometric symmetries and functionality in origami crease patterns. We apply this formalism to a particular class of periodic crease patterns with unit cells composed of four arbitrary parallelogram faces and establish that their mechanical response is characterized by an anticommuting symmetry. In particular, we show that the modes are eigenstates of this symmetry operator and that these modes are simultaneously diagonalizable with the symmetric strain operator and the antisymmetric curvature operator. This feature reveals that the anticommuting symmetry defines an equivalence class of crease pattern geometries which possess equal and opposite in-plane and out-of-plane Poisson's ratios.

cond-mat.soft

Locomotion without force, and impulse via dissipation: Robotic swimming in curved space via geometric phase

Locomotion by shape changes (spermatozoon swimming, snake slithering, bird flapping) or gas expulsion (rocket firing) is assumed to require environmental interaction, due to conservation of momentum. As first noted in (Wisdom, 2003) and later in (Guéron, 2009) and (Avron et al, 2006), in curved space or spacetime the non-commutativity of translations permits translation without momentum exchange, just as falling cats and lizards can self-deform to reorient in flat space without environmental interaction. Translation in curved space can occur not only in gravitationally induced curved spacetime (where translation is predicted to be on the order of $10^{-23}$ m per gait cycle) but also in the curved surfaces encountered by locomotors in real-world environments. Here we show that a precision robophysical apparatus consisting of motors driven on curved tracks (and thereby confined to a spherical surface without a solid substrate) can self-propel without environmental momentum exchange (impulse) via shape changes that can generate gauge potentials that manifest as translations. Our system produces shape changes comparable to the environment's inverse curvatures and generates from zero momentum forward movement of $10^{-1}$ cm per gait cycle even while resisted by weak gravitational and frictional forces. Dissipation via friction eventually arrests the robot but also imbues it with momentum which can be released upon a cessation of shape changes. This work demonstrates how the interaction between environmental curvature, active driving and geometric phases yields rich, exotic phenomena.

cond-mat.soft

Residual Overfit Method of Exploration

Exploration is a crucial aspect of bandit and reinforcement learning algorithms. The uncertainty quantification necessary for exploration often comes from either closed-form expressions based on simple models or resampling and posterior approximations that are computationally intensive. We propose instead an approximate exploration methodology based on fitting only two point estimates, one tuned and one overfit. The approach, which we term the residual overfit method of exploration (ROME), drives exploration towards actions where the overfit model exhibits the most overfitting compared to the tuned model. The intuition is that overfitting occurs the most at actions and contexts with insufficient data to form accurate predictions of the reward. We justify this intuition formally from both a frequentist and a Bayesian information theoretic perspective. The result is a method that generalizes to a wide variety of models and avoids the computational overhead of resampling or posterior approximations. We compare ROME against a set of established contextual bandit methods on three datasets and find it to be one of the best performing.

cs.LG

Hidden symmetries generate rigid folding mechanisms in periodic origami

We consider the zero-energy deformations of periodic origami sheets with generic crease patterns. Using a mapping from the linear folding motions of such sheets to force-bearing modes in conjunction with the Maxwell-Calladine index theorem we derive a relation between the number of linear folding motions and the number of rigid body modes that depends only on the average coordination number of the origami's vertices. This supports the recent result by Tachi which shows periodic origami sheets with triangular faces exhibit two-dimensional spaces of rigidly foldable cylindrical configurations. We also find, through analytical calculation and numerical simulation, branching of this configuration space from the flat state due to geometric compatibility constraints that prohibit finite Gaussian curvature. The same counting argument leads to pairing of spatially varying modes at opposite wavenumber in triangulated origami, preventing topological polarization but permitting a family of zero energy deformations in the bulk that may be used to reconfigure the origami sheet.

cond-mat.soft

Counterfactual Evaluation of Slate Recommendations with Sequential Reward Interactions

Users of music streaming, video streaming, news recommendation, and e-commerce services often engage with content in a sequential manner. Providing and evaluating good sequences of recommendations is therefore a central problem for these services. Prior reweighting-based counterfactual evaluation methods either suffer from high variance or make strong independence assumptions about rewards. We propose a new counterfactual estimator that allows for sequential interactions in the rewards with lower variance in an asymptotically unbiased manner. Our method uses graphical assumptions about the causal relationships of the slate to reweight the rewards in the logging policy in a way that approximates the expected sum of rewards under the target policy. Extensive experiments in simulation and on a live recommender system show that our approach outperforms existing methods in terms of bias and data efficiency for the sequential track recommendations problem.

cs.LG

Variational Tempering

Variational inference (VI) combined with data subsampling enables approximate posterior inference over large data sets, but suffers from poor local optima. We first formulate a deterministic annealing approach for the generic class of conditionally conjugate exponential family models. This approach uses a decreasing temperature parameter which deterministically deforms the objective during the course of the optimization. A well-known drawback to this annealing approach is the choice of the cooling schedule. We therefore introduce variational tempering, a variational algorithm that introduces a temperature latent variable to the model. In contrast to related work in the Markov chain Monte Carlo literature, this algorithm results in adaptive annealing schedules. Lastly, we develop local variational tempering, which assigns a latent temperature to each data point; this allows for dynamic annealing that varies across data. Compared to the traditional VI, all proposed approaches find improved predictive likelihoods on held-out data.

stat.ML