SearcharxivSearch

arXiv subjects

Atsushi Yamamura

Publications and source records attributed to Atsushi Yamamura.

10 recordsLinked to original sources

Reshaping Global Loop Structure to Accelerate Local Optimization by Smoothing Rugged Landscapes

Probabilistic graphical models with frustration exhibit rugged energy landscapes that trap iterative optimization dynamics. These landscapes are shaped not only by local interactions, but crucially also by the global loop structure of the graph. The famous Bethe approximation treats the graph as a tree, effectively ignoring global structure, thereby limiting its effectiveness for optimization. Loop expansions capture such global structure in principle, but are often impractical due to combinatorial explosion. The $M$-layer construction provides an alternative: make $M$ copies of the graph and reconnect edges between them uniformly at random. This provides a controlled sequence of approximations from the original graph at $M=1$, to the Bethe approximation as $M \rightarrow \infty$. Here we generalize this construction by replacing uniform random rewiring with a structured mixing kernel $Q$ that sets the probability that any two layers are interconnected. As a result, the global loop structure can be shaped without modifying local interactions. We show that, after this copy-and-reconnect transformation, there exists a regime in which layer-to-layer fluctuations decay, increasing the probability of reaching the global minimum of the energy function of the original graph. This yields a highly general and practical tool for optimization. Using this approach, the computational cost required to reach these optimal solutions is reduced across sparse and dense Ising benchmarks, including spin glasses and planted instances. When combined with replica-exchange Monte Carlo, the same construction increases the polynomial-time algorithmic threshold for the maximum independent set problem. A cavity analysis shows that structured inter-layer coupling significantly smooths rugged energy landscapes by collapsing configurational complexity and suppressing many suboptimal metastable states.

cond-mat.dis-nn

The geometry and dynamics of annealed optimization in the coherent Ising machine with hidden and planted solutions

The coherent Ising machine (CIM) is a nonconventional hardware architecture for finding approximate solutions to large-scale combinatorial optimization problems. It operates by annealing a laser gain parameter to adiabatically deform a high-dimensional energy landscape over a set of soft spins, going from a simple convex landscape to the more complex optimization landscape of interest. We address how the evolving energy landscapes guides the optimization dynamics against problems with hidden planted solutions. We study the Sherrington-Kirkpatrick spin-glass with ferromagnetic couplings that favor a hidden configuration by combining the replica method, random matrix theory, the Kac-Rice method and dynamical mean field theory. We characterize energy, number, location, and Hessian eigenspectra of global minima, local minima, and critical points as the landscape evolves. We find that low energy global minima develop soft-modes which the optimization dynamics can exploit to descend the energy landscape. Even when these global minima are aligned to the hidden configuration, there can be exponentially many higher energy local minima that are all unaligned with the hidden solution. Nevertheless, the annealed optimization dynamics can evade this cloud of unaligned high energy local minima and descend near to aligned lower energy global minima. Eventually, as the landscape is further annealed, these global minima become rigid, terminating any further optimization gains from annealing. We further consider a second optimization problem, the Wishart planted ensemble, which contains a hidden planted solution in a landscape with tunable ruggedness. We describe CIM phase transitions between recoverability and non-recoverability of the hidden solution. Overall, we find intriguing relations between high-dimensional geometry and dynamics in analog machines for combinatorial optimization.

cond-mat.dis-nn

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting

The deployment of artificial intelligence (AI) in critical decision-making and evaluation processes raises concerns about inherent biases that malicious actors could exploit to distort decision outcomes. We propose a systematic method to reveal such biases in AI evaluation systems and apply it to automated essay grading as an example. Our approach first identifies hidden neural activity patterns that predict distorted decision outcomes and then optimizes an adversarial input suffix to amplify such patterns. We demonstrate that this combination can effectively fool large language model (LLM) graders into assigning much higher grades than humans would. We further show that this white-box attack transfers to black-box attacks on other models, including commercial closed-source models like Gemini. They further reveal the existence of a "magic word" that plays a pivotal role in the efficacy of the attack. We trace the origin of this magic word bias to the structure of commonly-used chat templates for supervised fine-tuning of LLMs and show that a minor change in the template can drastically reduce the bias. This work not only uncovers vulnerabilities in current LLMs but also proposes a systematic method to identify and remove hidden biases, contributing to the goal of ensuring AI safety and security.

cs.CR

Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks

In this work, we reveal a strong implicit bias of stochastic gradient descent (SGD) that drives overly expressive networks to much simpler subnetworks, thereby dramatically reducing the number of independent parameters, and improving generalization. To reveal this bias, we identify invariant sets, or subsets of parameter space that remain unmodified by SGD. We focus on two classes of invariant sets that correspond to simpler (sparse or low-rank) subnetworks and commonly appear in modern architectures. Our analysis uncovers that SGD exhibits a property of stochastic attractivity towards these simpler invariant sets. We establish a sufficient condition for stochastic attractivity based on a competition between the loss landscape's curvature around the invariant set and the noise introduced by stochastic gradients. Remarkably, we find that an increased level of noise strengthens attractivity, leading to the emergence of attractive invariant sets associated with saddle-points or local maxima of the train loss. We observe empirically the existence of attractive invariant sets in trained deep neural networks, implying that SGD dynamics often collapses to simple subnetworks with either vanishing or redundant neurons. We further demonstrate how this simplifying process of stochastic collapse benefits generalization in a linear teacher-student framework. Finally, through this analysis, we mechanistically explain why early training with large learning rates for extended periods benefits subsequent generalization.

cs.LG

High-Dimensional Non-Convex Landscapes and Gradient Descent Dynamics

In these lecture notes we present different methods and concepts developed in statistical physics to analyze gradient descent dynamics in high-dimensional non-convex landscapes. Our aim is to show how approaches developed in physics, mainly statistical physics of disordered systems, can be used to tackle open questions on high-dimensional dynamics in Machine Learning.

cond-mat.dis-nn

Geometric landscape annealing as an optimization principle underlying the coherent Ising machine

Given the fundamental importance of combinatorial optimization across many diverse application domains, there has been widespread interest in the development of unconventional physical computing architectures that can deliver better solutions with lower resource costs. These architectures embed discrete optimization problems into the annealed, analog evolution of nonlinear dynamical systems. However, a theoretical understanding of their performance remains elusive, unlike the cases of simulated or quantum annealing. We develop such understanding for the coherent Ising machine (CIM), a network of optical parametric oscillators that can be applied to any quadratic unconstrained binary optimization problem. Here we focus on how the CIM finds low-energy solutions of the Sherrington-Kirkpatrick spin glass. As the laser gain is annealed, the CIM interpolates between gradient descent on the soft-spin energy landscape, to optimization on coupled binary spins. By exploiting spin-glass theory, we develop a detailed understanding of the evolving geometry of the high-dimensional CIM energy landscape as the laser gain increases, finding several phase transitions, from flat, to rough, to rigid. Additionally, we develop a cavity method that provides a precise geometric interpretation of supersymmetry breaking in terms of the response of a rough landscape to specific perturbations. We confirm our theory with numerical experiments, and find detailed information about critical points of the landscape. Our extensive analysis of phase transitions provides theoretically motivated optimal annealing schedules that can reliably find near-ground states. This analysis reveals geometric landscape annealing as a powerful optimization principle and suggests many further avenues for exploring other optimization problems, as well as other types of annealed dynamics, including chaotic, oscillatory or quantum dynamics.

cond-mat.dis-nn

The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural Networks

In this work, we explore the maximum-margin bias of quasi-homogeneous neural networks trained with gradient flow on an exponential loss and past a point of separability. We introduce the class of quasi-homogeneous models, which is expressive enough to describe nearly all neural networks with homogeneous activations, even those with biases, residual connections, and normalization layers, while structured enough to enable geometric analysis of its gradient dynamics. Using this analysis, we generalize the existing results of maximum-margin bias for homogeneous networks to this richer class of models. We find that gradient flow implicitly favors a subset of the parameters, unlike in the case of a homogeneous model where all parameters are treated equally. We demonstrate through simple examples how this strong favoritism toward minimizing an asymmetric norm can degrade the robustness of quasi-homogeneous models. On the other hand, we conjecture that this norm-minimization discards, when possible, unnecessary higher-order parameters, reducing the model to a sparser parameterization. Lastly, by applying our theorem to sufficiently expressive neural networks with normalization layers, we reveal a universal mechanism behind the empirical phenomenon of Neural Collapse.

cs.LG

Nonlinear Quantum Behavior of Ultrashort-Pulse Optical Parametric Oscillators

The quantum features of ultrashort-pulse optical parametric oscillators (OPOs) are investigated theoretically in the nonlinear regime near and above threshold. Viewing the pulsed OPO as a multimode open quantum system, we rigorously derive a general input-output model that features nonlinear coupling among many cavity (i.e., system) signal modes and a broadband single-pass (i.e., reservoir) pump field. Under appropriate assumptions, our model produces a Lindblad master equation with multimode nonlinear Lindblad operators describing two-photon dissipation and a multimode four-wave-mixing Hamiltonian describing a broadband, dispersive optical cascade, which we show is required to preserve causality. To simplify the multimode complexity of the model, we employ a supermode decomposition to perform numerical simulations in the regime where the pulsed supermodes experience strong single-photon nonlinearity. We find that the quantum nonlinear dynamics induces pump depletion as well as corrections to the below-threshold squeezing spectrum predicted by linearized models. We also observe the formation of non-Gaussian states with Wigner-function negativity and show that the multimode interactions with the pump, both dissipative and dispersive, can act as effective decoherence channels. Finally, we briefly discuss some experimental considerations for potentially observing such quantum nonlinear phenomena with ultrashort-pulse OPOs on nonlinear nanophotonic platforms.

quant-ph

Onset of non-Gaussian quantum physics in pulsed squeezing with mesoscopic fields

We study the emergence of non-Gaussian quantum features in pulsed squeezed light generation with a mesoscopic number (i.e., dozens to hundreds) of pump photons. Due to the strong optical nonlinearities necessarily involved in this regime, squeezing occurs alongside significant pump depletion, compromising the predictions made by conventional semiclassical models for squeezing. Furthermore, nonlinear interactions among multiple frequency modes render the system dynamics exponentially intractable in naïve quantum models, requiring a more sophisticated modeling framework. To this end, we construct a nonlinear Gaussian approximation to the squeezing dynamics, defining a "Gaussian interaction frame" (GIF) in which non-Gaussian quantum dynamics can be isolated and concisely described using a few dominant (i.e., principal) supermodes. Numerical simulations of our model reveal non-Gaussian distortions of squeezing in the mesoscopic regime, largely associated with signal-pump entanglement. We argue that the state of the art in nonlinear nanophotonics is quickly approaching this regime, providing an all-optical platform for experimental studies of the semiclassical-to-quantum transition in a rich paradigm of coherent, multimode nonlinear dynamics. Mesoscopic pulsed squeezing thus provides an intriguing case study of the rapid rise in dynamic complexity associated with semiclassical-to-quantum crossover, which we view as a correlate of the emergence of new information-processing capacities in the quantum regime.

quant-ph

A Quantum Model for Coherent Ising Machines: Discrete-time Measurement Feedback Formulation

Recently, the coherent Ising machine (CIM) as a degenerate optical parametric oscillator (DOPO) network has been researched to solve Ising combinatorial optimization problems. We formulate a theoretical model for the CIM with discrete-time measurement feedback processes, and perform numerical simulations for the simplest network, composed of two degenerate optical parametric oscillator pulses with the anti-ferromagnetic mutual coupling. We evaluate the extent to which quantum coherence exists during the optimization process.

quant-ph