SearcharxivSearch

arXiv subjects

Songbo Wang

Publications and source records attributed to Songbo Wang.

16 recordsLinked to original sources

Self-fictitious-play for Potential Monotone Ergodic Mean-field Games

We investigate long-time learning in ergodic, potential, monotone mean-field games (MFGs) via a self-fictitious-play (SFP) dynamics coupling an optimally controlled diffusion with a slowly evolving belief. At each time, the state follows the optimal feedback associated with the current belief, while the belief is updated using the player's own empirical occupation measure rather than the population distribution. For ergodic monotone potential MFGs on the torus, we prove that the SFP dynamics is contractive and admits a unique invariant law. Moreover, we show that this invariant law is quantitatively close to the MFG Nash equilibrium, with an error of order equal to the square root of the belief-update rate. The proof combines uniform-in-time regularity estimates for the ergodic Hamilton-Jacobi-Bellman equation with an energy argument based on the Lasry-Lions divergence. The linear-quadratic example shows that this rate is sharp, and the numerical experiments illustrate the predicted scaling.

math.OC

Rethinking Representations for Cross-Domain Infrared Small Target Detection: A Generalizable Perspective from the Frequency Domain

The accurate target-background separation in infrared small target detection (IRSTD) highly depends on the discriminability of extracted representations. However, most existing methods are confined to domain-consistent settings, while overlooking whether such discriminability can generalize to unseen domains. In practice, distribution shifts between training and testing data are inevitable due to variations in observational conditions and environmental factors. Meanwhile, the intrinsic indistinctiveness of infrared small targets aggravates overfitting to domain-specific patterns. Consequently, the detection performance of models trained on source domains can be severely degraded when deployed in unseen domains. To address this challenge, we propose a spatial-spectral collaborative perception network (S$^2$CPNet) for cross-domain IRSTD. Moving beyond conventional spatial learning pipelines, we rethink IRSTD representations from a frequency perspective and reveal inconsistencies in spectral phase as the primary manifestation of domain discrepancies. Based on this insight, we develop a phase rectification module (PRM) to derive generalizable target awareness. Then, we employ an orthogonal attention mechanism (OAM) in skip connections to preserve positional information while refining informative representations. Moreover, the bias toward domain-specific patterns is further mitigated through selective style recomposition (SSR). Extensive experiments have been conducted on three IRSTD datasets, and the proposed method consistently achieves state-of-the-art performance under diverse cross-domain settings.

cs.CV

Rethinking Semi-Supervised Node Classification with Self-Supervised Graph Clustering

The emergence of graph neural networks (GNNs) has offered a powerful tool for semi-supervised node classification tasks. Subsequent studies have achieved further improvements through refining the message passing schemes in GNN models or exploiting various data augmentation techniques to mitigate limited supervision. In real graphs, nodes often tend to form tightly-knit communities/clusters, which embody abundant signals for compensating label scarcity in semi-supervised node classification but are not explored in prior methods. Inspired by this, this paper presents NCGC that integrates self-supervised graph clustering and semi-supervised classification into a unified framework. Firstly, we theoretically unify the optimization objectives of GNNs and spectral graph clustering, and based on that, develop soft orthogonal GNNs (SOGNs) that leverage a refined message passing paradigm to generate node representations for both classification and clustering. On top of that, NCGC includes a self-supervised graph clustering module that enables the training of SOGNs for learning representations of unlabeled nodes in a self-supervised manner. Particularly, this component comprises two non-trivial clustering objectives and a Sinkhorn-Knopp normalization that transforms predicted cluster assignments into balanced soft pseudo-labels. Through combining the foregoing clustering module with the classification model using a multi-task objective containing the supervised classification loss on labeled data and self-supervised clustering loss on unlabeled data, NCGC promotes synergy between them and achieves enhanced model capacity. Our extensive experiments showcase that the proposed NCGC framework consistently and considerably outperforms popular GNN models and recent baselines for semi-supervised node classification on seven real graphs, when working with various classic GNN backbones.

cs.LG

Adaptive Overclocking: Dynamic Control of Thinking Path Length via Real-Time Reasoning Signals

Large Reasoning Models (LRMs) often suffer from computational inefficiency due to overthinking, where a fixed reasoning budget fails to match the varying complexity of tasks. To address this issue, we propose Adaptive Overclocking, a method that makes the overclocking hyperparameter $\alpha$ dynamic and context-aware. Our method adjusts reasoning speed in real time through two complementary signals: (1) token-level model uncertainty for fine-grained step-wise control, and (2) input complexity estimation for informed initialization. We implement this approach with three strategies: Uncertainty-Aware Alpha Scheduling (UA-$\alpha$S), Complexity-Guided Alpha Initialization (CG-$\alpha$I), and a Hybrid Adaptive Control (HAC) that combines both. Experiments on GSM8K, MATH, and SVAMP show that HAC achieves superior accuracy-latency trade-offs, reducing unnecessary computation on simple problems while allocating more resources to challenging ones. By mitigating overthinking, Adaptive Overclocking enhances both efficiency and overall reasoning performance.

cs.LG

Large-scale concentration and relaxation for mean-field Langevin particle systems

We study the Langevin dynamics of diffusive particles with regular pairwise interactions under mean-field scaling. By approximating empirical distributions with conditional distributions, we establish coercive and contractive properties for the modulated free energy functional. These properties yield near-optimal large-scale concentration and relaxation rates for the particle system throughout the subcritical regime. Furthermore, we derive generation of chaos estimates with the optimal order of particle approximation. As a simpler instance, we demonstrate long-time convergence of the independent projection of Langevin dynamics.

math.PR

Size of chaos for Gibbs measures of mean field interacting diffusions

We investigate Gibbs measures for diffusive particles interacting through a two-body mean field energy. By identifying a gradient structure for the conditional law, we derive sharp bounds on the size of chaos, providing a quantitative characterization of particle independence. To handle interaction forces that are unbounded at infinity, we study the concentration of measure phenomenon for Gibbs measures via a defective Talagrand inequality, which may hold independent interest. Our approach provides a unified framework for both the flat semi-convex and displacement convex cases. Additionally, we establish sharp chaos bounds for the quartic Curie-Weiss model in the sub-critical regime, demonstrating the generality of this method.

math.PR

Uniform log-Sobolev inequalities for mean field particles with flat-convex energy

The purpose of this short note is to demonstrate uniform logarithmic Sobolev inequalities for the mean field gradient particle systems associated to an energy functional that is convex in the flat sense. A defective log-Sobolev inequality was already established implicitly in a previous joint work with F. Chen and Z. Ren [arXiv:2212.03050 [math.PR]]. It remains only to tighten it by a uniform Poincar\'e inequality, which we prove by the method in a recent work of Guillin, W. Liu, L. Wu and C. Zhang [Ann. Appl. Probab., 32(3):1590-1614, 2022]. As an application, we show that the particle system exhibits the concentration of measure phenomenon in the long time.

math.PR

Sharp local propagation of chaos for mean field particles with $W^{-1,\infty}$ kernels

We present two methods to obtain $O(1/N^2)$ local propagation of chaos bounds for $N$ diffusive particles in $W^{-1,\infty}$ mean field interaction. This extends the recent finding of Lacker [Probab. Math. Phys., 4(2):377-432, 2023] to the case of singular interactions. The first method is based on a hierarchy of relative entropies and Fisher informations, and applies to the 2D viscous vortex model in the high temperature regime. Time-uniform local chaos bounds are also shown in this case. In the second method, we work on a hierarchy of $L^2$ distances and Dirichlet energies, and derive the desired sharp estimates for the same model in short time without restrictions on the temperature.

math.PR

Uniform-in-time propagation of chaos for kinetic mean field Langevin dynamics

We study the kinetic mean field Langevin dynamics under the functional convexity assumption of the mean field energy functional. Using hypocoercivity, we first establish the exponential convergence of the mean field dynamics and then show the corresponding $N$-particle system converges exponentially in a rate uniform in $N$ modulo a small error. Finally we study the short-time regularization effects of the dynamics and prove its uniform-in-time propagation of chaos property in both the Wasserstein and entropic sense. Our results can be applied to the training of two-layer neural networks with momentum and we include the numerical experiments.

math.PR

Time-uniform log-Sobolev inequalities and applications to propagation of chaos

Time-uniform log-Sobolev inequalities (LSI) satisfied by solutions of semi-linear mean-field equations have recently appeared to be a key tool to obtain time-uniform propagation of chaos estimates. This work addresses the more general settings of time-inhomogeneous Fokker-Planck equations. Time-uniform LSI are obtained in two cases, either with the bounded-Lipschitz perturbation argument with respect to a reference measure, or with a coupling approach at high temperature. These arguments are then applied to mean-field equations, where, on the one hand, sharp marginal propagation of chaos estimates are obtained in smooth cases and, on the other hand, time-uniform global propagation of chaos is shown in the case of vortex interactions with quadratic confinement potential on the whole space. In this second case, an important point is to establish global gradient and Hessian estimates, which is of independent interest. We prove these bounds in the more general situation of non-attractive logarithmic and Riesz singular interactions.

math.PR

Self-interacting approximation to McKean-Vlasov long-time limit: a Markov chain Monte Carlo method

For a certain class of McKean-Vlasov processes, we introduce proxy processes that substitute the mean-field interaction with self-interaction, employing a weighted occupation measure. Our study encompasses two key achievements. First, we demonstrate the ergodicity of the self-interacting dynamics, under broad conditions, by applying the reflection coupling method. Second, in scenarios where the drifts are negative intrinsic gradients of convex mean-field potential functionals, we use entropy and functional inequalities to demonstrate that the stationary measures of the self-interacting processes approximate the invariant measures of the corresponding McKean-Vlasov processes. As an application, we show how to learn the optimal weights of a two-layer neural network by training a single neuron.

math.PR

PETformer: Long-term Time Series Forecasting via Placeholder-enhanced Transformer

Recently, the superiority of Transformer for long-term time series forecasting (LTSF) tasks has been challenged, particularly since recent work has shown that simple models can outperform numerous Transformer-based approaches. This suggests that a notable gap remains in fully leveraging the potential of Transformer in LTSF tasks. Consequently, this study investigates key issues when applying Transformer to LTSF, encompassing aspects of temporal continuity, information density, and multi-channel relationships. We introduce the Placeholder-enhanced Technique (PET) to enhance the computational efficiency and predictive accuracy of Transformer in LTSF tasks. Furthermore, we delve into the impact of larger patch strategies and channel interaction strategies on Transformer's performance, specifically Long Sub-sequence Division (LSD) and Multi-channel Separation and Interaction (MSI). These strategies collectively constitute a novel model termed PETformer. Extensive experiments have demonstrated that PETformer achieves state-of-the-art performance on eight commonly used public datasets for LTSF, surpassing all existing models. The insights and enhancement methodologies presented in this paper serve as valuable reference points and sources of inspiration for future research endeavors.

cs.LG

Logarithmic Sobolev inequalities for non-equilibrium steady states

We consider two methods to establish log-Sobolev inequalities for the invariant measure of a diffusion process when its density is not explicit and the curvature is not positive everywhere. In the first approach, based on the Holley-Stroock and Aida-Shigekawa perturbation arguments [J. Stat. Phys., 46(5-6):1159-1194, 1987, J. Funct. Anal., 126(2):448-475, 1994], the control on the (non-explicit) perturbation is obtained by stochastic control methods, following the comparison technique introduced by Conforti [Ann. Appl. Probab., 33(6A):4608-4644, 2023]. The second method combines the Wasserstein-2 contraction method, used in [Ann. Henri Lebesgue, 6:941-973, 2023] to prove a Poincar\'e inequality in some non-equilibrium cases, with Wang's hypercontractivity results [Potential Anal., 53(3):1123-1144, 2020].

math.PR

Mean Field Optimization Problem Regularized by Fisher Information

Recently there is a rising interest in the research of mean field optimization, in particular because of its role in analyzing the training of neural networks. In this paper by adding the Fisher Information as the regularizer, we relate the regularized mean field optimization problem to a so-called mean field Schrodinger dynamics. We develop an energy-dissipation method to show that the marginal distributions of the mean field Schrodinger dynamics converge exponentially quickly towards the unique minimizer of the regularized optimization problem. Remarkably, the mean field Schrodinger dynamics is proved to be a gradient flow on the probability measure space with respect to the relative entropy. Finally we propose a Monte Carlo method to sample the marginal distributions of the mean field Schrodinger dynamics.

math.PR

Entropic Fictitious Play for Mean Field Optimization Problem

We study two-layer neural networks in the mean field limit, where the number of neurons tends to infinity. In this regime, the optimization over the neuron parameters becomes the optimization over the probability measures, and by adding an entropic regularizer, the minimizer of the problem is identified as a fixed point. We propose a novel training algorithm named entropic fictitious play, inspired by the classical fictitious play in game theory for learning Nash equilibriums, to recover this fixed point, and the algorithm exhibits a two-loop iteration structure. Exponential convergence is proved in this paper and we also verify our theoretical results by simple numerical examples.

math.OC

Uniform-in-time propagation of chaos for mean field Langevin dynamics

We study the mean field Langevin dynamics and the associated particle system. By assuming the functional convexity of the energy, we obtain the $L^p$-convergence of the marginal distributions towards the unique invariant measure for the mean field dynamics. Furthermore, we prove the uniform-in-time propagation of chaos in both the $L^2$-Wasserstein metric and relative entropy.

math.PR