SearcharxivSearch

arXiv subjects

Jianfeng Lu

Publications and source records attributed to Jianfeng Lu.

At least 19 recordsLinked to original sources

Accelerated High-Accuracy Sampling from a Warm Start via the Proximal Bouncy Particle Sampler

We study the problem of sampling from $\mu(\mathrm{d}x)\propto e^{-V(x)}\,\mathrm{d}x$ on $\mathbb{R}^d$, where $V$ is $\alpha$-strongly convex and $\beta$-smooth, and write $\kappa:=\beta/\alpha$. We design and analyze the Proximal Bouncy Particle Sampler (Proximal BPS), a new sampler that combines ideas from the proximal sampler and the bouncy particle sampler. From a warm start initialization with $ O(1) $ R\'enyi divergence w.r.t. $\mu$, Proximal BPS returns a sample whose law is $\varepsilon$-close to $\mu$ in total variation distance using $\widetilde O(\sqrt\kappa\,d^{1/4} \,\mathrm{polylog}(1/\varepsilon))$ gradient queries in expectation.

math.ST

Smoothed Picard Hamiltonian Monte Carlo

We develop a new low-accuracy sampler, called \emph{smoothed Picard Hamiltonian Monte Carlo}, which combines Gaussian smoothing, Picard iteration, and higher-order discretization. For a log-concave target $\pi \propto \exp(-V)$ in dimension $d$ satisfying $0 \prec \alpha I \preceq \nabla^2 V \preceq \beta I$, with condition number $\kappa := \beta/\alpha$, smoothed Picard HMC returns a sample with $\sqrt \alpha\,W_2(\cdot,\pi) \le \varepsilon$ using $\widetilde O(\kappa^2 + \kappa^{7/6} d^{1/6}/\varepsilon^{1/3})$ gradient queries. We also prove stronger $W_q$ bounds, and then develop an algorithmic framework, the recursive warm start generator, to upgrade these $W_q$ bounds to stronger divergence guarantees. This produces a warm start for the proximal bouncy particle sampler, introduced in a companion work, leading to a high-accuracy log-concave sampler with complexity $\widetilde O((\kappa^{7/6} d^{1/6} + \kappa^{1/2} d^{1/4})\mathrm{polylog}(1/\varepsilon))$.

math.ST

Windowed thinning and query complexity for the bouncy particle and Zigzag samplers

Let $\mu(d x)\propto e^{-U(x)} d x$ on $\R^d$, where $U$ is $m$-strongly convex and $L$-smooth, and denote by $\kappa=L/m$ the condition number. We consider windowed thinning, an exact simulation method for the bouncy particle sampler and the coordinate Zigzag process. The method divides a trajectory into deterministic windows and uses a gradient evaluation at the beginning of each window to construct a tractable local envelope for the event rate. Combining this construction with quantitative mixing estimates and finite-time bounds on the expected numbers of bounces and flips yields query complexity guarantees from a Gaussian cold start. For total-variation error $\varepsilon$, the expected query counts are $O(\kappa^{1/2}d\,(d\log\kappa+\log\frac1\varepsilon))$ gradient queries for the bouncy particle sampler and $O(\kappa d^{1/4}(d\log\kappa+\log\frac1\varepsilon))$ full-gradient equivalents for Zigzag, where $d$ coordinate-partial queries count as one equivalent.

math.NA

Log-Sobolev inequalities for boundary-driven anharmonic chains

We study the non-equilibrium steady state of a weakly anharmonic chain of $N$ oscillators driven at its boundary by Langevin thermostats at unequal temperatures. Under a perturbative weak-anharmonicity condition, we prove a full-gradient logarithmic Sobolev inequality whose constant is independent of the chain length $N$. For homogeneous pinned chains, an additional quantitative regularity assumption yields a boundary space-time logarithmic Sobolev inequality and relative-entropy decay on the same $O(N^3)$ relaxation time scale as the harmonic chain. The proof extracts a finite-dimensional Gaussian component from the boundary noise and compares conditional terminal-state laws by a change of variables. The estimates are uniform over bounded positive temperatures and require no near-equilibrium assumption on their difference.

math-ph

A Mathematical Introduction to Diffusion Models

These notes give a proof-oriented introduction to diffusion models from the viewpoint of sampling, tracing a single arc from classical sampling dynamics to modern diffusion samplers, their error analysis, and inference-time control. Throughout, the material is layered into core definitions and identities proved in full, representative estimates proved under simplifying assumptions, and research-level theorems stated with a proof roadmap. The intended audience is beginning graduate students with a background in probability but no prior exposure to stochastic differential equations, stochastic numerics, or diffusion models.

cs.LG

Empowering Feed-Forward Reconstruction Models with Metric Scale via Satellite Images

Feed-forward 3D reconstruction models have recently shown strong generalization across diverse scenes, yet most of them recover geometry only up to an unknown global scale. This scale ambiguity limits their use in applications that require metric understanding of the environment. Existing metric reconstruction methods commonly rely on large-scale metric annotations or accurate camera calibration, both of which are costly or unreliable in many real-world settings. We propose a satellite-guided framework for resolving scale ambiguity in feed-forward 3D reconstruction. The key idea is to use readily available satellite imagery as a global metric reference. Given a coarse camera pose, our method retrieves a local satellite patch and integrates it with a feed-forward reconstruction backbone through bidirectional cross-view interaction. By enforcing consistency between the reconstructed scene and the satellite reference, the model infers absolute scale, refines scene geometry, and estimates camera pose in a metric coordinate frame. Experiments on KITTI, nuScenes, and Oxford RobotCar show consistent improvements in metric depth estimation, multi-view point-cloud reconstruction, and cross-view camera localization, while preserving strong generalization across datasets and geographic regions.

cs.CV

Space-Time Log-Sobolev Inequality and Hypocoercive Hypercontractivity for Underdamped Langevin Dynamics

We study hypercontractivity for underdamped Langevin dynamics with a convex confining potential whose spatial Gibbs marginal satisfies a logarithmic Sobolev inequality (LSI). Unlike in the overdamped case, the noise acts only on the velocity variable, so the usual LSI-based argument does not apply. Nevertheless, under an additional tame Hessian-growth assumption on the potential, we prove that, when the spatial LSI constant is $\rho$ and the friction parameter is of order $\sqrt{\rho}$, the semigroup satisfies a Gross-type $L^p$-to-$L^q$ estimate in which the integrability exponent grows exponentially on the kinetic time scale $t \sim \rho^{-1/2}$. The central ingredient, of independent interest, is a space-time LSI, valid for controlled kinetic paths of finite action, which quantifies how dissipation in the velocity variable is transferred to the position variable. We derive it from a controlled version of the hypocoercive entropy-decay estimate, and we then restore the Gross hypercontractivity mechanism through a duality argument built on a non-reversible forward/backward interpolation of the underdamped Langevin semigroup. As a corollary, R\'{e}nyi divergences of every order decay exponentially at the sharp hypocoercive rate $\mathcal{O}(\sqrt{\rho})$.

math.AP

Long-time reverse transportation inequalities for non-globally-dissipative Langevin dynamics

We establish a dimension-free, uniform-in-time reverse transportation inequality for Langevin dynamics with non-convex potentials. This inequality controls the Rényi divergence of arbitrary order between the process distributions starting from distinct initial points and serves as the dual version of the Harnack inequality. Notably, we prove that this inequality retains exponential decay in the long-time regime, thereby extending existing results for log-concave sampling to the non-convex setting.

math.PR

HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

While hyperspectral imaging provides rich spatial-spectral information across hundreds of narrow wavelength bands for precise material identification, ground-based hyperspectral pre-trained backbones remain absent, constrained by varying spectral configurations across sensors, limited annotations and heterogeneous labeling schemes, and the limited scale and scene diversity of existing datasets. To address these challenges and enable universal perception, we propose HyperVision, the first ground-based hyperspectral pre-trained backbone. First, to handle varying spectral configurations, HyperVision adopts a channel-adaptive dynamic embedding mechanism to map heterogeneous inputs into a unified token space. Second, we develop an unsupervised representation learning framework. Specifically, to address limited annotations and heterogeneous labeling schemes, a multi-source pseudo-labeling method is introduced to fuse spatial structures from SAM2 and fine-grained spectral material information from HyperFree. Furthermore, to enrich scene diversity and compensate for limited dataset scale, a cross-modal knowledge distillation mechanism is utilized to transfer rich semantic representations from a pre-trained RGB vision model to our backbone. Pre-trained on a collection of 15k images from 26 diverse ground-based datasets, HyperVision demonstrates exceptional generalization. Requiring only efficient head-only adaptation without adjusting backbone parameters, it outperforms state-of-the-art task-specific methods across three downstream tasks under varying sensor configurations, yielding up to a 16.3% relative improvement in hyperspectral semantic segmentation $\mathrm{Acc}_{\mathrm{M}}$, a 2.1% relative gain in object tracking AUC, and a 35.5% reduction in salient object detection MAE. The source code and pre-trained models are available at https://github.com/lronkitty/HyperVision.

cs.CV

A sharp hypocoercive entropy decay estimate for underdamped Langevin dynamics

We study the underdamped Langevin dynamics with invariant measure $μ(\,\mathrm{d}x\,\mathrm{d}v)\propto \mathrm{e}^{-U(x)-\lvert v\rvert^2/2}\,\mathrm{d}x\,\mathrm{d}v$. Assume that the position marginal $μ_x(\,\mathrm{d}x)\propto \mathrm{e}^{-U(x)}\,\mathrm{d}x$ satisfies a logarithmic Sobolev inequality with constant $ρ>0$, and that $U$ is convex on $\mathbb{R}^d$ and satisfies some growth conditions. We introduce a modified entropy approach with a Wasserstein entropy-current corrector \begin{equation*} \mathcal H_ε(g)=\operatorname{Ent}_μ(g) +ε\int Π_v(v\,g)\cdot\bigl(x-T_q(x)\bigr)\,μ_x(\mathrm{d}x), \end{equation*} where $Π_v$ denotes averaging over the velocity variable against the standard Gaussian $κ(\mathrm{d}v)=(2π)^{-d/2}\mathrm{e}^{-\lvert v\rvert^2/2}\,\mathrm{d}v$, $q=Π_v g$ is the position marginal density of $g$, and $T_q$ is the Brenier optimal transport map from $qμ_x$ to $μ_x$. For friction $γ=Γ\sqrtρ$ with $Γ>0$, and for any initial law $p_0$ with finite relative entropy, if $p_t$ denotes the law of underdamped Langevin dynamics at time $t$, we establish the explicit entropy decay \begin{equation*} \operatorname{Ent}(p_t\midμ) \leq \frac{1+θ}{1-θ}\,\mathrm{e}^{-Λt}\,\operatorname{Ent}(p_0\midμ), \qquad t\ge0, \end{equation*} with rate \begin{equation*} Λ=\fracθ{2(1+θ)}\sqrtρ, \qquad θ=\min\Bigl\{\tfracΓ{12},\tfrac{1}{4Γ}\Bigr\}. \end{equation*} In particular, the entropy convergence rate has optimal $\sqrtρ$ order.

math.AP

CDFCI: High-Performance Parallel Software for Many-Body Large-Scale Eigenvalue Problems

CDFCI is a shared-memory parallel numerical program for computing low-lying eigenpairs of large-scale, non-relativistic fermionic Hamiltonians. The software is designed to handle a broad class of many-body quantum models, including both ab initio electronic structure Hamiltonians and lattice-based Hamiltonians arising in condensed matter physics. CDFCI combines an efficient coordinate-descent-based selected configuration interaction algorithm with dedicated parallelization strategies, achieving high performance on modern multi-core architectures. Benchmark results on representative quantum chemistry and condensed matter test cases demonstrate that CDFCI attains state-of-the-art accuracy with competitive performance compared to established selected configuration interaction (such as CIPSI or SHCI) and DMRG implementations. The software is open-source, extensively documented, and provides a Python interface for seamless integration with PySCF and other many-body simulation workflows.

physics.comp-ph

Mirror Descent for Deterministic Optimal Control

We study an explicit mirror-descent method for finite-horizon deterministic optimal control problems. The method is motivated by Pontryagin's maximum principle: at each iteration, one solves the state and adjoint equations and updates the control by maximizing a first-order approximation of the regularized Hamiltonian penalized by a Bregman divergence. In the Euclidean case, the update reduces to a projected gradient step in the control variable. Under global smoothness assumptions and uniform convexity of the mirror map, we prove a relative smoothness estimate for the cost functional and derive an energy dissipation inequality for sufficiently small step sizes. Under an additional concavity assumption on the unregularized Hamiltonian and convexity of the terminal cost, we establish relative convexity of the regularized objective. These estimates yield an $O(1/n)$ convergence rate in the unregularized convex case and a geometric rate when the control regularization parameter is positive. Numerical examples illustrate the behavior of the method in linear-quadratic, degenerate convex, and nonlinear high-dimensional settings.

math.OC

Guidance for twisted particle filter: a continuous-time perspective

The particle filter (PF), also known as sequential Monte Carlo (SMC), approximates high-dimensional probability distributions and their normalizing constants in the discrete-time setting. To reduce the variance of the Monte Carlo approximation, various twisted particle filters (TPFs) have been proposed, in which a twisting function is chosen or learned to modify the Markov transition kernel. Guided by existing control-based importance sampling algorithms in the continuous-time setting, we propose a novel algorithm called the ``Twisted-Path Particle Filter'' (TPPF), in which the twisting function is parameterized by a neural network and trained to minimize a specific KL-divergence between path measures. Numerical experiments illustrate the capability of the proposed algorithm.

stat.CO

Sharp hypocoercive convergence estimates for underdamped Langevin dynamics via the modified $L^2$ method

In this note, we consider the underdamped Langevin dynamics with invariant measure $\mu(\mathrm{d}x\,\mathrm{d}v) \propto e^{-U(x)-|v|^2/2}\,\mathrm{d}x\,\mathrm{d}v$. Assume that the position marginal $\mu_x(\mathrm{d}x)\propto e^{-U(x)}\,\mathrm{d}x$ satisfies a Poincar\'{e} inequality with constant $m>0$, and that $\nabla^2 U\ge -K\,\mathrm{Id}$ for some $K\ge 0$. We revisit the modified $L^2$ method of Dolbeault--Mouhot--Schmeiser, employing a shifted corrector \begin{equation*} A_\alpha=(\alpha- L_{\mathrm o})^{-1}(L_a\Pi_v)^*, \qquad \alpha\ge0, \end{equation*} where ${L}_{\mathrm{o}}=\Delta_x-\nabla U\cdot\nabla_x$ is the overdamped generator, ${L}_a$ is the generator of the Hamiltonian flow, and $\Pi_v$ denotes averaging over the velocity variable. We establish an explicit hypocoercive $L^2$-convergence rate $\Lambda_{\alpha,\gamma}$ for every shift $\alpha\ge0$ and friction coefficient $\gamma>0$, and show that, for each fixed $\gamma$, the rate is maximized at $\alpha=0$. Optimizing further over $\gamma$ gives \begin{equation*} \Lambda_{0,\gamma_*} = \frac{\sqrt m} {2\left(2+\sqrt{2+\frac{2K}{m}} +\sqrt{6+\frac{2K}{m}}\right)}\,, \qquad \text{at} \quad \gamma_*=\sqrt{6m+2K}. \end{equation*} For convex $U$, this recovers the optimal $O(\sqrt m)$ rate.

math.AP

Quantum Fisher information matrix via its classical counterpart from random measurements

Preconditioning with the quantum Fisher information matrix (QFIM) is a popular approach in quantum variational algorithms. Yet the QFIM is costly to obtain directly, usually requiring more state preparation than its classical counterpart: the classical Fisher information matrix (CFIM). It is known that averaging the classical Fisher information matrix over Haar-random measurement bases yields $\mathbb{E}_{U\simμ_H}[F^U(\boldsymbolθ)] = \frac{1}{2}Q(\boldsymbolθ)$ for pure states in $\mathbb{C}^N$. In this paper, we review this identity by revealing its connection to covariant measurement in quantum metrology. Furthermore, we go beyond this and obtain the exact variance of CFIM ($O(N^{-1})$), estimate its moment, and establish non-asymptotic concentration bounds ($\exp(-Θ(N)t^2)$), demonstrating that using few random measurement bases is sufficient to approximate the QFIM accurately in high-dimensional settings. This work establishes a solid theoretical foundation for efficient quantum natural gradient methods via randomized measurements.

quant-ph

Quantum Gibbs sampling through the detectability lemma

Gibbs state preparation is an important subroutine in quantum computing. In this work we use the detectability lemma to improve Gibbs state preparation. Specifically, we design new Gibbs state preparation methods that do not rely on simulating Lindbladian evolution, thus avoiding the overhead from it. For local Lindbladians consisting of $M$ terms, this approach reduces the cost by a factor of $O(M)$. We also combine the detectability lemma operator and quantum singular value transformation to implement ground state projection operators of frustration-free Hamiltonians, resulting in a quadratic speedup in the spectral gap dependence. Applying this method to Lindbladians for the Gibbs state of local commuting Hamiltonians, we achieve quadratically better dependence on the Lindbladian spectral gap.

quant-ph

State-space models through the lens of ensemble control

State-space models (SSMs) are effective architectures for sequential modeling, but a rigorous theoretical understanding of their training dynamics is still lacking. In this work, we formulate the training of SSMs as an ensemble optimal control problem, where a shared control law governs a population of input-dependent dynamical systems. We derive Pontryagin's maximum principle (PMP) for this ensemble control formulation, providing necessary conditions for optimality. Motivated by these conditions, we introduce an algorithm based on the method of successive approximations. We prove convergence of this iterative scheme along a subsequence and establish sufficient conditions for global optimality. The resulting framework provides a control-theoretic perspective on SSM training.

math.OC

Multimodal Classification via Total Correlation Maximization

Multimodal learning integrates data from diverse sensors to effectively harness information from different modalities. However, recent studies reveal that joint learning often overfits certain modalities while neglecting others, leading to performance inferior to that of unimodal learning. Although previous efforts have sought to balance modal contributions or combine joint and unimodal learning, thereby mitigating the degradation of weaker modalities with promising outcomes, few have examined the relationship between joint and unimodal learning from an information-theoretic perspective. In this paper, we theoretically analyze modality competition and propose a method for multimodal classification by maximizing the total correlation between multimodal features and labels. By maximizing this objective, our approach alleviates modality competition while capturing inter-modal interactions via feature alignment. Building on Mutual Information Neural Estimation (MINE), we introduce Total Correlation Neural Estimation (TCNE) to derive a lower bound for total correlation. Subsequently, we present TCMax, a hyperparameter-free loss function that maximizes total correlation through variational bound optimization. Extensive experiments demonstrate that TCMax outperforms state-of-the-art joint and unimodal learning approaches. Our code is available at https://github.com/hubaak/TCMax.

cs.CV