SearcharxivSearch

arXiv subjects

Hongxu Chen

Publications and source records attributed to Hongxu Chen.

At least 19 recordsLinked to original sources

On Medial Quandle Coloring Link Invariant and Detecting Causality

We investigate the capability of medial quandles to detect causality in (2+1) dimensional globally hyperbolic spacetime by determining if their coloring link invariants can distinguish between the connected sum of two Hopf links and an infinite series of relevant three-component links constructed by Allen and Swenberg in 2020, who suggested that any link invariant must be able to distinguish those links for them to detect causality in the given setting. We show that these quandles fail to do so as long as $a\sim b\Leftrightarrow a*b=a$ defines an equivalence relation. The Alexander quandles are an example to which this result applies. We also give a sufficient condition for a medial quandle to distinguish between the connected sum of two Hopf links and all links in Allen-Swenberg series, and give an example of such a quandle. Inspired by this result, we also derive a generalized theorem about the coloring of medial quandles on a specific type of tangle, which helps to determine whether these quandles can distinguish between a wider range of knots and links or not.

math.GT

Quantum-Limited Symbol-Blind Channel Estimation for Coherent State Discrimination

Residual dispersion breaks temporal-mode matching in photon-starved coherent links. For equiprobable $M$-ary PSK coherent states in a known spectral mode, with unknown symbols and carrier phase, we establish the quantum limit for blind joint estimation of group delay and second-order dispersion: after eliminating the common phase, it is $4N_s\mathbf{C}$, set by the covariance of the centered generators alone. A multi-output quantum pulse gate with photon-number-resolving detection locally attains it and supports reception below the standard quantum limit under turbulent fading.

quant-ph

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pretrained main network, has shown remarkable success in visual-condition controllable generation. Despite its widespread adoption, the role of the side branch and its training efficiency remain underexplored. In this paper, we first revisit this mainstream paradigm through the lens of score-based generative modeling: 1) The main network preserves visual perceptual quality by providing a prior unconditional score. 2) The side network steers conditional control by implicitly contributing a likelihood score. Guided by this perspective, we propose LIkelihood Score Alignment (LISA), an effective regularization method that explicitly aligns the intermediate feature of the side network with an approximated likelihood score. Specifically, we first hook features from a designated layer of the side network and project them into the score latent space by a lightweight decoder. Then, we construct an approximated likelihood score target and calculate the distance between the decoder's output and this target as an additional regularization loss. Finally, we jointly optimize the side network and decoder with both standard diffusion loss and our regularization loss. Experiments across various image/video tasks, architectures, and diffusion/flow models demonstrated that LISA can not only consistently accelerate the training convergence and improve final synthetic results, but also encourage the side network's features to be more disentangled for conditional modeling with negligible additional training cost and zero extra inference cost.

cs.CV

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models

Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the fundamental feature discrepancy between semantic perception and pixel-level reconstruction. Bridging this gap requires overcoming two core challenges: endowing semantic encoders with high-fidelity reconstruction capabilities, and effectively aligning generative models with semantic spaces without relying on external teachers. To this end, we propose a novel unified multimodal framework featuring \textbf{S}emantic-\textbf{P}ixel self-alignment and \textbf{A}daptive \textbf{R}outing (\textbf{SPAR}). First, to reconcile semantic perception with pixel-level reconstruction, we introduce an asymmetric dual-stream unified tokenizer. A lightweight semantic stream anchors discriminative features, while a Transformer-augmented pixel stream recovers fine-grained visual details into a unified compact latent space. Second, to eliminate external dependencies, we propose a self-aligned generation paradigm that natively leverages this optimized tokenizer as an internal alignment teacher for the diffusion model. Furthermore, to facilitate flexible multimodal interaction within this unified space, we introduce Dynamic Token Routing, which enables each token to adaptively aggregate multi-layer MLLM features based on its distinct semantic demands. Extensive experiments demonstrate that SPAR establishes the state-of-the-art for unified architectures, achieving exceptional generation and reconstruction quality while preserving foundational visual understanding capabilities.

cs.CV

Stationary Vlasov-Poisson-Boltzmann system in a convex domain

We study the stationary and dynamical Vlasov-Poisson-Boltzmann system in a bounded, convex domain subject to a confining external potential field. For the stationary problem, we construct a unique stationary solution with an inflow boundary condition. A key difficulty is to obtain pointwise regularity for stationary solutions due to the intricate coupling between the self-consistent electric field and the Boltzmann collision operator. To overcome this issue, we establish a $W^{1,p}_{x,v}$--$\alpha C^1_{x,v}$ bootstrap framework and derive an unweighted $C^1_v$ estimate by exploiting the structure of the external potential field. We then investigate the dynamical Vlasov-Poisson-Boltzmann system near the stationary solution. We prove the global existence and uniqueness of solutions for small perturbations and establish exponential convergence toward the stationary state in weighted $L^\infty$ norms. Our results reveal the stabilizing effect of the external potential field and provide a framework for the stationary and dynamical theories of the Vlasov-Poisson-Boltzmann system in bounded domains.

math.AP

HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression

Large language models achieve remarkable performance via extended chain-of-thought (CoT) reasoning, yet this lengthy process incurs substantial inference overhead. Existing CoT compression methods struggle with inflexible manual length budgets, computationally expensive multi-stage training pipelines, and fragile scalability restricted to small models. We propose HMPO (Hybrid Median-length Policy Optimization), a cost-effective, single-stage reinforcement learning framework. HMPO efficiently compresses CoT via three synergistic components: an adaptive median-based budget derived from successful rollouts to eliminate manual tuning, a cosine-decay token reward for smooth length penalization, and a multiplicative reward formulation that substantially mitigates trivial reward hacking by strictly prioritizing answer correctness. Trained exclusively on mathematical data, HMPO generalizes seamlessly across math, code, science, and instruction-following tasks. Extensive experiments scaling from 9B to 122B parameters across dense and Mixture-of-Experts (MoE) architectures demonstrate that HMPO achieves 19%--46% token compression with negligible accuracy degradation, all while drastically reducing training costs compared to existing multi-stage baselines.

cs.LG

Half-space problem on the Boltzmann equation with zero Mach number at infinity

We study the long-time dynamics of the time-evolutionary Boltzmann equation with hard sphere collisions in the three-dimensional half-space \( \mathbb{R}^2 \times \mathbb{R}^+\), subject to diffuse reflection boundary conditions and small perturbations around a global Maxwellian equilibrium. The far-field velocity is assumed to be at rest; namely, we take the zero Mach number at infinity. In the first goal, we construct global-in-time low-regularity solutions near Maxwellians. We leverage time-decay properties along the two-dimensional tangential direction to establish polynomial decay rates of solutions matching the 2D heat equation. In the second goal, we further prove the propagation of Gevrey regularity: analyticity (Gevrey index 1) in the tangential spatial variable \(x_\parallel\), and Gevrey class with index 2 in the tangential velocity variable \(v_\parallel\), under suitably regular initial data. The proofs combine an \(L^1_k \cap L^p_k\) Fourier-space approach for decay estimates, macro-micro decomposition with \(L^2 - L^\infty\) frameworks adapted to unbounded domains, and weighted Gevrey norms to control regularity propagation, overcoming challenges from boundary effects and nonlinear interactions.

math.AP

Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation

Recent flow matching (FM) methods improve the few-shot adaptation of vision-language models, by modeling cross-modal alignment as a continuous multi-step flow. In this paper, we argue that existing FM methods are inherently constrained by incompatible geometric priors on pre-trained cross-modal features, resulting in suboptimal adaptation performance. We first analyze these methods from a polar decomposition perspective (i.e., radial and angular sub-manifolds). Under this new geometric view, we identify three overlooked limitations in them: 1) Angular dynamics distortion: The radial-angular coupling induces non-uniform speed on the angular sub-manifold, leading to regression training difficulty and extra truncation errors. 2) Radial dynamics neglect: Feature normalization discards modality confidence, failing to distinguish out-of-distribution and in-distribution data, and abandoning crucial radial dynamics. 3) Context-agnostic unconditional flow: Dataset-specific information loss during pre-trained cross-modal feature extraction remains unrecovered. To resolve these issues, we propose warped product flow matching (WP-FM), a unified Riemannian framework that reformulates alignment on a warped product manifold. Within this framework, we derive direct product flow matching (DP-FM) by introducing a constant-warping metric, which yields a decoupled cylindrical manifold (i.e., direct product manifold). DP-FM enables independent radial evolution and constant-speed angular geodesic transport, effectively eliminating angular dynamics distortion while preserving radial consistency. Meanwhile, we incorporate classifier-free guidance by conditioning the flow on the pre-trained VLMs' hidden states to inject missing dataset-specific information. Extensive results across 11 benchmarks have demonstrated that DP-FM achieves a new state-of-the-art for multi-step few-shot adaptation.

cs.CV

Universal 2-Local Symmetry-Preserving Quantum Neural Networks for Fermionic Systems

Simulating quantum many-body systems represents a fundamental challenge where classical machine learning methods are severely bottlenecked by the exponential curse of dimensionality. Variational Quantum Algorithms (VQAs) offer a native paradigm to tackle this by optimizing parameterized unitary evolutions to find the ground states of problem Hamiltonians. However, the efficacy of these VQA is deeply hindered by the challenge of balancing the preservation of critical physical symmetries with the strict constraints of hardware implementability. In this work, we address this dilemma by proposing a hardware-efficient, symmetry-preserving ansatz fortified with complete theoretical guarantees for fermionic systems, termed the Hamming Weight Preserving (HWP) ansatz. We establish the necessary and sufficient conditions for 2-local HWP operators to achieve subspace universality, formally debunking the prevailing assumption that truncation-free simulation requires complex high-order interactions. Empirical validations corroborate our theoretical guarantees, showcasing the exact approximation of arbitrary unitary matrices within the HWP subspace. Crucially, we demonstrate the exceptional versatility of the proposed approach by deploying the exact same ansatz across distinct fermionic models, including diverse molecular electronic structures and the Fermi-Hubbard model. Our proposed HWP ansatz consistently suppresses ground-state energy errors below $1 \times 10^{-10}$ Ha, achieving a level of precision that surpasses the stringent threshold of chemical accuracy by multiple orders of magnitude. This work establishes a complete, theoretically fortified 2-local framework for symmetry-preserving computation, offering a highly universal and hardware-efficient building block for advancing quantum machine learning and fermionic many-body simulations.

quant-ph

Diffusive limit of the Boltzmann equation around Rayleigh profile in the half space

This paper concerns the diffusive limit of the time evolutionary Boltzmann equation in the half space $\mathbb{T}^2\times\mathbb{R}^+$ for a small Knudsen number $\varepsilon>0$. For boundary conditions in the normal direction, it involves diffuse reflection moving with a tangent velocity proportional to $\varepsilon$ on the wall, whereas the far field is described by a global Maxwellian with zero bulk velocity. The incompressible Navier-Stokes equations, as the corresponding formal fluid dynamic limit, admit a specific time-dependent shearing solution known as the Rayleigh profile, which accounts for the effect of the tangentially moving boundary on the flow at rest in the far field. Using the Hilbert expansion method, for well-prepared initial data we construct the Boltzmann solution around the Rayleigh profile without initial singularity over any finite time interval.

math.AP

Decentralized Non-convex Stochastic Optimization with Heterogeneous Variance

Decentralized optimization is critical for solving large-scale machine learning problems over distributed networks, where multiple nodes collaborate through local communication. In practice, the variances of stochastic gradient estimators often differ across nodes, yet their impact on algorithm design and complexity remains unclear. To address this issue, we propose D-NSS, a decentralized algorithm with node-specific sampling, and establish its sample complexity depending on the arithmetic mean of local standard deviations, achieving tighter bounds than existing methods that rely on the worst-case or quadratic mean. We further derive a matching sample complexity lower bound under heterogeneous variance, thereby proving the optimality of this dependence. Moreover, we extend the framework with a variance reduction technique and develop D-NSS-VR, which under the mean-squared smoothness assumption attains an improved sample complexity bound while preserving the arithmetic-mean dependence. Finally, numerical experiments validate the theoretical results and demonstrate the effectiveness of the proposed algorithms.

math.OC

Bi-Anchor Interpolation Solver for Accelerating Generative Modeling

Flow Matching (FM) models have emerged as a leading paradigm for high-fidelity synthesis. However, their reliance on iterative Ordinary Differential Equation (ODE) solving creates a significant latency bottleneck. Existing solutions face a dichotomy: training-free solvers suffer from significant performance degradation at low Neural Function Evaluations (NFEs), while training-based one- or few-steps generation methods incur prohibitive training costs and lack plug-and-play versatility. To bridge this gap, we propose the Bi-Anchor Interpolation Solver (BA-solver). BA-solver retains the versatility of standard training-free solvers while achieving significant acceleration by introducing a lightweight SideNet (1-2% backbone size) alongside the frozen backbone. Specifically, our method is founded on two synergistic components: \textbf{1) Bidirectional Temporal Perception}, where the SideNet learns to approximate both future and historical velocities without retraining the heavy backbone; and 2) Bi-Anchor Velocity Integration, which utilizes the SideNet with two anchor velocities to efficiently approximate intermediate velocities for batched high-order integration. By utilizing the backbone to establish high-precision ``anchors'' and the SideNet to densify the trajectory, BA-solver enables large interval sizes with minimized error. Empirical results on ImageNet-256^2 demonstrate that BA-solver achieves generation quality comparable to 100+ NFEs Euler solver in just 10 NFEs and maintains high fidelity in as few as 5 NFEs, incurring negligible training costs. Furthermore, BA-solver ensures seamless integration with existing generative pipelines, facilitating downstream tasks such as image editing.

cs.CV

Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise

The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most existing works on this phenomenon focus on the convergence of optimization errors, while the analysis for generalization bounds under the heavy-tailed gradient noise remains limited. In this paper, we develop a general framework for establishing generalization bounds under heavy-tailed noise. Specifically, we introduce a truncation argument to achieve the generalization error bound based on the algorithmic stability under the assumption of bounded $p$th centered moment with $p\in(1,2]$. Building on this framework, we further provide the stability and generalization analysis for several popular stochastic algorithms under heavy-tailed noise, including clipped and normalized stochastic gradient descent, as well as their mini-batch and momentum variants.

cs.LG

BGK model for rarefied gas in a bounded domain

We study the Bathnagar-Gross-Krook (BGK) equation in a smooth bounded domain featuring a diffusive reflection boundary condition with general collision frequency. We prove that the BGK equation admits a unique global solution with an exponential convergence rate if the initial condition is a small perturbation around the global Maxwellian in the $L^\infty$ space. For the proof, we utilize the dissipative nature from the linearized BGK operator and establish an $L^2$ coercive estimate. Next, we derive the a priori estimate by obtaining an $L^\infty$ bound on the nonlinear operator; this requires a delicate analysis to manage its intrinsic nonlinear structure. Finally, we establish the $L^\infty$ stability estimate and introduce sequential arguments for the nonlinear BGK operator, thereby concluding both well-posedness and positivity.

math.AP

Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data

This paper studies the policy mirror descent (PMD) method, which is a general policy optimization framework in reinforcement learning and can cover a wide range of policy gradient methods by specifying difference mirror maps. Existing sample complexity analysis for policy mirror descent either focuses on the generative sampling model, or the Markovian sampling model but with the action values being explicitly approximated to certain pre-specified accuracy. In contrast, we consider the sample complexity of policy mirror descent with temporal difference (TD) learning under the Markovian sampling model. Two algorithms called Expected TD-PMD and Approximate TD-PMD have been presented, which are off-policy and mixed policy algorithms respectively. Under a small enough constant policy update step size, the $\tilde{O}(\varepsilon^{-2})$ (a logarithm factor about $\varepsilon$ is hidden in $\tilde{O}(\cdot)$) sample complexity can be established for them to achieve average-time $\varepsilon$-optimality. The sample complexity is further improved to $O(\varepsilon^{-2})$ (without the hidden logarithm factor) to achieve the last-iterate $\varepsilon$-optimality based on adaptive policy update step sizes.

math.OC

The Boltzmann equation in an infinite layer: spectrum and asymptotics toward the heat equation

In the paper, we develop spectral theory to analyze the sharp asymptotic behavior of solutions to the Boltzmann equation around global Maxwellians in a three-dimensional infinite layer $\mathbb{R}^2\times (-1,1)$. The isothermal diffuse reflection boundary condition is imposed on two parallel infinite planes at $x_3=\pm 1$. The main difficulties lie in the fact that the direct Fourier transform is not applicable to the vertical $x_3$-variable, and the linear collision operator $K$ loses its compactness on $L^2((-1,1)\times \R^3_v)$ although it is compact on $L^2(\R^3_v)$. By introducing a regularization operator $K_n$ via the finite-dimensional Fourier series truncation in $L^2(-1,1)$, we study the spectrum of the linearized initial-boundary value approximation problem, establish the resolvent estimates, and identify the leading diffusive eigenvalue. This spectral structure governs the sharp asymptotic dynamics of the original linear problem as $n\to \infty$, enabling us to construct the large-time behavior for the nonlinear problem and rigorously prove that the solution converges with a faster rate toward that of the two-dimensional heat equation in the horizontal direction.

math.AP

Hypocoercivity for the Linear Semiconductor Boltzmann Equation with Boundaries and Uncertainties

In this paper, we establish hypocoercivity for the semiconductor Boltzmann equation with the presence of an external electrical potential under the Maxwell boundary condition. We will construct a modified entropy Lyapunov functional, which is proved to be equivalent to some weighted norm of the corresponding function space. We then show that the entropy functional dissipates along the solutions, and the exponential decay to the equilibrium state of the system follows by a Gronwall type inequality. We also generalize our arguments to situations where uncertainties in our model arise,and the hypocoercivity method we have established is adopted to analyze the regularity of the solutions along the random space.

math.AP

Global dynamics of isothermal rarefied gas flows in an infinite layer

Let rarefied gas be confined in an infinite layer with diffusely reflecting boundaries that are isothermal and non-moving. The initial-boundary value problem on the nonlinear Boltzmann equation governing the rarefied gas flow in such setting is challenging due to unboundedness of both domain and its boundaries as well as the presence of physical boundary conditions. In the paper, we establish the global-in-time dynamics of such rarefied gas flows near global Maxwellians in three or two-dimensions. For the former case, we also prove that the solutions decay in time at a polynomial rate which is the same as that of solutions to the two-dimensional heat equation. This is the first result on global solutions of the Boltzmann equation with non-compact and diffuse boundaries.

math.AP