SearcharxivSearch

arXiv subjects

Hong Zhang

Publications and source records attributed to Hong Zhang.

At least 19 recordsLinked to original sources

Renewal process's guide to fractional Navier-Stokes equations

The Navier-Stokes equations, which remain unsolved, are crucial equations in fluid mechanics. Discovering the solutions to the Navier-Stokes equations is one of the challenging Millennium problems. In 1900, Hilbert proposed a potential approach to tackle this problem by establishing the relationship between microscopic dynamics and the macroscopic continuum equations. The key bridge is the derivation of Boltzmann equation and the theory of probability. In this paper, we shall use the collision renewal process with arbitrarily distributed waiting times to derive the Boltzmann equation for the time evolution of the probability of the velocity and the displacement of the particle, based on which we prove that the renewal process with exponential collision waiting time is equivalent to the classical Navier-Stokes equations, and that with power-law waiting time is equivalent to the fractional Navier-Stokes equations. Since the collision renewal process with arbitrarily distributed waiting times is a random process and is easy to perform the stochastic simulations of trajectories to obtain the corresponding solution, we actually find a stochastic approach to solve the classical and fractional Navier-Stokes equations.

cond-mat.stat-mech

Intermittent continuous-time random walks under renewal reset mechanism

Stochastic resetting as a practical and efficient search strategy in complex and disordered environments has long been a topic of interest to researchers. Based on the competition between jumping and resetting, this article proposes and investigates intermittent continuous-time random walks (CTRWs) under stochastic resetting, using the smaller waiting time for jump and reset as the renewal time, where the waiting times for both jump and reset can have arbitrary distributions. After each renewal event, the system will proceed with new waiting times for jump and reset regardless of their previous histories. We study the governing equation and Montroll-Weiss equation with renewal resetting, as well as the Markovian resetting for intermittent CTRWs. We prove the existence of non-equilibrium stationary states within the renewal reset mechanism when the jump and reset waiting times follow any exponential and power law distributions. For exponential and Gaussian distributed jump lengths, we examine the mean square displacements (MSDs) of particles to determine their monotonicity and asymptotic stability. Moreover, we calculate the first-arrival time to quantify search efficiency, and validate the intermittent CTRWs under renewal resetting lead to a finite mean first-arrival time (MFAT) to any fixed position for exponential jump and reset waiting time distributions (WTDs), power-law jump and exponential reset WTDs, as well as exponential jump and power-law reset WTDs. However, the MFAT diverges for power-law jump and reset WTDs. The intermittent CTRW model, which is based on the competition mechanism, can be applied to many physical scenarios, such as the foraging strategy of animals that return to their nests after an unsuccessful foraging attempt, or the work planning of intelligent robots that return to energy replenishment points after prolonged operation.

cond-mat.stat-mech

The Dynamical Instability of Rotating Boson Stars

We investigate the dynamical instability of rotating boson stars described by the Gross--Pitaevskii--Poisson equations with contact self-interactions. Through three-dimensional simulations, we confirm that the rotating boson star undergoes a quasiperiodic conversion between ring-like and twin-star-like density configurations in the early nonlinear stage. We develop a systematic linear stability analysis to identify the modes driving this instability and show that repulsive self-interactions could significantly increase the lifetime of the rotating boson stars. We further construct a three-mode Hamiltonian to describe the early nonlinear stage, which explains the quasiperiodic conversion. This analytical framework agrees reasonably well with the simulation results and provides a clear picture to understand the dynamics of rotating boson stars.

gr-qc

Retrospective Causal Attribution under Case-Control Sampling

The probability of necessity PN quantifies the probability that an exposed individual who experienced an outcome would not have experienced it in the absence of exposure. Case-control studies are an important resource for investigating etiologic questions, but their sampling design can introduce selection bias and complicate the statistical inference for PN.In this paper, we develop a nonparametric framework for identification and efficient estimation of PN under case-control sampling. With an externally supplied population outcome prevalence, we derive an exact identification formula that identifies PN under standard causal assumptions and monotonicity and yields a valid lower bound without monotonicity. For rare outcomes, we derive a more tractable approximation that requires no external prevalence information and prove that its approximation error vanishes at the order of the population outcome prevalence. We further establish the semiparametric efficiency theory for the exact and approximate functionals, propose asymptotically efficient estimators, and construct confidence intervals for the corresponding targets. The proposed approach has potential applications in biomedical and epidemiological studies where causal attribution is investigated using retrospective data.

stat.ME

Breakdown of Aharonov-Bohm cage in Rydberg synthetic lattices: the roles of inhomogeneity and long-range exchange

While the interaction-induced breakdown of Aharonov-Bohm (AB) cage is typically attributed to uniform bound-pair transport, systems with inhomogeneous exchange interactions realized with Rydberg synthetic lattices exhibit more complex dynamics. Employing the evolution-path symmetry (EPS) framework developed recently, we analyze the two-particle dynamics via path interference in Fock space. We find that a homogeneous nearest-neighbor exchange interaction cannot break the AB cage, regardless of whether the long-range exchange interaction is present or not. In contrast, we demonstrate that inhomogeneous nearest-neighbor exchange interaction breaks the destructive-interference EPS, and lifts the degeneracy of many-body compact localized states, thereby generating non-local dispersive eigenstates. Consequently, the initial state gains a non-zero overlap with these dispersive states, enabling delocalized transport. Furthermore, while long-range exchange interaction alone preserves the AB cage, its coupling with nearest-neighbor inhomogeneous exchange interaction opens non-canceling pathways that alter the diffusion profile. Our work connects microscopic path interference with macroscopic spectral reorganization, offering an analytical understanding of the mechanism underlying exchange-interaction-induced transport in Rydberg synthetic lattices.

cond-mat.quant-gas

A fully discrete LBRFD-IPDG method for linear fourth-order parabolic equations

We propose a fully discrete method for linear fourth-order parabolic equations with Dirichlet boundary conditions, combining an implicit LBRFD multistep scheme in time with a mixed interior penalty discontinuous Galerkin (IPDG) method in space. The temporal discretization employs equispaced linear barycentric rational interpolants and incorporates a startup procedure. To facilitate the spatial discretization, the original problem is reformulated through an auxiliary variable. For certain parameter pairs $(n,d)$, the LBRFD method is shown to be $A(\alpha)$-stable and to possess a wider stability angle than the corresponding BDF$p$ method of the same order. Stability and a priori error estimates are established via a $G$-energy technique and the discrete Gr\"onwall lemma. The theoretical analysis yields a total $L^2$ error estimate of order $h^{k-1}+\tau^p$, where $p=d$ if $n-d$ is even and $p=d+1$ if $n-d$ is odd. The reduced spatial convergence rate is attributed to boundary contributions on $\partial\Omega$. Despite this theoretical prediction, numerical experiments confirm the stability and demonstrate optimal convergence of order $h^{k+1}+\tau^p$.

math.NA

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research lines promote vanilla OPSD by choosing which tokens to learn from and by controlling how much privileged information the teacher receives, respectively. However, we show that each line optimizes one variable while holding the other fixed, which leads to a suboptimal solution. We argue that the two variables are coupled through the student's learning capacity: the privileged information sets the per-token divergence the teacher prescribes, while token weighting selects which of these the student must absorb. We formalize the two lines of work into a unified optimization framework, which maximizes the aggregate teacher--student divergence, subject to a budget on the aggregate learning difficulty the student can absorb. Under this modelling, we propose Unified On-Policy Self-Distillation (USD), a lightweight online algorithm to solve the Lagrangian. USD reveals that a single dual variable governs both decisions: at one price for learning difficulty, it simultaneously sets the token-selection threshold and the direction of privileged-information adjustment, keeping supervision matched to the student's evolving capacity. Through extensive experiments, USD consistently demonstrates superior performance over OPSD and token- and PI-side baselines across various model scales on various reasoning benchmarks. Code is available at https://github.com/lauvlalala/USD.

cs.AI

HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation

Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has emerged: the robot queries a multimodal LLM to understand its observations and plan the next action. However, long-horizon navigation based on either image streams or dense map inevitably introduces a growing memory and reasoning bottleneck. We present HAM-VLN, a decision-coupled, agent-authored memory that equips the robot with a persistent, depth-grounded world graph. In the same model call used to select the next action, HAM-VLN also records semantic and reflective information---including room type, objects, navigation progress, and failure notes. Recent waypoints remain verbatim within a bounded window, while older history re-enters the context only through retrieval scored by relevance, recency, and salience, together with one-hop topological expansion. This design requires no additional LLM calls beyond the per-waypoint decision. Compared to previous methods, HAM-VLN not only improves various navigation metrics but also reduces the context length by more than 65%. Specifically, HAM-VLN achieves 61.0% Success Rate (SR) on VLN-CE R2R, 52.7% SR on VLN-CE RxR, and 79.7% SR on HM3D-v2 ObjectNav without any training.

cs.RO

A Portable and Versatile Limited-Memory BFGS Implementation in PETSc/TAO

The limited-memory BFGS (L-BFGS) Hessian update scheme is the critical kernel in many quasi-Newton optimization algorithms. The most common approach to implementing L-BFGS uses $2m$ sequential rank-1 updates as part of solving a linear system when there are $m$ history steps. The performance of this approach suffers when the latency of synchronization is significant, and its poor temporal locality increases the memory traffic when vectors do not fit in cache. The compact dense representation of L-BFGS results in an approach that has minimal synchronization latency and better temporal locality, but it requires an additional pass over the basis vectors and an additional basis that must be recomputed when the $B_0$ matrix changes as in variable-metric methods. In the Portable Extensible Toolkit for Scientific Computation and the Toolkit for Advanced Optimization (PETSc/TAO), we have implemented an intermediate dense formulation of BFGS that retains most of the good characteristics of both the recursive and compact dense approaches. We report single-node performance tests of these implementations on the U.S. Department of Energy's Polaris and Frontier machines, testing both GPU-based and CPU-based computations.

cs.DC

AnchorMark: Robust Diffusion Watermarking via Latent-Space Rotation Synchrony

Inversion-based watermarking embeds watermark payloads directly into the generative process, avoiding a separate post-hoc image-domain embedding stage while preserving the native visual fidelity of synthesized images. However, existing methods remain vulnerable to compound lossy post-processing, particularly when rotation is involved, as it disrupts the spatial correspondence required for latent-space decoding. To overcome this limitation, we introduce AnchorMark, a training-free, robust inversion-based watermarking. We uncover a latent-space property termed Rotation Synchrony: image-domain rotations and their counterparts in the recovered initial latent share the same angle. Building on this property, AnchorMark embeds a synchronization anchor in the central region of the initial latent, enabling accurate estimation and correction of the rotation angle during extraction. Experiments show that AnchorMark substantially improves bit accuracy under rotation and combined attacks, with limited impact on image quality.

cs.CR

BlindPSNR: A No-Reference Fidelity Predictor for Low-Light Image Enhancement

Low-light image enhancement (LLIE) methods involve tunable parameters that are typically fixed, often leading to performance degradation when applied across scenes. Manually selecting the best configuration, however, can be time-consuming and not always practical. Peak signal-to-noise ratio (PSNR) is the natural fidelity criterion for automating parameter selection, yet it requires a ground-truth reference that is typically unavailable. To our knowledge, no learning-based method addresses no-reference PSNR prediction for low-light image enhancement; the natural surrogate, no-reference image quality assessment (NR-IQA), targets perceptual quality rather than signal fidelity, and all seven baselines we test achieve 0% top-1 selection accuracy on our benchmark. With paired training data, the ground-truth PSNR is analytically computable, providing exact supervision without a separate teacher network. Building on this, we propose BlindPSNR, a lightweight no-reference network that fuses the enhanced image with the degraded low-light input via windowed cross-attention and estimates PSNR through heteroscedastic regression. While a scalar-regression baseline achieves top-1 accuracy of 54.4%, BlindPSNR raises this to 89.5% with regret dropping from 1.62 dB to 0.026 dB, and generalizes to unseen datasets (SRCC = 0.61-0.67).

cs.CV

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera. Our model first selects the target from an indexed set of bounding boxes, then decodes tracking waypoints conditioned on this image-grounded decision. To preserve target motion cues over time, ReferTrack maintains a sliding-window queue of previously selected bounding boxes, injecting their geometric features into the visual history via temporal-viewpoint-bbox indicator (TVBI) tokens. We further enhance target identification by co-training on a custom Refer-QA dataset. On EVT-Bench, ReferTrack achieves state-of-the-art single-view performance with success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity tracking splits, respectively -- matching or even surpassing several multi-camera baselines on identification-heavy tasks. Finally, real-world deployments on legged and humanoid robots validate its robust sim-to-real transfer capabilities. Code is available at https://github.com/MedlarTea/referTrack.

cs.RO

Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation

Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models (LLMs) have shown unprecedented generalization capabilities in many research fields. Recently, projecting visual embeddings into the language space via vision-language models (VLMs) to achieve sim-toreal and cross-scene generalization has become a prevailing paradigm in the field of Vision-and-Language Navigation in Continuous Environments (VLN-CE). VLN requires an embodied agent to navigate through unseen environments following natural linguistic instructions. We emphasize that a VLN task can be decomposed into a sequence of sub-tasks, each corresponding to a process of 3D spatial interaction with the environments described by instructions such as "walk to the end of the sofa and turn left." However, such spatial interactions involving moving into the image along the direction of depth sensing are puzzling for VLMs as they were predominantly trained on conversations with RGB images. Rather than incorporating depth or 3D geometric information-which VLMs rarely encounter during pretrainingwe propose an alternative approach: fine-tuning VLMs to learn navigation interactions directly in 2D pixel space through autoregressive trajectory generation. Given a linguistic instruction and historical observations, our model sequentially predicts a series of pixel coordinates, drawing a trajectory from the bottom center of the current observation. While prior work has proved that pixel-goal supervision outperforms learning of discrete actions, our experiments further verify that the supervision of pixel-space trajectory significantly enhances VLN performance. Moreover, we demonstrate that our flagship model achieves state-of-the-art level performance with relatively limited computational resources and training data.

cs.CV

Quasi-bound states and late-time evolution of a massive fermion around a Reissner-Nordstr\"{o}m black hole

A massive fermion around a charged black hole provides a gravitational analogue of atomic bound states and their relaxation. In this work, we study this system by formulating the radial equation as a coupled matrix system and constructing the Green's function with ingoing boundary conditions at the horizon and decaying boundary conditions at infinity. In the weak-coupling scenario $|qQ|\sim mM<1$, a matrix matching scheme gives an improved analytic expression of quasi-bound-state spectrum, including fine-structure corrections and more accurate decay widths. The extremal Reissner-Nordstr\"{o}m case ($|Q|=M$) is treated separately and shown to be the smooth limiting result of the non-extremal spectrum. We further analyze the branch-cut contribution to the time-domain Green's function in the late-time limit. We confirm an oscillatory power-law behavior in intermediate late-time regime $1/m < t < 1/m^3M^2$. In the far late-time regime $t>1/m^3M^2$, the activation of the quasi-bound states produces an $t^{-5/6}\exp(-\eta t^{1/3})$ suppression with a chirping phase before the asymptotic $t^{-5/6}$ tail previously found in the limit $t\to\infty$. Direct time-domain simulations support this distinction and show how the quasi-bound contribution coexists with the familiar power-law component.

gr-qc

Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?

Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reasoning and real-to-sim digital twins. However, robot-mounted cameras naturally rotate during manipulation and navigation, while learned single-view reconstruction models often rely on view-dependent priors and may generalize poorly to out-of-distribution camera rotations. Such rotations can introduce 3D inconsistencies, incorrect layouts, and violations of physical constraints, but this failure mode remains under-evaluated. We introduce an evaluation protocol with controlled axis-wise roll, pitch, and yaw sweeps to trace errors in monocular depth estimation (MDE), canonical object meshes, camera-space layout, and physical plausibility within a representative SAM3D-style pipeline. On the Aria Digital Twin dataset and a real Franka wrist-camera sequence, camera rotations induce MDE distortion, layout drift, and collision penetration, while canonical mesh predictions remain relatively stable. A two-stage SAM3D+FoundationPose pipeline is more robust than one-stage feed-forward layout prediction, and our Gravity-Aware Refinement reduces one-stage pairwise ICP-based layout-orientation error by 47.1$\%$. Our evaluation reveals that current single-view mesh reconstruction methods generalize poorly to robot camera rotation, and suggests that explicit gravity cues are important for reliable robotic single-view mesh reconstruction.

cs.CV

Anisotropic 2D FUP and quantum open baker's map

We prove an essential spectral gap for 2D anisotropic quantum open baker's map. This extends the 1D results of Dyatlov--Jin 2017 and the isotropic 2D results of Cohen 2025a. The key ingredient is the anisotropic discrete fractal uncertainty principle (FUP) associated with a 2D anisotropic fractal set called the Bedford--McMullen carpet. We also study the relation between our anisotropic discrete FUP and its continuous counterpart in the spirit of Dyatlov--Jin 2018 and Cohen 2025a. In particular, we prove {continuous FUP} for 2D {anisotropic porous} sets, extending the (high-dimensional) isotropic results of Cohen 2025b. To the best of our knowledge, the anisotropic (line) porosity condition -- a variant of Cohen's line porosity and stronger than ball porosity -- appears to be new to the literature.

math.CA

FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction

Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training relies on voxel classification, which imposes only local constraints and lacks global supervision on the distribution of the primitives. Therefore, they inevitably predict spurious primitives in empty regions, undermining both representational and computational efficiency. To address this, we propose Feed-forward Likelihood Maximization (FLM), a novel framework that reformulates occupancy prediction as voxel distribution estimation. In FLM, a network is trained to predict a mixture model that maximizes the likelihood over ground-truth occupied voxels in a feed-forward manner. To enable end-to-end training of networks and voxelization of a standard mixture model, we define mixture weights as normalized primitive volumes to implicitly enforce simplex constraints and derive novel voxelization formulas. Based on FLM, our FLM-Occ, a novel method that is capable of relocating randomly initialized primitives over long distances to model a scene. On Occ-ScanNet, FLM-Occ achieves superior accuracy using only 32 superquadrics, 2.7% of the prior SoTA, while running 3.7 times faster.

cs.CV

GCNGrasp-VP: Affordance-Guided View Planning for Efficient Task-Oriented Grasping

Task-oriented grasping performance degrades significantly when object views suffer from occlusions. Existing task-oriented grasping methods typically assume task-relevant regions are visible in the initial frame, while view planning approaches enable active perception but often ignore task semantics and rely on time-consuming scene reconstruction. To address these limitations, we present GCNGrasp-VP, an efficient framework integrating affordance field prediction with active view planning. Central to this framework is GCNGrasp-v2, a task-oriented grasp model that simultaneously supports grasp evaluation and affordance field prediction, achieving constant-time inference complexity. Leveraging this capability, our Affordance-guided View Planner (Affordance-VP) utilizes the affordance field as an information gain metric to guide camera observation of task-relevant regions without requiring scene reconstruction. View planning results show that our method significantly outperforms scene-uncertainty-driven baselines with only one view adjustment. Real-world validation further confirms substantial improvements in grasp success rates for single-object scenarios while maintaining millisecond-level computational latency. Code and models are available at https://github.com/Instinct323/GCNGrasp-VP.

cs.RO