SearcharxivSearch

arXiv subjects

Jiao Xu

Publications and source records attributed to Jiao Xu.

13 recordsLinked to original sources

Learning to Track from Privileged Target Appearances

Target templates define what a visual tracker searches for, yet the templates available at inference trade off localization certainty with appearance freshness: the initial ground-truth template is exact but becomes stale, whereas recent templates better reflect the current appearance but are cropped from uncertain predictions. We quantify this bottleneck with a non-deployable oracle that supplies an exact current-frame target crop, improving AUC on LaSOT by 15.2 percentage points. This gap reveals a training-only opportunity: frame-level ground truths provide exact current- and future-frame target crops, although such crops are unavailable at deployment. We introduce Privileged Appearance Transfer for Tracking (PATT), a teacher-student training framework that transfers these privileged appearances to a deployable tracker through multi-level representation prediction. The privileged teacher observes exact target crops from past, current, and future frames, whereas the student receives only past-frame templates and learns to predict the teacher's search representations. To avoid transferring unreliable teacher signals, PATT weights this transfer by the teacher's relative localization advantage over the student and its absolute localization accuracy. After training, the teacher, latent predictor, reliability weights, and privileged crops are removed, leaving standard student-only inference. Across seven benchmarks at two model scales, PATT achieves consistent gains under both long- and short-term tracking protocols.

cs.CV

RELO: Reinforcement Learning to Localize for Visual Object Tracking

Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly aligned with tracking optimization and evaluation metrics, such as intersection over union (IoU) and area under the success curve (AUC). Here, we introduce RELO, a REinforcement-learning-to-LOcalize method for visual object tracking that formulates target localization as a Markov decision process. Specifically, RELO replaces handcrafted spatial priors with a localization policy learned over spatial positions via reinforcement learning, with rewards combining frame-level IoU and sequence-level AUC. We additionally introduce layer-aligned temporal token propagation to improve semantic consistency across frames, with negligible computational overhead. Across multiple benchmarks, RELO achieves superior results, attaining 57.5% AUC on LaSOText without template updates. This confirms that reward-driven localization provides an effective alternative to prior-driven localization for visual object tracking.

cs.CV

PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To bridge this gap, we introduce PulseMind, a new family of multi-modal diagnostic models that integrates a systematically curated dataset, a comprehensive evaluation benchmark, and a tailored training framework. Specifically, we first construct a diagnostic dataset, MediScope, which comprises 98,000 real-world multi-turn consultations and 601,500 medical images, spanning over 10 major clinical departments and more than 200 sub-specialties. Then, to better reflect the requirements of real-world clinical diagnosis, we develop the PulseMind Benchmark, a multi-turn diagnostic consultation benchmark with a four-dimensional evaluation protocol comprising proactiveness, accuracy, usefulness, and language quality. Finally, we design a training framework tailored for multi-modal clinical diagnostics, centered around a core component named Comparison-based Reinforcement Policy Optimization (CRPO). Compared to absolute score rewards, CRPO uses relative preference signals from multi-dimensional com-parisons to provide stable and human-aligned training guidance. Extensive experiments demonstrate that PulseMind achieves competitive performance on both the diagnostic consultation benchmark and public medical benchmarks.

cs.CV

Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentation

In this paper, we present a new dynamic collaborative network for semi-supervised 3D vessel segmentation, termed DiCo. Conventional mean teacher (MT) methods typically employ a static approach, where the roles of the teacher and student models are fixed. However, due to the complexity of 3D vessel data, the teacher model may not always outperform the student model, leading to cognitive biases that can limit performance. To address this issue, we propose a dynamic collaborative network that allows the two models to dynamically switch their teacher-student roles. Additionally, we introduce a multi-view integration module to capture various perspectives of the inputs, mirroring the way doctors conduct medical analysis. We also incorporate adversarial supervision to constrain the shape of the segmented vessels in unlabeled data. In this process, the 3D volume is projected into 2D views to mitigate the impact of label inconsistencies. Experiments demonstrate that our DiCo method sets new state-of-the-art performance on three 3D vessel segmentation benchmarks. The code repository address is https://github.com/xujiaommcome/DiCo

cs.CV

Robust Sparse Signal Recovery with Outliers: A Hard Thresholding Pursuit Approach Based on LAD

Recovering a sparse signal from outlier-contaminated measurements is a fundamental challenge in many applications. While existing algorithms predominantly address scenarios with bounded noise or assume known signal sparsity, few methods tackle the more practical problem of sparse recovery from gross outliers without prior knowledge of sparsity. To bridge this gap, we study the sparsity-constrained Least Absolute Deviations (LAD) minimization problem. This paper proposes the Graded Fast Hard Thresholding Pursuit (GFHTP$_1$) algorithm with a quantile-truncated step size for $\ell_1$-loss minimization. In contrast to most state-of-the-art methods, our GFHTP$_1$ requires no prior knowledge of the signal's sparsity level. We establish a theoretical convergence analysis under mild conditions and further prove that an $s$-sparse signal can be recovered exactly within at most $s$ iterations. To our knowledge, these results provide the first efficient recovery guarantees for sparse signal reconstruction from outlier-corrupted measurements without a sparsity prior. Numerical experiments demonstrate that GFHTP$_1$ consistently outperforms competing algorithms in robustness to varying signal sparsity and outlier support size, while also achieving less computational time.

cs.IT

On the global well-posedness and scattering of the 3D Klein-Gordon-Zakharov system

In this paper we are interested in the global well-posedness of the 3D Klein-Gordon-Zakharov equations with small initial data. We show the uniform boundedness of the energy for the global solution without any compactness assumptions on the initial data. The main novelty of our proof is to apply a modified Alinhac's ghost weight method together with a newly developed normal-form type estimate to remedy the lack of the space-time scaling vector field; moreover, we give a clear description of the smallness conditions on the initial data.

math.AP

Tunable optical bistability in grapheme Tamm plasmon/Bragg reflector hybrid structure at terahertz frequencies

We propose a composite multilayer structure consist of graphene Tamm plasmon and Bragg reflector with defect layer to realize the low threshold and tunable optical bistability (OB) at the terahertz frequencies. This low-threshold OB originates from the couple of the Tamm plasmon (TP) and the defect mode (DM). We discuss the influence of graphene and the DM on the hysteretic response of the TM-polarized reflected light. It is found that the switch-up and switch-down threshold required to observe the optical bistable behavior are lowered markedly due to the excitation of the TP and DM. Besides, the switching threshold value can be further reduced by coupling the TP and DM. We believe these results will provide a new avenue for realizing the low threshold and tunable optical bistable devices and other nonlinear optical devices.

physics.optics

Stability and convergence of Strang splitting. Part II: tensorial Allen-Cahn equations

We consider the second-order in time Strang-splitting approximation for vector-valued and matrix-valued Allen-Cahn equations. Both the linear propagator and the nonlinear propagator are computed explicitly. For the vector-valued case, we prove the maximum principle and unconditional energy dissipation for a judiciously modified energy functional. The modified energy functional is close to the classical energy up to $\mathcal O(τ)$ where $τ$ is the splitting step. For the matrix-valued case, we prove a sharp maximum principle in the matrix Frobenius norm. We show modified energy dissipation under very mild splitting step constraints. We exhibit several numerical examples to show the efficiency of the method as well as the sharpness of the results.

math.NA

Stability and convergence of Strang splitting. Part I: Scalar Allen-Cahn equation

We consider a class of second-order Strang splitting methods for Allen-Cahn equations with polynomial or logarithmic nonlinearities. For the polynomial case both the linear and the nonlinear propagators are computed explicitly. We show that this type of Strang splitting scheme is unconditionally stable regardless of the time step. Moreover we establish strict energy dissipation for a judiciously modified energy which coincides with the classical energy up to $\mathcal O(τ)$ where $τ$ is the time step. For the logarithmic potential case, since the continuous-time nonlinear propagator no longer enjoys explicit analytic treatments, we employ a second order in time two-stage implicit Runge--Kutta (RK) nonlinear propagator together with an efficient Newton iterative solver. We prove a maximum principle which ensures phase separation and establish energy dissipation law under mild restrictions on the time step. These appear to be the first rigorous results on the energy dissipation of Strang-type splitting methods for Allen-Cahn equations.

math.NA

Recall and Learn: A Memory-augmented Solver for Math Word Problems

In this article, we tackle the math word problem, namely, automatically answering a mathematical problem according to its textual description. Although recent methods have demonstrated their promising results, most of these methods are based on template-based generation scheme which results in limited generalization capability. To this end, we propose a novel human-like analogical learning method in a recall and learn manner. Our proposed framework is composed of modules of memory, representation, analogy, and reasoning, which are designed to make a new exercise by referring to the exercises learned in the past. Specifically, given a math word problem, the model first retrieves similar questions by a memory module and then encodes the unsolved problem and each retrieved question using a representation module. Moreover, to solve the problem in a way of analogy, an analogy module and a reasoning module with a copy mechanism are proposed to model the interrelationship between the problem and each retrieved question. Extensive experiments on two well-known datasets show the superiority of our proposed algorithm as compared to other state-of-the-art competitors from both overall performance comparison and micro-scope studies.

cs.CL

Energy-dissipation for time-fractional phase-field equations

We consider a class of time-fractional phase field models including the Allen-Cahn and Cahn-Hilliard equations. We establish several weighted positivity results for functionals driven by the Caputo time-fractional derivative. Several novel criterions are examined for showing the positive-definiteness of the associated kernel functions. We deduce strict energy-dissipation for a number of non-local energy functionals, thereby proving fractional energy dissipation laws.

math.AP

Uniform boundedness of highest norm for 2D quasilinear wave

We consider the two-dimensional quasilinear wave equations with quadratic nonlinearities. We introduce a new class of null forms and prove uniform boundedness of the highest order norm of the solution for all time. This class of null forms include several prototypical strong null conditions as special cases. To handle the critical decay near the light cone we inflate the nonlinearity through a new normal form type transformation which is based on a deep cancelation between the tangential and normal derivatives with respect to the light cone. Our proof does not employ the Lorentz boost and can be generalized to systems with multiple speeds.

math.AP