SearcharxivSearch

arXiv subjects

Jingdong Zhang

Publications and source records attributed to Jingdong Zhang.

At least 19 recordsLinked to original sources

Event-Triggered Stabilisation of Desynchronisation in Networked Oscillatory Systems

Pathological neuronal synchrony provides one practical motivation for studying sparse desynchronisation control in networked oscillatory systems, particularly in applications where actuation and communication are resource constrained, as in deep brain stimulation. Motivated by this challenge, we study how to stabilise desynchronisation in coupled oscillatory dynamical systems without continuously updated control, even when the oscillators' phase is unavailable. We develop a general event-triggered control framework for stabilising the desynchronised state of coupled limit-cycle oscillatory dynamical systems. Our analysis establishes a unified theoretical result showing that desynchronisation can be achieved by various feedback controllers that act sparsely and depend only on an order parameter observable. The proposed controllers admit a gradient-descent interpretation and stabilise the desynchronisation state of general phase-reduced oscillator networks. We further prove the controlled systems under event-triggered mechanisms possess a strictly positive lower bound of the inter-event dwell-time, excluding Zeno behaviour. To address the practical unavailability of exact phase reductions, we introduce a pseudo-phase construction that yields a computable order parameter from state measurements alone. Numerical studies on representative oscillator networks demonstrate the effectiveness and robustness of the proposed framework.

math.DS

CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment

Multimodal learning has significantly advanced survival prediction by integrating pathology images with genomic data. However, clinical information, despite its critical role in reflecting a patient' s overall health, remains underutilized due to its discrete, sparse, and low-dimensional nature. Furthermore, the inherent heterogeneity across these modalities pose significant challenges in modeling cross-modal interactions. In this paper, we propose CIGTSurv, a Clinical Information Guided Tri-modal framework for Survival prediction. Specifically, we first design a holistic text template and use pretrained foundation models to transform clinical tabular data into high-dimensional tokenized embeddings. Using clinical information as an anchor, we then introduce a dual-level interaction mechanism: 1) a local prototype association (LPA) module based on cross-attention to explicitly learn token-level correspondences between different modalities, and 2) a global feature alignment (GFA) loss based on Maximum Mean Discrepancy (MMD) to implicitly enhance cross-modal distribution consistency. Extensive experiments on five TCGA cancer cohorts demonstrate that CIGTSurv achieves state-of-the-art (SOTA) survival prediction performance. Our source code is publicly available at https://github.com/Daijing-ai/CIGT-Surv.git.

cs.CV

Ptolemy's Equant Equates to a Universal Dynamical Clock via Machine Learning

Oscillatory dynamics arise ubiquitously in nonlinear systems, yet identifying a physically interpretable phase and phase dynamics in nonlinear, high-dimensional oscillations remains a central unresolved problem. Here we establish the principle of a universal dynamical clock, a physical perspective in which oscillations of arbitrary dimensionality and geometry are equivalently represented as uniform rotation through an equant-induced nonlinear viewing coordinate, inspired by Ptolemy's equant and formalised through an areal-uniformity principle reminiscent of Kepler's second law. Using a machine-learning framework, we demonstrate the existence of such an equant for a broad class of oscillatory dynamics and construct the associated dynamical clock and phase dynamics under additive forces, including noise, periodic perturbations, and coupling. Its value in uncovering new physical rules and phenomena is demonstrated by four findings: (i) collective oscillations in Escherichia coli populations obey a previously unexplained superlinear scaling law, resolving a long-standing open problem posed in 2004; (ii) the response mechanisms of engineered genetic circuits to changes in gene expression and environmental conditions; (iii) a classical-mechanics counterpart of the Berry geometric phase emerges naturally from the phase of the dynamical clock; and (iv) optimal equant non-uniformity provides a geometric early-warning signal for critical transitions and enables prediction of critical parameters. By providing operational and system-agnostic phase dynamics that can be constructed directly from data, the dynamical clock enables principled classification, comparison, and control of oscillatory systems, and offers a new route to understanding how specific dynamical regimes support distinct functional behaviours in networked systems.

math.DS

Lyapunov Guidance: A Unified Framework for Stabilizing Generative Flows

Flow matching has emerged as an effective framework for learning complex data distributions, but adapting pretrained flow models to new tasks often requires computationally expensive retraining. Post-training guidance provides a more efficient alternative, but existing methods are largely heuristic and offer no explicit stability guarantees. We address this limitation by proposing LyaGuide, a unified Lyapunov-guided framework that formulates flow guidance as a Lyapunov control problem. Our main theoretical result establishes an equivalence between guided flow matching and Lyapunov control, thereby unifying common guidance strategies, such as classifier guidance, reward guidance, and energy-based guidance, within a single control-theoretic framework. To enforce the Lyapunov condition, we introduce a pseudo-projection operator with a closed-form expression that endows learned or heuristic guidance terms with explicit stability guarantees. LyaGuide supports two practical settings: a model-driven setting, where the target guidance distribution is specified through a known Lyapunov function, and a data-driven setting, where the guidance is adapted from task-specific downstream data. LyaGuide is compatible with existing guidance methods, introduces minimal additional computational overhead, and is straightforward to integrate in practice. Extensive experiments on synthetic benchmarks, image inverse problems, reinforcement learning planning, and energy-based modeling demonstrate consistent improvements in sample quality, guidance fidelity, and robustness, while maintaining computational efficiency.

cs.LG

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search

Humanoid Visual Search (HVS) requires agents to actively explore immersive 360$^\circ$ environments. While prior methods treat this as a monolithic task relying on cumulative, multi-turn Chain-of-Thought (CoT) reasoning, they impose heavy cognitive burdens and require expensive trajectory-level annotations. In this paper, we propose Imagining in 360$^\circ$, a novel framework that decouples the exploration process into a specialized Imaginator and an Actor. The Imaginator functions as a probabilistic predictor of spatial priors; instead of maintaining a cumulative reasoning chain, it infers the semantic layout of both observed and unobserved regions in a single step. By sampling multiple hypotheses within this semantic space, we provide the Actor with a distribution of effective spatial information, offering robust guidance that hedges against uncertainty during active search. This decoupled architecture significantly lowers data engineering costs by eliminating the need for full-trajectory CoT annotations, enabling the generation of over 1.96 million curated training samples. Extensive experiments demonstrate that explicitly modeling semantic spatial priors drastically improves search efficiency and success rates in complex, in-the-wild environments.

cs.CV

UniSER: A Foundation Model for Unified Soft Effects Removal

Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially visible. The prevailing works address these degradations in isolation, developing highly specialized, specialist models that lack scalability and fail to exploit the shared underlying essences of these restoration problems. Meanwhile, although recent large-scale generalist models (e.g., GPT-4o, Flux Kontext, Nano Banana) offer powerful text-driven editing capabilities, they heavily rely on detailed prompts and often fail to achieve robust removal on such fine-grained tasks while preserving the scene's identity. Leveraging the common essence of soft effects, i.e., semi-transparent occlusions, we introduce a foundational versatile model UniSER, capable of addressing diverse degradations caused by soft effects within a single framework. Our methodology centers on curating a massive 3.8M-pair dataset to ensure robustness and generalization, which includes novel, physically-plausible data to fill critical gaps in public benchmarks, and a tailored training pipeline that fine-tunes a Diffusion Transformer to learn robust restoration priors from this diverse data, integrating fine-grained mask and strength controls. This synergistic approach allows UniSER to significantly outperform both specialist and generalist models, achieving robust, high-fidelity restoration in the wild.

cs.CV

MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors

Comprehensive panoramic scene understanding is critical for immersive applications, yet it remains challenging due to the scarcity of high-resolution, multi-task annotations. While perspective foundation models have achieved success through data scaling, directly adapting them to the panoramic domain often fails due to severe geometric distortions and coordinate system discrepancies. Furthermore, the underlying relations between diverse dense prediction tasks in spherical spaces are underexplored. To address these challenges, we propose MTPano, a robust multi-task panoramic foundation model established by a label-free training pipeline. First, to circumvent data scarcity, we leverage powerful perspective dense priors. We project panoramic images into perspective patches to generate accurate, domain-gap-free pseudo-labels using off-the-shelf foundation models, which are then re-projected to serve as patch-wise supervision. Second, to tackle the interference between task types, we categorize tasks into rotation-invariant (e.g., depth, segmentation) and rotation-variant (e.g., surface normals) groups. We introduce the Panoramic Dual BridgeNet, which disentangles these feature streams via geometry-aware modulation layers that inject absolute position and ray direction priors. To handle the distortion from equirectangular projections (ERP), we incorporate ERP token mixers followed by a dual-branch BridgeNet for interactions with gradient truncation, facilitating beneficial cross-task information sharing while blocking conflicting gradients from incompatible task attributes. Additionally, we introduce auxiliary tasks to fertilize the cross-task learning process. Extensive experiments demonstrate that MTPano achieves state-of-the-art performance on multiple benchmarks and delivers competitive results against task-specific panoramic specialist foundation models.

cs.CV

Evidence of Long-Lived Powerful Gyrosynchrotron Radio Emission in the Close Binary FF UMa

RS Canum Venaticorum (RS CVn) close binaries, characterized by tidal locking, rapid rotations, and strong magnetic fields, are ideal laboratories for high-resolution radio observations to probe emission processes, magnetic field configurations, and interaction activity. Despite their importance, only a few RS CVn sources have been explored by polarimetric observations of very long baseline interferometry (VLBI). To expand the effort, we have analyzed the existing Very Long Baseline Array (VLBA) astrometric data for the RS CVn binary FF Ursae Majoris (FF UMa). In the 5GHz VLBA experiments conducted between 2021 and 2024, both total intensity and circularly polarized emission were clearly detected at six of seven epochs. The consistently high brightness temperatures (10^7 K) and the moderate fractional circular polarization (10%-30%) over about three years indicate that the radio emission is mainly produced by gyrosynchrotron radiation from mildly relativistic electrons in the highly-ordered magnetic field. The radio luminosities are also comparable to those of previously studied powerful RS CVn binaries and show a significant anti-correlation with fractional circular polarization. A mean centroid offset of 13.4 +/- 3.1 solar radii between the Stokes I and V emission was found across multiple epochs, indicating a possible additional contribution from the secondary star via a magnetically active corona, a giant magnetic loop, or significant interaction activity with the primary star in the quiescent state.

astro-ph.SR

HGP-Mamba: Integrating Histology and Generated Protein Features for Mamba-based Multimodal Survival Risk Prediction

Recent advances in multimodal learning have significantly improved cancer survival risk prediction. However, the joint prognostic potential of protein markers and histopathology images remains underexplored, largely due to the high cost and limited availability of protein expression profiling. To address this challenge, we propose HGP-Mamba, a Mamba-based multimodal framework that efficiently integrates histological with generated protein features for survival risk prediction. Specifically, we introduce a protein feature extractor (PFE) that leverages pretrained foundation models to derive high-throughput protein embeddings directly from Whole Slide Images (WSIs), enabling data-efficient incorporation of molecular information. Together with histology embeddings that capture morphological patterns, we further introduce the Local Interaction-aware Mamba (LiAM) for fine-grained feature interaction and the Global Interaction-enhanced Mamba (GiEM) to promote holistic modality fusion at the slide level, thus capture complex cross-modal dependencies. Experiments on four public cancer datasets demonstrate that HGP-Mamba achieves state-of-the-art performance while maintaining superior computational efficiency compared with existing methods. Our source code is publicly available at https://github.com/Daijing-ai/HGP-Mamba.git.

cs.CV

VLBI astrometry of radio stars to link radio and optical celestial reference frames - II. 11 radio stars

The alignment between the radio-based International Celestial Reference Frame (ICRF) and the optical Gaia Celestial Reference Frame (Gaia-CRF) is critical for multi-waveband astronomy, yet systematic offsets at the optical bright end (G<13) limit their consistency. While radio stars offer a potential link between these frames, their utility has been restricted by the scarcity of precise Very Long Baseline Interferometry (VLBI) astrometry. In this study, we present new VLBI astrometry of 11 radio stars using the Very Long Baseline Array (VLBA), expanding the existing sample with positions, parallaxes, and proper motions measured. All 11 radio stars were detected, for 10 of which parallaxes and proper motions can be estimated, achieving median uncertainties better than 0.1 mas and 0.1 mas/yr, respectively. These new samples greatly contribute to the link between ICRF and Gaia-CRF at the optical bright end.

astro-ph.SR

SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation

Existing single-view 3D generative models typically adopt multiview diffusion priors to reconstruct object surfaces, yet they remain prone to inter-view inconsistencies and are unable to faithfully represent complex internal structure or nontrivial topologies. In particular, we encode geometry information by projecting it onto a bounding sphere and unwrapping it into a compact and structural multi-layer 2D Spherical Projection (SP) representation. Operating solely in the image domain, SPGen offers three key advantages simultaneously: (1) Consistency. The injective SP mapping encodes surface geometry with a single viewpoint which naturally eliminates view inconsistency and ambiguity; (2) Flexibility. Multi-layer SP maps represent nested internal structures and support direct lifting to watertight or open 3D surfaces; (3) Efficiency. The image-domain formulation allows the direct inheritance of powerful 2D diffusion priors and enables efficient finetuning with limited computational resources. Extensive experiments demonstrate that SPGen significantly outperforms existing baselines in geometric quality and computational efficiency.

cs.CV

Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense Predictions

In recent years, simultaneous learning of multiple dense prediction tasks with partially annotated label data has emerged as an important research area. Previous works primarily focus on leveraging cross-task relations or conducting adversarial training for extra regularization, which achieve promising performance improvements, while still suffering from the lack of direct pixel-wise supervision and extra training of heavy mapping networks. To effectively tackle this challenge, we propose a novel approach to optimize a set of compact learnable hierarchical task tokens, including global and fine-grained ones, to discover consistent pixel-wise supervision signals in both feature and prediction levels. Specifically, the global task tokens are designed for effective cross-task feature interactions in a global context. Then, a group of fine-grained task-specific spatial tokens for each task is learned from the corresponding global task tokens. It is embedded to have dense interactions with each task-specific feature map. The learned global and local fine-grained task tokens are further used to discover pseudo task-specific dense labels at different levels of granularity, and they can be utilized to directly supervise the learning of the multi-task dense prediction framework. Extensive experimental results on challenging NYUD-v2, Cityscapes, and PASCAL Context datasets demonstrate significant improvements over existing state-of-the-art methods for partially annotated multi-task dense prediction.

cs.CV

Neural Event-Triggered Control with Optimal Scheduling

Learning-enabled controllers with stability certificate functions have demonstrated impressive empirical performance in addressing control problems in recent years. Nevertheless, directly deploying the neural controllers onto actual digital platforms requires impractically excessive communication resources due to a continuously updating demand from the closed-loop feedback controller. We introduce a framework aimed at learning the event-triggered controller (ETC) with optimal scheduling, i.e., minimal triggering times, to address this challenge in resource-constrained scenarios. Our proposed framework, denoted by Neural ETC, includes two practical algorithms: the path integral algorithm based on directly simulating the event-triggered dynamics, and the Monte Carlo algorithm derived from new theoretical results regarding lower bound of inter-event time. Furthermore, we propose a projection operation with an analytical expression that ensures theoretical stability and schedule optimality for Neural ETC. Compared to the conventional neural controllers, our empirical results show that the Neural ETC significantly reduces the required communication resources while enhancing the control performance in constrained communication resources scenarios.

math.OC

Validating the bright Gaia celestial reference frame with new VLBI astrometry of radio stars

There exist inconsistencies between the bright and faint Gaia Celestial Reference Frame 3 (Gaia-CRF3), which manifests as a systematic rotation and needs to be independently estimated then corrected in future data releases. We collected 64 radio stars with VLBI astrometry, of which 16 have new VLBI observations with reference epochs close to Gaia. We estimated the orientation and spin biases of the bright Gaia-CRF3 by comparing VLBI radio star astrometry with their Gaia DR3 counterparts. We also attempted to estimate orientation by utilizing the a priori magnitude-dependent spin parameters derived from Gaia internal estimation. Our independent estimation of the orientation at G<10.5 is [-15\pm119, +330\pm139,+218\pm109] uas (J2016.0), and the spin ([+21\pm18, +52\pm20,-7\pm20] uas/yr) agrees with Gaia internal estimation within 1-sigma range. The orientation-only estimation suggests that the orientation bias of the bright Gaia-CRF3 may also be magnitude-dependent.

astro-ph.SR

Serial MultiView: an efficient approach to mitigating atmospheric spatial-structure errors for VLBI astrometry

Atmospheric propagation errors are a main constraint on the accuracy of Very Long Baseline Interferometry (VLBI) astrometry. For relative astrometry, differential techniques can mitigate these errors, but their effectiveness diminishes with decreasing elevation and increasing angular separations between target and calibrator, among others. The MultiView technique addresses atmospheric spatial-structure errors by observing multiple calibrators around the target and interpolating at the target position, thereby reducing atmospheric errors more effectively than phase-referencing with only one calibrator. The first MultiView realisation at 1.6GHz involved cyclically observing all calibrators and the target, fitting a phase plane from calibrator solutions in each cycle, and is a well-established technique. This implementation reduces on-target time and is constricted by the short atmospheric coherence time at high frequencies. We propose a new realisation, serial MultiView, which rotates the phase plane iteratively based on the time series of calibrator residual phases. This new strategy obviates the necessity of observing all calibrators within each cycle, thereby shortening the observing cycle and offering considerable potential at higher frequencies where the temporal structure is the dominant source of errors. Additionally, by incorporating time-domain information in the iterations, phase ambiguities can be accurately and automatically identified. We verify the astrometric accuracy of serial MultiView at 5GHz by comparing it to conventional MultiView, achieving <10uas error in RA direction, and show the calibration overhead can be reduced in both approaches. This approach enables efficient, high-accuracy differential astrometry and artifact-reduced imaging for astrophysical studies, and we provide a user-friendly tool for it.

astro-ph.IM

3R-GS: Best Practice in Optimizing Camera Poses Along with 3DGS

3D Gaussian Splatting (3DGS) has revolutionized neural rendering with its efficiency and quality, but like many novel view synthesis methods, it heavily depends on accurate camera poses from Structure-from-Motion (SfM) systems. Although recent SfM pipelines have made impressive progress, questions remain about how to further improve both their robust performance in challenging conditions (e.g., textureless scenes) and the precision of camera parameter estimation simultaneously. We present 3R-GS, a 3D Gaussian Splatting framework that bridges this gap by jointly optimizing 3D Gaussians and camera parameters from large reconstruction priors MASt3R-SfM. We note that naively performing joint 3D Gaussian and camera optimization faces two challenges: the sensitivity to the quality of SfM initialization, and its limited capacity for global optimization, leading to suboptimal reconstruction results. Our 3R-GS, overcomes these issues by incorporating optimized practices, enabling robust scene reconstruction even with imperfect camera registration. Extensive experiments demonstrate that 3R-GS delivers high-quality novel view synthesis and precise camera pose estimation while remaining computationally efficient. Project page: https://zsh523.github.io/3R-GS/

cs.CV

The orbital period of the long-period and colliding-wind binary WR 146 from radio interferometry of the shock cone

We report the first measurement of the orbital period of a long-period colliding-wind binary (CWB) system WR 146, derived by tracing the rotational morphology of its wind-colliding region (WCR) and the relative orientation of the two binary components. This result is based on our imaging observations using the Very Long Baseline Array (VLBA) and the European Very Long Baseline Interferometry (VLBI) Network (EVN), combined with archival data from VLBA, EVN, the Very Large Array (VLA), the enhanced Multi-Element Radio-Linked Interferometer Network (eMERLIN) arrays, and optical images from the Hubble Space Telescope (HST). We evaluated two methods for determining the binary's orbital period based on the images of the WCR: (I) fitting the shock cone of the WCR and (II) stacking images using the cross-correlation function. Using these techniques, we find orbital period estimates of 810+120-90 years from method I and 1120+540-270 years from method II, both of which support a long orbital period of approximately 1,000 years. Furthermore, we analyzed archival spectral data of WR 146 to estimate the stellar wind velocities of the binary components, finding no significant orbital phase lag between the binary orientation and the WCR rotation. We also estimate the range of the binary's mass using the currently measured parameters.

astro-ph.SR

SolidGS: Consolidating Gaussian Surfel Splatting for Sparse-View Surface Reconstruction

Gaussian splatting has achieved impressive improvements for both novel-view synthesis and surface reconstruction from multi-view images. However, current methods still struggle to reconstruct high-quality surfaces from only sparse view input images using Gaussian splatting. In this paper, we propose a novel method called SolidGS to address this problem. We observed that the reconstructed geometry can be severely inconsistent across multi-views, due to the property of Gaussian function in geometry rendering. This motivates us to consolidate all Gaussians by adopting a more solid kernel function, which effectively improves the surface reconstruction quality. With the additional help of geometrical regularization and monocular normal estimation, our method achieves superior performance on the sparse view surface reconstruction than all the Gaussian splatting methods and neural field methods on the widely used DTU, Tanks-and-Temples, and LLFF datasets.

cs.CV