SearcharxivSearch

arXiv subjects

Xia Li

Publications and source records attributed to Xia Li.

At least 19 recordsLinked to original sources

Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search

Optimization-based jailbreak attacks such as Greedy Coordinate Gradient (GCG) achieve strong effectiveness and transferability by optimizing adversarial suffixes on white-box source models. However, existing GCG-based methods rely on averaged adversarial loss and deep greedy search, which can over-emphasize easy-to-jailbreak behaviors and overlook promising regions of the suffix space. We propose BOSS, a plug-and-play framework that improves GCG-based jailbreak optimization through breadth-oriented suffix search. BOSS uses Tail-Focused Adversarial Loss (TFAL), standard source loss, and behavior coverage to select terminal suffixes, then explores multiple short trajectories and selectively continues promising suffixes. Experiments on public benchmarks show that BOSS improves attack success rates across multiple GCG-based methods while reducing optimization time.

cs.CL

Portable surrogate-free 4D MRI from standard fast multi-slice 2D MRI via implicit neural representations

Four-dimensional MRI (4D MRI) characterizes respiratory organ motion, yet existing reconstruction pipelines are tightly coupled to specific acquisition platforms (e.g., non-Cartesian trajectories with self-gating, vendor-specific navigators, or external respiratory hardware), limiting broad adoption across diverse clinical and research settings, including low-field, open-bore, and non-supine imaging. We present SIMPLE-4D (Surrogate-free, IMplicit, PortabLE 4D MRI), a software-first portable workflow that operates entirely on reconstructed slices from standard fast multi-slice 2D MRI and requires no pulse-sequence modification, no non-Cartesian trajectory, no navigator, and no external hardware. SIMPLE-4D combines an acquisition-agnostic front end consuming standard 2D protocols, a surrogate-free variational motion encoder that extracts a compact motion code directly from each 2D slice, and a physics-aware continuous spatio-temporal reconstruction based on a hash-encoded implicit neural representation (INR) with a SIREN deformation network producing bidirectional cycle-consistent DVFs and motion-dependent Gauss-Legendre thick-slice quadrature. Bidirectionality yields a complete inter-frame motion model by composition, supporting downstream tasks such as dose accumulation without retraining. We validate the identical pipeline on two contrasting datasets: a 1.5 T clinical bSSFP dataset (5 volunteers, 3 sessions each) and a 0.5 T open-bore HASTE dataset (5 volunteers, supine and upright). To our knowledge, this is the first per-frame 4D volumetric respiratory MRI reconstruction on a weight-bearing upright open-bore low-field scanner from reconstructed 2D Cartesian slices alone. On low-field data the INR template additionally acts as an implicit denoiser, yielding +132% SNR. Systematic ablations isolate each component's contribution.

physics.med-ph

Causal-Privacy Audit Workflow for Synthetic and Distilled Data in Dropout Support

Synthetic and distilled student data are increasingly used to enable privacy-conscious learning analytics, yet their suitability for decision-facing institutional support remains uncertain. In dropout support, generated data must preserve not only predictive utility or distributional resemblance, but also the financial-status evidence used to guide advising, payment-plan assistance, and scholarship-related decisions. Method: This study introduces CaP-Eval, a decision-facing causal-privacy audit workflow for evaluating generated student data under a fixed estimand, timing-aware adjustment design, estimator set, and empirical privacy-governance screen. The workflow compares original, distilled, adversarial synthetic, statistical synthetic, and DPGNet privacy-oriented generated data on predictive utility, treatment-effect fidelity, robustness to alternative estimators, and local training-record proximity. Results: DPGNet and distilled data preserved the original financial-status treatment-effect structure more reliably than the adversarial and Gaussian Copula baselines. DPGNet preserved full direction and rank agreement across epsilon levels; epsilon = 10 produced the smallest non-original IPW and DML deviations, while epsilon = 1 and epsilon = 5 amplified several financial-status contrasts. Distilled data remained highly faithful but retained the strongest local training-record proximity signal. TabularGNet preserved qualitative directions with moderate attenuation, and Gaussian Copula compressed effect magnitudes. Conclusions: Predictive utility, privacy orientation, empirical disclosure signals, and causal fidelity diverged; generated student data require joint audits of direction, magnitude, overlap, and release-governance risk before decision use.

cs.LG

Edge-directed geometric partitioning for versatile video coding

To improve the coding performance, geometric partition (GEO) was proposed for the upcoming VVC standard. GEO provides 140 partition candidates. The index of optimal GEO mode needs to be signaled explicitly. Considering different structural characteristics of different CUs and the correlation between spatial adjacent blocks and temporal collocated blocks, we propose a GEO mode prediction strategy by constructing a Most Probable Mode (MPM) list to reduce the overhead of GEO index and improve coding efficiency. Based on the observation of the high correlation between the partition mode and object boundaries, an edge-directed geometric partition scheme is proposed to construct the MPM list according to spatio-temporal edge information. The proposed method provides an objective BD-rate gain of 0.58% and 1.00% on average for RA and LDB configurations compared to VTM-6.0. Besides, it also promotes the visual quality of object boundaries.

cs.CV

Forget Less, Generalize More: Unifying Temporal and Structural Adaptation for Dynamic Graphs

Representation learning on dynamic graphs requires capturing complex dependencies that evolve across both time and structure. Existing approaches typically adopt fixed temporal decay schemes or predetermined structural propagation depths, limiting their ability to generalize across graphs with diverse interaction frequencies and topological characteristics. We propose Dual-Scale Retentive Dynamics (DSRD), a unified framework that maintains a retentive representation state encoding both temporal memory and structural context. DSRD introduces two key components: (i) a retentive state with dual-scale adaptation that jointly models temporal dynamics and structural propagation within a single recurrent formulation, and (ii) adaptive decay kernels with learnable time-sensitivity parameters that automatically balance short-term responsiveness and long-term retention based on the underlying interaction patterns. We provide theoretical analysis establishing the equivalence between event-wise parallel aggregation and efficient recurrent state updates, as well as stability and boundedness guarantees for the learned dynamics. Extensive experiments on 14 real-world benchmarks demonstrate that DSRD consistently achieves state-of-the-art performance on both link prediction and node classification tasks, with strong generalization across transductive and inductive settings.

cs.LG

Multiple populations detection with the Chinese Space Station Survey Telescope main survey camera

Multiple stellar populations (MPs), characterized by star-to-star light-element abundance variations, are ubiquitous in globular clusters (GCs). Spectroscopy directly reveals these anomalies, while photometric studies, especially with the \textit{Hubble Space Telescope} (\textit{HST}), have been essential for tracing MP sequences in colour-magnitude diagrams (CMDs). However, the limited field of view of \textit{HST} confines most studies to cluster centres. The upcoming \textit{Chinese Space Station Survey Telescope} (CSST), with its wide field of view and UV-optical coverage, will enable systematic MP studies over entire clusters. We assess the capability of the CSST wide-field camera to detect and characterize MPs in GCs using realistic simulations. Synthetic stellar population models with different helium abundances ($\Delta Y$) and CNO variations were used to simulate CSST observations of GCs at distances of 9.6 and 20~kpc under different exposure times. MP detectability was evaluated using CMDs in seven CSST bands and UV-optical pseudo-colour diagrams. For a GC at 9.6~kpc, the $NUV-u$ colour is highly sensitive to $\Delta Y$ and CNO variations, with separations of $\Delta(NUV-u)\approx0.16$ mag for red giants and up to 0.44 mag for dwarfs. MPs can be resolved when the total UV exposure exceeds $\sim1000$~s and the optical exposure exceeds $\sim300$~s. At 20~kpc, encompassing $\sim80\%$ of Galactic GCs, CSST still retains strong diagnostic power, resolving populations with $\Delta Y\geq0.06$ and $\delta[\mathrm{N/Fe}]\geq0.64$, and separating MPs down to $i\sim19.5$ mag in clusters with large chemical spreads. The $NUV$-$u$-$g$ combination provides diagnostic performance comparable to the \textit{HST} F275W--F336W--F438W system. CSST will enable homogeneous MP surveys across the full spatial extent of star clusters in the Milky Way and nearby galaxies.

astro-ph.SR

Weakly Supervised Camouflaged Object Detection Based on the SAM Model and Mask Guidance

Camouflaged object detection (COD) from a single image is a challenging task due to the high similarity between objects and their surroundings. Existing fully supervised methods require labor-intensive pixel-level annotations, making weakly supervised methods a viable compromise that balances accuracy and annotation efficiency. However, weakly supervised methods often experience performance degradation due to the use of coarse annotations. In this paper, we introduce a new weakly supervised approach for camouflaged object detection to overcome these limitations. Specifically, we propose a novel network, MGNet, which tackles edge ambiguity and missed detections by utilizing initial masks generated by our custom-designed Cascaded Mask Decoder (CMD) to guide the segmentation process and enhance edge predictions. We introduce a Context Enhancement Module(CEM) to reduce the missing detection, and a Mask-guided Feature Aggregation Module (MFAM) for effective feature aggregation. For the weak supervision challenge, we propose BoxSAM, which leverages the Segment Anything Model (SAM) with bounding-box prompts to generate pseudo-labels. By employing a redundant processing strategy, high quality pixel-level pseudo-labels are provided for training MGNet. Extensive experiments demonstrate that our method delivers competitive performance against current state-of-the-art methods.

cs.CV

Learning to Evolve: Multi-modal Interactive Fields for Robust Humanoid Navigation in Dynamic Environments

Safe manipulation-oriented navigation for humanoid robots requires scene memory that remains reliable under locomotion-induced perceptual distortion, environmental changes, and interaction-level geometric safety constraints. Existing semantic mapping and scene-graph systems are difficult to deploy directly in this setting because they often assume stable camera trajectories, static environments, or coarse object geometry. We introduce the Multi-modal Interactive Field (MIF), a humanoid-oriented system that integrates confidence-aware semantic 3D Gaussian Splatting, discrepancy-triggered spatial memory updates, and task-driven geometric reconstruction within a closed-loop perception-adaptation pipeline. MIF couples three fields: an uncertainty-aware 3DGS Appearance Field that suppresses gait-induced blur, a Spatial Field that maintains topological memory, and a Geometry Field that supports Interaction Pose Safety (IPS) before manipulation. A discrepancy detection score is introduced to separate locomotion-induced false-positive changes from persistent changes and updates only locally inconsistent regions. On a Unitree-G1 humanoid in a real dynamic office, MIF improves relocation success in non-static environments from 12% to 94% compared with static scene-graph memory, while reducing semantic memory footprint by 91.4% through feature distillation for practical online operation. Project page and code: https://ziya-jiang.github.io/MIF-homepage/

cs.RO

Design, Testing, and Commissioning of the Sun Yat-sen University (SYSU) 80 cm Infrared Telescope

The Sun Yat-sen University (SYSU) 80 cm telescope is a new generation near-infrared (NIR) facility in China dedicated to time-domain astronomy, while also serving as a testbed for emerging NIR cameras. Commissioned in October 2024 at the 4100 m Lenghu site on the Tibetan Plateau in China, the telescope adopts a reflective Cassegrain design with two Nasmyth foci for J and K bands. The J band imaging system, initially equipped with a 640 x 512 off-the-shelf InGaAs camera (INS Mars640) and upgraded in June 2025 to a 1280 x 1024 science-grade, deeply cooled camera (YNAOIR), achieves background-limited performance with a dark current of ~ 14 e-/s/pix and a readout noise of ~ 11 e-. The system reaches a limiting magnitude of J ~ 17 mag (Vega system) in single 20 s exposures and depths of J ~ 19.4 mag with stacked 30 minute exposures. For a variable with J ~ 14 mag during on-sky tests, the system delivers millimagnitude-level photometric precision. Since commissioning, the telescope observed transients such as gamma-ray bursts (GRBs), supernovae and comets, variables including active galactic nuclei (AGNs), high-redshift quasars (z > 6), and brown dwarfs, as well as deep-field imaging reaching J ~ 20.5 mag. This validates the feasibility of using InGaAs cameras for astronomical observations, encouraging other institutions to develop dedicated infrared telescopes or integrate infrared cameras into existing optical telescopes.

astro-ph.IM

LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment

We present LoD-Loc v3, a novel method for generalized aerial visual localization in dense urban environments. While prior work LoD-Loc v2 achieves localization through semantic building silhouette alignment with low-detail city models, it suffers from two key limitations: poor cross-scene generalization and frequent failure in dense building scenes. Our method addresses these challenges through two key innovations. First, we develop a new synthetic data generation pipeline that produces InsLoD-Loc - the largest instance segmentation dataset for aerial imagery to date, comprising 100k images with precise instance building annotations. This enables trained models to exhibit remarkable zero-shot generalization capability. Second, we reformulate the localization paradigm by shifting from semantic to instance silhouette alignment, which significantly reduces pose estimation ambiguity in dense scenes. Extensive experiments demonstrate that LoD-Loc v3 outperforms existing state-of-the-art (SOTA) baselines, achieving superior performance in both cross-scene and dense urban scenarios with a large margin. The project is available at https://nudt-sawlab.github.io/LoD-Locv3/.

cs.CV

Mock Observations of Multiple Stellar Populations in Tidal Streams of Palomar 5 for the Chinese Space Station Survey Telescope

Observations show that multiple stellar populations (MPs) are ubiquitous in globular clusters. The Hubble Space Telescope (HST) has been a pivotal tool for previous photometric studies of MPs. The Chinese Space Station Survey Telescope (CSST) is a two-meter telescope scheduled for launch. One of its imaging instruments, the Survey Camera (SC), combines ultraviolet sensitivity comparable to that of HST with a significantly larger field of view, making it well-suited for conducting large-scale photometric surveys of MPs within extensive stellar stream structures. In this work, we perform mock observations of the stellar stream Palomar 5 to assess the feasibility of detecting MPs with the CSST/SC. The results indicate that the CSST/SC cannot resolve MPs in stellar streams at distances comparable to Palomar 5 ($\gtrsim 20$ kpc) with one or ten 150 s exposures. This fundamental limitation arises from the absence of the precise proper motions required to disentangle stream members. We estimate that successful resolution would require the target stream to be $\lesssim$ 8 kpc under a 150 s exposure. Furthermore, using theoretical color-magnitude diagrams, we find that the CSST/SC $g$-band provides an optimal balance between contamination rate and completeness rate for member identification in the cluster's core. However, this approach fails in the stream due to severe field star contamination. Therefore, future CSST observations of Palomar 5 and its tidal tails will employ multiple epochs across several bands to obtain the deep photometry and proper motion data for a definitive MP analysis.

astro-ph.IM

PivotAttack: Rethinking the Search Trajectory in Hard-Label Text Attacks via Pivot Words

Existing hard-label text attacks often rely on inefficient "outside-in" strategies that traverse vast search spaces. We propose PivotAttack, a query-efficient "inside-out" framework. It employs a Multi-Armed Bandit algorithm to identify Pivot Sets-combinatorial token groups acting as prediction anchors-and strategically perturbs them to induce label flips. This approach captures inter-word dependencies and minimizes query costs. Extensive experiments across traditional models and Large Language Models demonstrate that PivotAttack consistently outperforms state-of-the-art baselines in both Attack Success Rate and query efficiency.

cs.CL

Spatial Property of Multiple Metallic Populations in the Tidal Stream of {\omega} Centauri

{\omega} Centauri, the remnant nucleus of an accreted dwarf galaxy, is a unique laboratory for studying complex stellar populations. The recently discovered Fimbulthul stream provides a fossil record of its ongoing tidal dissolution. In this work, we investigate the spatial distributions of metal-rich and metal-poor populations within {\omega} Centauri and its stream to constrain the cluster's formation history. Using synthetic photometry from Gaia DR3 XP spectra, we classify stars via a Support Vector Classifier (SVC). The spatial distributions are then compared to a scaling N-body simulation performed with the PeTar code. Our analysis reveals no significant radial gradient in population ratios within the cluster, though the metal-rich stars may be slightly more extended. The population ratio in the tidal stream is consistent with that of the present-day cluster, albeit with large uncertainties. Our simulation indicates that any initial radial gradient must have been shallow, with a maximum fraction difference less than 0.15. Both observational and dynamical results suggest that the metal-rich population is not formed centrally concentrated. By combining our results and existing literature, we propose a new formation scenario for {\omega} Centauri.

astro-ph.GA

RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology

Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as assistive systems that fail to provide both modalities simultaneously offer limited value to medical practitioners. To alleviate this limitation, we first introduce RadDiagSeg-D, a dataset combining abnormality detection, diagnosis, and multi-target segmentation into a unified and hierarchical task. RadDiagSeg-D covers multiple imaging modalities and is precisely designed to support the development of models that produce descriptive text and corresponding segmentation masks in tandem. Subsequently, we leverage the dataset to propose a novel vision-language model, RadDiagSeg-M, capable of joint abnormality detection, diagnosis, and flexible segmentation. RadDiagSeg-M provides highly informative and clinically useful outputs, effectively addressing the need to enrich contextual information for assistive diagnosis. Finally, we benchmark RadDiagSeg-M and showcase its strong performance across all components involved in the task of multi-target text-and-mask generation, establishing a robust and competitive baseline.

cs.CV

A fast powerful X-ray transient from possible tidal disruption of a white dwarf

Stars captured by black holes (BHs) can be torn apart by strong tidal forces, producing electromagnetic flares. To date, more than 100 tidal disruption events (TDEs) have been observed, each involving invariably normal gaseous stars whose debris falls onto the BH, sustaining the flares over years. White dwarfs (WDs), which are the most prevalent compact stars and a million times denser--and therefore tougher--than gaseous stars, can only be disrupted by intermediate-mass black holes (IMBHs) of 10^2--10^5 solar masses. WD-TDEs are considered to generate more powerful and short-lived flares, but their evidence has been lacking. Here we report observations of a fast and luminous X-ray transient EP250702a detected by Einstein Probe. Its one-day-long X-ray peak as luminous as 10^(47-49) erg/s showed strong recurrent flares with hard spectra extending to several tens of MeV gamma-rays, as detected by Fermi/GBM and Konus-Wind, indicating relativistic jet emission. The jet's X-ray dropped sharply from 3 x 10^49 erg/s to around 10^44 erg/s within 20 days (10 days in the source rest frame). These characteristics are inconsistent with any known transient phenomena other than a jetted-TDE evolving over an unprecedentedly short timescale, indicating the disruption of a WD by an IMBH. At late times, a new soft component progressively dominates the X-ray spectrum, exhibiting an extreme super-Eddington luminosity, which possibly originates from an accretion disc. WD-TDEs open a new window for investigating the elusive IMBHs and their surrounding stellar environments, and they are prime sources of gravitational waves in the band of space-based interferometers.

astro-ph.HE

Contour-informed inter-patient deformable registration of Head-and-Neck patients

Background and Purpose: Voxel-based analysis (VBA) helps to identify dose-sensitive regions by aligning individual dose distributions within a common coordinate system (CCS). Accurate deformable image registration (DIR) is essential for addressing anatomical variability across patients. To improve both global and region-specific alignment, we enhanced our in-house DIR algorithm (CPT-DIR) with contour-informed regularisations. We tested its performance for head-and-neck (HN) CT images. Materials and Methods: We developed and evaluated contour-informed CPT-DIR on 37 HN CTs, including 7 with ground-truth dose for warped dose validation. Bone contours were generated using TotalSegmentator, while other organs at risk (OARs) were manually delineated. Contour-based constraints, such as Dice Similarity, were integrated to enhance registration outcome. The global registration results were evaluated using MAE, SSIM and PSNR. Geometric accuracy and warped dose accuracy were assessed using Dice Similarity Coefficient (DSC) and Dose-Organ Overlap (DOO). Constrained and unconstrained CPT-DIR were compared to B-spline. Results: CPT-DIR achieved superior accuracy with a MAE of 98.9\pm6.3 HU, lower than 179.1\pm17.8 HU for B-spline. Incorporating brainstem contours as regularisation improved the DSC from 0.604\pm0.116 to 0.878\pm0.017 and DOO from 0.430\pm0.117 to 0.753\pm0.043 for brainstem. Across all metrics, the enhanced CPT-DIR outperformed the B-spline, confirming its advantages in geometric accuracy. Conclusions: The integration of contour-informed regularisation in CPT-DIR improved DIR accuracy, particularly in dosimetrically relevant regions. This enhanced spatial alignment enabled more precise dose mapping for VBA and demonstrated strong potential for advancing reliable inter-patient dosimetric studies in HN radiotherapy.

physics.med-ph

CPT-4DMR: Continuous sPatial-Temporal Representation for 4D-MRI Reconstruction

Four-dimensional MRI (4D-MRI) is an promising technique for capturing respiratory-induced motion in radiation therapy planning and delivery. Conventional 4D reconstruction methods, which typically rely on phase binning or separate template scans, struggle to capture temporal variability, complicate workflows, and impose heavy computational loads. We introduce a neural representation framework that considers respiratory motion as a smooth, continuous deformation steered by a 1D surrogate signal, completely replacing the conventional discrete sorting approach. The new method fuses motion modeling with image reconstruction through two synergistic networks: the Spatial Anatomy Network (SAN) encodes a continuous 3D anatomical representation, while a Temporal Motion Network (TMN), guided by Transformer-derived respiratory signals, produces temporally consistent deformation fields. Evaluation using a free-breathing dataset of 19 volunteers demonstrates that our template- and phase-free method accurately captures both regular and irregular respiratory patterns, while preserving vessel and bronchial continuity with high anatomical fidelity. The proposed method significantly improves efficiency, reducing the total processing time from approximately five hours required by conventional discrete sorting methods to just 15 minutes of training. Furthermore, it enables inference of each 3D volume in under one second. The framework accurately reconstructs 3D images at any respiratory state, achieves superior performance compared to conventional methods, and demonstrates strong potential for application in 4D radiation therapy planning and real-time adaptive treatment.

cs.CV

Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning

This paper aims to model 3D human motion across domains, where a single model is expected to handle multiple modalities, tasks, and datasets. Existing cross-domain models often rely on domain-specific components and multi-stage training, which limits their practicality and scalability. To overcome these challenges, we propose a new setting to train a unified cross-domain model through a single process, eliminating the need for domain-specific components and multi-stage training. We first introduce Pose-in-Context (PiC), which leverages in-context learning to create a pose-centric cross-domain model. While PiC generalizes across multiple pose-based tasks and datasets, it encounters difficulties with modality diversity, prompting strategy, and contextual dependency handling. We thus propose Human-in-Context (HiC), an extension of PiC that broadens generalization across modalities, tasks, and datasets. HiC combines pose and mesh representations within a unified framework, expands task coverage, and incorporates larger-scale datasets. Additionally, HiC introduces a max-min similarity prompt sampling strategy to enhance generalization across diverse domains and a network architecture with dual-branch context injection for improved handling of contextual dependencies. Extensive experimental results show that HiC performs better than PiC in terms of generalization, data scale, and performance across a wide range of domains. These results demonstrate the potential of HiC for building a unified cross-domain 3D human motion model with improved flexibility and scalability. The source codes and models are available at https://github.com/BradleyWang0416/Human-in-Context.

cs.CV