SearcharxivSearch

arXiv subjects

Ruizhi Zhang

Publications and source records attributed to Ruizhi Zhang.

At least 19 recordsLinked to original sources

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time

While artificial intelligence has mastered structured games like chess and Go, vision-language agents still struggle in visually-driven 3D games without access to game states. Existing game environments typically evaluate a fixed agent configuration, rather than an agent's ability to improve its configuration across consecutive episodes of the same task---a paradigm known as test-time learning (TTL). Furthermore, current TTL methods typically optimize single modalities---such as text prompts or actions---in isolation, ignoring the synergy between perception, reasoning, and control. To bridge these gaps, we first introduce \textbf{PokeGym}, a long-horizon benchmark built upon the 3D open-world game Pokémon Legends: Z-A, where agents act from visual observations without access to game states, designed to evaluate an agent's ability to learn and adapt across consecutive episodes of the task. To tackle this challenging environment, we propose Graph-Guided Evolutionary Multimodal Agent Configuration (\textbf{G-EvoMAC}), a graph-guided framework that jointly optimizes visual perception, strategy, and action set synergistically. Extensive experiments show that G-EvoMAC achieves a 60.18\% average success rate on PokeGym, outperforming the strongest baseline by over 11 percentage points, validating the power of cross-modal co-evolution.

cs.CV

Multi-User Localization via Active Sensing with Electromagnetically Reconfigurable Antennas

This paper investigates multi-user localization in uplink wireless systems assisted by electromagnetically reconfigurable antennas (ERAs). Unlike traditional localization schemes, we formulate an active sensing problem where a base station (BS) exploits historical pilot observations accumulated over previous sensing stages to adapt the shared ERA configuration and progressively refine position estimates. To capture both theoretical flexibility and practical hardware constraints, we establish a unified wideband geometric signal model accommodating two complementary ERA paradigms: a synthesis-based model utilizing spherical-harmonic basis functions, and a finite-state model based on measured radiation codebooks. Because analytically solving the resulting joint design problem is highly intractable due to the high-dimensional observation and the shared-aperture coupling among multiple users, we develop a learning-based active sensing framework. Specifically, pilot-matched wideband observations are compressed into compact user-wise features and sequentially accumulated by a long short-term memory (LSTM) module. These temporal features are then processed by a graph neural network (GNN) to capture multi-user shared-aperture coupling. Model-specific output heads generate either continuous synthesis coefficients or finite-state ERA selections, while a localization head produces stage-wise position estimates. Numerical results under a specific channel distribution show that the proposed ERA-assisted active sensing framework achieves progressive localization refinement across sensing stages and obtains better performance than conventional non-reconfigurable arrays and representative ablation baselines.

eess.SP

Disentangling the Distant Stellar Halo Using K-Giants in the DESI Year 3 Data

We present a sample of 88,959 K-giants from DESI Milky Way Survey Year 3 data, which we use to characterize the chemo-dynamical properties of the stellar halo at Galactocentric distances of 12 to ~100 kpc. Using HDBSCAN, we identify five prominent stellar halo substructures: Aleph, the Sagittarius stream, Gaia Sausage-Enceladus (GSE), Cetus-Palca and the Orphan-Chenab stream. We present the properties of each of these structures as they appear in our catalog, and examine how uncertainties on distance affect the characterization of substructure with this approach. We also examine regions associated with previously reported overdensities (such as the Virgo Overdensities and the Sagittarius spur) that we do not recover with HDBSCAN. The size and distance range of our catalog allows us to explore in detail the residual stellar halo, comprising stars that we do not associate with any substructure. We find that samples of ~2000 outer halo stars with both highly prograde and highly retrograde angular momenta have similar metallicity distribution functions (MDFs), which do not resemble the MDFs of either GSE or Sagittarius. Both the prograde and retrograde residual halo MDFs are bimodal, with a metal-poor peak at [Fe/H] ~ -2 and a metal-rich peak at [Fe/H] ~ -1.3 (prograde) or -1.5 (retrograde). The MDF for lower angular momentum residual halo K-giants does not show clear evidence for a metal-poor peak, and broadly resembles the MDF of GSE, even at much lower binding energies than GSE itself. We discuss possible interpretations of these findings for GSE accretion scenarios.

astro-ph.GA

LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. https://github.com/BDML-lab/LongNovel

cs.CL

Cross-Stage Attention Multi-Expert Network for Radiologist-Inspired Breast Ultrasound Diagnosis

Breast ultrasound imaging is an important noninvasive method for early breast cancer diagnosis, but automatic benign/malignant classification remains challenging due to tumor heterogeneity, blurred boundaries, and data imbalance. To improve feature representation and classification accuracy, this paper proposes the Cross-Stage Attention Mixture-of-Experts Network (CSA-MoE-Net). It adopts a Cross-Stage Attention-enhanced ResNet-18 as the backbone, in which the Cross-Stage Attention module adaptively recalibrates multi-level features, thereby enhancing key tumor features and suppressing redundancy. A three-branch Mixture of Experts (MoE) Block learns complementary features from the Whole Tumor Image, Tumor Core, and Boundary, and an Adaptive Gating Network fuses them to capture morphological, textural, and contextual information. The fused features are denoted as Fused Expert Feature (FEF) in the architecture. Experiments on a balanced dataset of 2,129 breast ultrasound images show that, averaged over 20 independent runs, the model achieves an accuracy of 96.33\%, precision of 94.09\%, recall of 98.53\%, F1-score of 96.25\%, and AUC of 99.50\%. Compared to the baseline ResNet-18, these metrics improve by 3.01, 0.70, 5.37, 2.98, and 5.42 percentage points, respectively. The proposed mechanism requires no invasive modification and can be seamlessly embedded into VGG-16, DenseNet-121, etc., yielding stable performance gains, thus providing reliable support for computer-aided diagnosis.

cs.CV

Exploring the ex-situ components within Gaia DR3

The presence of Gaia DR3 provides a large sample of stars with complete 6D information, offering a fertile ground for the exploration of stellar objects that were accreted to the Milky Way through ancient merger events. In this study, we developed a deep learning methodology to identify ex-situ stars within the Gaia DR3 catalogue. After two phases of training, our neural network (NN) model was capable of performing binary classification of stars based on input data consisting of 3D position and velocity, as well as actions. From the target sample of 27,085,748 stars, our NN model managed to identify 160,146 ex-situ stars. The metallicity distribution suggests that this ex-situ sample comprises multiple components but appears to be predominated by the Gaia-Sausage-Enceladus. We identified member stars of the Magellanic Clouds, Sagittarius, and 20 globular clusters throughout our examination. Furthermore, an extensive group of member stars from Gaia-Sausage-Enceladus, Thamnos, Sequoia, Helmi streams, Wukong, and Pontus were meticulously selected, constituting an ideal sample for the comprehensive study of substructures. Finally, we conducted a preliminary estimation to determine the proportions of ex-situ stars in the thin disc, thick disc, and halo, which resulted in percentages of 0.1%, 1.6%, and 63.2%, respectively. As the vertical height from the Galactic disc and distance from the Galactic centre increased, there was a corresponding upward trend in the ex-situ fraction of the target sample.

astro-ph.GA

Sequential Change Detection for Multiple Data Streams with Differential Privacy

Sequential change-point detection seeks to rapidly identify distributional changes in streaming data while controlling false alarms. Existing multi-stream detection methods typically rely on non-private access to raw observations or intermediate statistics, limiting their usage in privacy-sensitive settings. We study sequential change-point detection for multiple data streams under differential privacy constraints. We consider multiple independent streams undergoing a synchronized change at an unknown time and in an unknown subset of streams, and propose DP-SUM-CUSUM, a differentially private detection procedure based on the summation of per-stream CUSUM statistics with calibrated Laplace noise injection. We show that DP-SUM-CUSUM satisfies sequential $\varepsilon$-differential privacy and derive bounds on the average run length to false alarm and the worst-case average detection delay, explicitly characterizing the privacy--efficiency tradeoff. A truncation-based extension is also presented to handle distributional shifts with unbounded log-likelihood ratios. Simulations and experiments on an Internet of Things (IoT) botnet dataset validate the proposed approach.

math.ST

DW-DGAT: Dynamically Weighted Dual Graph Attention Network for Neurodegenerative Disease Diagnosis

Parkinson's disease (PD) and Alzheimer's disease (AD) are the two most prevalent and incurable neurodegenerative diseases (NDs) worldwide, for which early diagnosis is critical to delay their progression. However, the high dimensionality of multi-metric data with diverse structural forms, the heterogeneity of neuroimaging and phenotypic data, and class imbalance collectively pose significant challenges to early ND diagnosis. To address these challenges, we propose a dynamically weighted dual graph attention network (DW-DGAT) that integrates: (1) a general-purpose data fusion strategy to merge three structural forms of multi-metric data; (2) a dual graph attention architecture based on brain regions and inter-sample relationships to extract both micro- and macro-level features; and (3) a class weight generation mechanism combined with two stable and effective loss functions to mitigate class imbalance. Rigorous experiments, based on the Parkinson Progression Marker Initiative (PPMI) and Alzheimer's Disease Neuroimaging Initiative (ADNI) studies, demonstrate the state-of-the-art performance of our approach.

cs.CV

Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging

Recent advances in Large Language Model (LLM)-based agents have shown remarkable progress in code generation. However, current agent methods mainly rely on text-output-based feedback (e.g. command-line outputs) for multi-round debugging and struggle in graphical user interface (GUI) that involve visual information. This is mainly due to two limitations: 1) GUI programs are event-driven, yet existing methods cannot simulate user interactions to trigger GUI element logic 2) GUI programs possess visual attributes, making it difficult for text-based approaches to assess whether the rendered interface meets user needs. To systematically address these challenges, we first introduce InteractGUI Bench, a novel benchmark comprising 984 commonly used real-world desktop GUI application tasks designed for fine-grained evaluation of both interaction logic and visual structure. Furthermore, we propose VF-Coder, a vision-feedback-based multi-agent system for debugging GUI code. By perceiving visual information and directly interacting with program interfaces, VF-Coder can identify potential logic and layout issues in a human-like manner. On InteractGUI Bench, our VF-Coder approach increases the success rate of Gemini-3-Flash from 21.68% to 28.29% and raises the visual score from 0.4284 to 0.5584, indicating the effectiveness of visual feedback in GUI debugging.

cs.SE

A 3D Chemodynamical Census of Inner-Galaxy Metal-poor Giants to [Fe/H]~-3.5

The earliest assembly of the Milky Way remains poorly understood, yet the spatial, chemical, and kinematic properties of its most metal-poor stars provide a unique fossil record of its proto-Galaxy phase. Understanding how this ancient component formed is essential for linking near-field Galactic archaeology to high-redshift galaxy evolution. We construct the currently largest 3D map of inner-Galaxy metal-poor giants by combining several narrow/medium-band photometric surveys, reaching metallicities down to $\mathrm{[Fe/H]}\sim-3.5$. Our final sample contains 5,095,676 giants, including 1,717,610 stars with $\mathrm{[Fe/H]}<-1$. Across $-4\le \mathrm{[Fe/H]}<-1$, the density distribution reveals a centrally concentrated, flattened spheroidal component extending to $r_{\rm gc}\sim15$ kpc, together with a prominent overdensity near $X\sim-5$ kpc that is dominated by metal-poor stars on disklike orbits, with a kinematically hot background also present. The selection-function-corrected metallicity distribution function shows a distinct, very metal-poor component around $\mathrm{[Fe/H]}\sim-2.7$ that becomes most prominent at 1$<r_{\rm gc}<$3 kpc. Stars with $-3.5\lesssim\mathrm{[Fe/H]}\lesssim-1.4$ exhibit weak net rotation and low rotational support within $r_{\rm gc}<15$ kpc. Finally, we briefly note that the centrally enhanced very metal-poor component could be qualitatively consistent with one or more early dissipative build-up episodes (e.g., high-$z$ compaction/"blue-nugget" phases) as one possible interpretation.

astro-ph.GA

Physics Informed Deep Unfolded Full Waveform Inversion for Edema Detection

Edema is a potential indicator of underlying pathological changes. However, its low-contrast signature is often masked in conventional B-mode imaging by strong scatterers, making reliable detection challenging. Ultrasound (US) provides a non-invasive, non-ionizing, and cost-efficient imaging option that is widely used. Conventional techniques, which rely on beamforming, often lack sufficient physical interpretability. Quantitative US (QUS) can estimate physical properties such as the speed of sound (SoS) and density by solving a physics-based inverse problem directly on the measured US wavefields, i.e., the raw per-element channel data (CD), to recover their spatial distribution. However, state-of-the-art physics-based inversion methods, including full waveform inversion (FWI) and model-based quantitative radar and US (MB-QRUS), are computationally intensive and susceptible to local minima, which constrains their clinical utility. We introduce deep unfolded FWI (DUFWI), a physics-faithful unfolded iterative inversion method that exhibits FWI-like refinement behavior while learning the update rule from data, requiring only a small number of iterations for real-time SoS reconstruction. Across both simulated datasets and hardware measurements acquired with a Verasonics US system, the DUFWI significantly outperforms classical FWI and MB-QRUS in reconstruction quality while maintaining high computational efficiency. These results demonstrate real-time edema diagnosis in both simulation and hardware experiments, with phantom-based validation using cylindrical rods, supporting practical deployment under typical US imaging setting.

eess.SP

Sim2Sea: Sim-to-Real Policy Transfer for Maritime Vessel Navigation in Congested Waters

Autonomous navigation in congested maritime environments is a critical capability for a wide range of real-world applications. However, it remains an unresolved challenge due to complex vessel interactions and significant environmental uncertainties. Existing methods often fail in practical deployment due to a substantial sim-to-real gap, which stems from imprecise simulation, inadequate situational awareness, and unsafe exploration strategies. To address these, we propose \textbf{Sim2Sea}, a comprehensive framework designed to bridge simulation and real-world execution. Sim2Sea advances in three key aspects. First, we develop a GPU-accelerated parallel simulator for scalable and accurate maritime scenario simulation. Second, we design a dual-stream spatiotemporal policy that handles complex dynamics and multi-modal perception, augmented with a velocity-obstacle-guided action masking mechanism to ensure safe and efficient exploration. Finally, a targeted domain randomization scheme helps bridge the sim-to-real gap. Simulation results demonstrate that our method achieves faster convergence and safer trajectories than established baselines. In addition, our policy trained purely in simulation successfully transfers zero-shot to a 17-ton unmanned vessel operating in real-world congested waters. These results validate the effectiveness of Sim2Sea in achieving reliable sim-to-real transfer for practical autonomous maritime navigation.

cs.RO

Swimming Under Constraints: A Safe Reinforcement Learning Framework for Quadrupedal Bio-Inspired Propulsion

Bio-inspired aquatic propulsion offers high thrust and maneuverability but is prone to destabilizing forces such as lift fluctuations, which are further amplified by six-degree-of-freedom (6-DoF) fluid coupling. We formulate quadrupedal swimming as a constrained optimization problem that maximizes forward thrust while minimizing destabilizing fluctuations. Our proposed framework, Accelerated Constrained Proximal Policy Optimization with a PID-regulated Lagrange multiplier (ACPPO-PID), enforces constraints with a PID-regulated Lagrange multiplier, accelerates learning via conditional asymmetric clipping, and stabilizes updates through cycle-wise geometric aggregation. Initialized with imitation learning and refined through on-hardware towing-tank experiments, ACPPO-PID produces control policies that transfer effectively to quadrupedal free-swimming trials. Results demonstrate improved thrust efficiency, reduced destabilizing forces, and faster convergence compared with state-of-the-art baselines, underscoring the importance of constraint-aware safe RL for robust and generalizable bio-inspired locomotion in complex fluid environments.

cs.RO

Give me scissors: Collision-Free Dual-Arm Surgical Assistive Robot for Instrument Delivery

During surgery, scrub nurses are required to frequently deliver surgical instruments to surgeons, which can lead to physical fatigue and decreased focus. Robotic scrub nurses provide a promising solution that can replace repetitive tasks and enhance efficiency. Existing research on robotic scrub nurses relies on predefined paths for instrument delivery, which limits their generalizability and poses safety risks in dynamic environments. To address these challenges, we present a collision-free dual-arm surgical assistive robot capable of performing instrument delivery. A vision-language model is utilized to automatically generate the robot's grasping and delivery trajectories in a zero-shot manner based on surgeons' instructions. A real-time obstacle minimum distance perception method is proposed and integrated into a unified quadratic programming framework. This framework ensures reactive obstacle avoidance and self-collision prevention during the dual-arm robot's autonomous movement in dynamic environments. Extensive experimental validations demonstrate that the proposed robotic system achieves an 83.33% success rate in surgical instrument delivery while maintaining smooth, collision-free movement throughout all trials. The project page and source code are available at https://give-me-scissors.github.io/.

cs.RO

User Localization via Active Sensing with Electromagnetically Reconfigurable Antennas

This paper presents an end-to-end deep learning framework for electromagnetically reconfigurable antenna (ERA)-aided user localization with active sensing, where ERAs provide additional electromagnetic reconfigurability to diversify the received measurements and enhance localization informativeness. To balance sensing flexibility and overhead, we adopt a two-timescale design: the digital combiner is updated at each stage, while the ERA patterns are reconfigured at each substage via a spherical-harmonic representation. The proposed mechanism integrates attention-based feature extraction and LSTM-based temporal learning, enabling the system to learn an optimized sensing strategy and progressively refine the UE position estimate from sequential observations. Simulation results show that the proposed approach consistently outperforms conventional digital beamforming-only and single-stage sensing baselines in terms of localization accuracy. These results highlight the effectiveness of ERA-enabled active sensing for user localization in future wireless systems.

eess.SP

Sequential Change Detection with Differential Privacy

Sequential change detection is a fundamental problem in statistics and signal processing, with the CUSUM procedure widely used to achieve minimax detection delay under a prescribed false-alarm rate when pre- and post-change distributions are fully known. However, releasing CUSUM statistics and the corresponding stopping time directly can compromise individual data privacy. We therefore introduce a differentially private (DP) variant, called DP-CUSUM, that injects calibrated Laplace noise into both the vanilla CUSUM statistics and the detection threshold, preserving the recursive simplicity of the classical CUSUM statistics while ensuring per-sample differential privacy. We derive closed-form bounds on the average run length to false alarm and on the worst-case average detection delay, explicitly characterizing the trade-off among privacy level, false-alarm rate, and detection efficiency. Our theoretical results imply that under a weak privacy constraint, our proposed DP-CUSUM procedure achieves the same first-order asymptotic optimality as the classical, non-private CUSUM procedure. Numerical simulations are conducted to demonstrate the detection efficiency of our proposed DP-CUSUM under different privacy constraints, and the results are consistent with our theoretical findings.

math.ST

HR-GO II: chemical abundances of low-$E$ retrograde dynamically-tagged-groups: Revealing Thamnos as a very metal-poor substructure

Milky Way halo substructures identified in dynamical space are known to suffer from contamination from the Milky Way in-situ stars, which makes their accreted origins uncertain. We present detailed chemical abundances of 35 stars belonging to two sets of dynamically tagged groups, Rg8 and Rg9, to investigate their accreted nature. Both groups are composed of stars with low orbital energy and very retrograde orbits. We find that Rg8 and Rg9 are chemically indistinguishable across all elements, from C to Eu, strongly indicating that they belong to the same structure. The iron-abundance distribution of this low-$E$ retrograde group has a prominent peak at [Fe/H] $\approx-2.1$, revealing that its main population is very metal-poor, and a secondary peak at [Fe/H] $\approx-1.5$, very likely due to contamination from Milky Way in-situ stars. These groups also heavily overlap with the Thamnos substructure in dynamical space, and we thus use them to investigate the chemical properties of Thamnos. The dominant, low-metallicity population provides strong evidence for the ex-situ origin of Thamnos, as well as its very metal-poor nature. We do not see any evidence of an $α$ knee in our sample, which is consistent with previous studies. Comparison with the Cetus-Palca stream in the chemical space shows similar abundance distributions, and thus it suggests that the Thamnos progenitor dwarf galaxy had a truncated star formation history due to its early merger with the Milky Way.

astro-ph.GA

Wearable-informed generative digital avatars predict task-conditioned post-stroke locomotion

Dynamic prediction of locomotor capacity after stroke could enable more individualized rehabilitation, yet current assessments largely provide static impairment scores and do not indicate whether patients can perform specific tasks such as slope walking or stair climbing. Here, we present a wearable-informed data-physics hybrid generative framework that reconstructs a stroke survivor's locomotor control from wearable inertial sensing and predicts task-conditioned post-stroke locomotion in new environments. From a single 20 m level-ground walking trial recorded by five IMUs, the framework personalizes a physics-based digital avatar using a healthy-motion prior and hybrid imitation learning, generating dynamically feasible, patient-specific movements for inclined walking and stair negotiation. Across 11 stroke inpatients, predicted postures reached 82.2% similarity for slopes and 69.9% for stairs, substantially exceeding a physics-only baseline. In a multicentre pilot randomized study (n = 21; 28 days), access to scenario-specific locomotion predictions to support task selection and difficulty titration was associated with larger gains in Fugl-Meyer lower-extremity scores than standard care (mean change 6.0 vs 3.7 points; $p < 0.05$). These results suggest that wearable-informed generative digital avatars may augment individualized gait rehabilitation planning and provide a pathway toward dynamically personalized post-stroke motor recovery strategies.

cs.CE