SearcharxivSearch

arXiv subjects

Rui Zhang

Publications and source records attributed to Rui Zhang.

At least 19 recordsLinked to original sources

LeCor: Learning to Be Corrected by Meta-Learned Test-Time Training for Interactive 3D Lung-Tumour Segmentation

Delineating lung tumours on computed tomography (CT) takes a considerable share of the time spent on radiotherapy planning, and a contour proposed by a model can be refined interactively by the clinician. Promptable foundation models such as SAM 3 support this workflow by writing each correction into a session memory that conditions the remaining slices, while the model weights stay fixed. On 690 test cases from five public CT cohorts, fine-tuning SAM 3 on lung tumours raises the Dice obtained from a single point prompt from 0.298 to 0.757, and seven rounds of corrections raise it further to 0.765, but under memory conditioning alone the accuracy on slices the annotator has not touched stops improving after six rounds. We therefore treat each correction as a training signal and propose LeCor, which performs test-time training on a small set of case adapters that are reset for every case and meta-learned such that a single gradient step driven by a click improves the slices that were not clicked. On the 133 test cases that span at least eight slices, LeCor raises the Dice reached after seven correction rounds from 0.787 with the fine-tuned model to 0.827, reduces the number of cases that never reach a Dice of 0.80 from 47 to 27, and reaches in three correction rounds the accuracy that the fine-tuned model attains in seven.

cs.CV

Physical-Field Reconstruction from Sparse Observations: When Are Diffusion Models Preferable to Deterministic Regression?

Reconstructing physical fields from sparse observations is central to system identification, forecasting, and control, yet sparse measurements generally underdetermine the full field. This makes reconstruction an ill-posed inverse problem rather than simple interpolation. Although many deterministic and generative methods have been developed, there is still no clear consensus on when a single point estimate is sufficient and when a distribution of plausible reconstructions is more useful. We conduct a fair comparison of a deterministic U-Net, conditional diffusion, and prior-guided diffusion under matched experimental settings, including 2D Poisson equation, 2D Navier-Stokes flow, and 1D Kuramoto-Sivashinsky dynamics. Through this comparison, we make three observations. First, accuracy is field- and regime-dependent, with no systematic advantage for diffusion under higher complexity or sparser observations. Second, ensemble means improve phase-aligned accuracy, whereas individual samples better preserve variability and can retain high-wavenumber power in selected regimes. Third, conditional diffusion provides more reliable uncertainty estimates at lower cost, while prior-guided diffusion is more robust to mask-distribution shifts but requires substantially higher inference cost and guidance tuning. These results clarify when generative reconstruction is useful and provide guidance for improving uncertainty estimation, fine-scale sample fidelity, robustness, and computational efficiency in sparse field reconstruction.

physics.comp-ph

Exploiting LLM Agents for Trustworthy AutoResearch in Wireless Communications

Large language model (LLM) agent-enabled AutoResearch is attracting growing interest across scientific disciplines, in which LLMs are leveraged for knowledge synthesis, multistep planning, code generation, tool invocation, and iterative refinement, thus automating the research lifecycle, from hypothesis generation and experimentation to analysis and manuscript preparation. Wireless communications is particularly suitable to this paradigm. This is due to the fact that advances in this field often rely on fundamental-limit analysis, system optimization, and protocol design, all supported by mature mathematical, simulation, and optimization toolchains, as well as standards, measurement, and digital-twin platforms. However, integrating autonomous agents into rigorous wireless research workflows requires traceable processes and verifiable evidence. This article presents a general trustworthy wireless AutoResearch framework. This framework employs typed research contracts, version-controlled artifacts, independent validators, and bounded agent authority to connect hypothesis generation with system modeling, mathematical formu lation, algorithm design, code generation, simulation, reproducible claims, and manuscript preparation. We conduct a case study on integrated sensing and communication (ISAC) for unmanned aerial vehicles (UAVs). It is shown that the proposed framework can generate innovative research ideas and, with appropriate expert intervention, produce a manuscript whose evaluation score is comparable to or higher than those of related IEEE conference and letter papers. These findings demonstrate the effectiveness of LLM agents in orchestrating authoritative scientific tools and versioned research artifacts, while underscoring the importance of human-AI collaboration.

eess.SP

Scratchy: Visual-Scratchpad Multimodal Reasoning for Cryptographic Proof Generation in EasyCrypt

Large language models (LLMs) have recently made substantial progress in formal proof generation, yet presenting distinctive challenges in cryptographic area. Computational security arguments posit that a valid proof must coordinate probability, adversarial games, invariants, assumptions and bounds, which can be provided by a machine-checked framework named EasyCrypt. Although all objects may appear in available context, LLMs still struggle because proof-theoretic dependencies are typically implicit in a linear representation and distributed across multiple programs. So, this paper presents Scratchy, a visual-scratchpad approach that exposes these dependencies for multimodal generation. Given the natural-language security description, with formal context and target propositions, the proof objects can be normalized into a typed proof-relation graph. Then a structure-preserving visual compiler transforms the graph into the formula-rich visual proof state that guides a multimodal model in generating the EasyCrypt proof. Also, the Scratchy-eval, a 114-task dataset derived from reliable official EasyCrypt files, has been introduced. It contains 64 security-form proof generations and 50 multiple-choice knowledge tests. After a series of evaluations, covering semantic grounding, relational invariants, and game reductions, classical LLMs like GPT-5.6-Sol and Claude-Opus-5 have gained a clear advantage from Scratchy's structured visual proof states. This contrast suggests that explicit proof structure can make the improvement and multimodal proof-state representation as a promising direction for computer-aided cryptography.

cs.CR

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral learning rapidly improves policy performance while enhancing online experience quality. As autonomous experience accumulates, global return propagation and local preference ranking progressively calibrate value estimates, yielding relative action advantages for reference-regularized policy improvement while suppressing drift. To enable ACoB on large VLAs, we develop ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, delivering up to 10.9$\times$ improvements in throughput and computational efficiency. Extensive evaluations on nine high-precision chemistry tasks across four categories and four robot embodiments show that VLA-Precision achieves 98.3\% mean success rate in 45.8 min/task, with 27.6 s episodes running at 1.2$\times$ and 1.8$\times$ the speeds of VLA and RL baselines. Resources are available at https://vla-precision.github.io.

cs.RO

A New Paradigm of 6G Networks: Proactive Channel Cognition and Reconfiguration

The sixth-generation (6G) wireless networks are expected to enable the deep integration of communication, sensing, computing, control, and intelligence in highly dynamic environments. This evolution drives a fundamental transition from conventional passive channel adaptation to proactive channel cognition and reconfiguration, wherein wireless channels are no longer regarded as uncontrollable propagation media but as network resources that can be learned, predicted, and actively reconfigured. This paper presents a comprehensive overview of this emerging paradigm. We first review channel cognition through the channel knowledge map (CKM) as a systematic framework for learning and exploiting channel characteristics across space, time, and frequency domain. The definitions, construction methods, and applications in wireless networks of CKMs are comprehensively reviewed. Building upon channel cognition, we then review channel reconfiguration technologies from two complementary perspectives: transceiver-side reconfiguration enabled by movable antennas (MAs) and environment-side reconfiguration enabled by intelligent reflecting surfaces (IRSs). For both MA- and IRS-enabled wireless systems, we review their architectures, performance advantages, and key design challenges. Finally, we discuss several promising research directions to inspire further innovations in this burgeoning field.

eess.SP

ScentEcho: Exploring Adsorbent Materials for Accurate Odor Collection and Playback

Delivering odors that feel realistic and recognizable remains a core challenge for olfactory interaction systems, particularly in applications that demand precise scent delivery. A key limitation lies in the difficulty of capturing, preserving, and playing back real-world scent sources in a reliable and scalable manner. This study explores the potential of adsorbent materials for supporting realistic scent playback. We present ScentEcho, a portable system that enables modular scent collection and release. Through user evaluations, we identify which adsorbent materials tend to perform better for specific odors, and observe that perceived intensity strongly influences similarity ratings. In addition, odor recognition follows a graded pattern, with users moving from broad category identification to more specific source recognition as similarity increases. These findings offer practical insights for designing olfactory interfaces that are both expressive and perceptually aligned with user expectations.

cs.HC

Automated Shape-Model-Based Astrometry of Phobos from Mars Express SRC Images

High-resolution spacecraft images provide important astrometric constraints for orbit refinement, but measurements of resolved bodies are often limited by labor-intensive control-point selection and the difficulty of achieving consistent reductions over large image archives. We present an automated shape-model-based astrometric pipeline for Phobos and apply it to Mars Express Super Resolution Channel (SRC) images. For each exposure, a synthetic image is rendered from a high-resolution 3D shape model under the nominal spacecraft-target-Sun geometry. Feature correspondences between the observed and synthetic images are established using SuperPoint and SuperGlue, followed by RANSAC filtering. The matched synthetic-image keypoints are then associated with surface points through ray-shape intersection. The geometric adjustment fixes the adopted body orientation, spacecraft state, and corrected camera pointing and estimates only two effective plane-of-sky position offsets using the exact perspective-projection model. These offsets are used to derive the center-of-figure position of Phobos. We first test the method on an image set previously analysed with a control-point approach and obtain comparable astrometric performance. We then extend the analysis to a larger SRC dataset spanning 2007-2025 and obtain 1113 successful measurements. Relative to the JPL MAR099 ephemeris, the resulting observed-minus-computed residuals have mean values of 0.186 km in $\alpha \times cos(\delta)$ and 0.053 km in $\delta$, with corresponding standard deviations of 0.609 km and 0.583 km. These results demonstrate that the proposed pipeline provides a practical approach to large-scale, homogeneous astrometric reduction of archival spacecraft images of Phobos, with potential application to other resolved bodies.

astro-ph.IM

Adaptive Payload-Aided Near-Field Beam Tracking via Thompson Sampling

Extremely large antenna arrays at high frequencies substantially extend the radiative near-field region in 6G networks, making mobile beam alignment depend jointly on user angle and range. The added range dimension enlarges the beam-search space, making repeated pilot-based sweeping costly under mobility, while sensing-assisted tracking depends on propagation conditions, echo quality, and target reflectivity. We propose an adaptive payload-aided near-field beam-tracking framework based on maximum likelihood estimation (MLE) and Thompson sampling (TS). Selected received payload samples are fed back and reused as tracking observations, avoiding dedicated beam-sweeping symbols during tracking. Within each sliding window, local angle and range trajectories are modeled by low-order polynomials to capture velocity, acceleration, and higher-order motion variations, and are estimated by MLE using the spherical-wave channel model. A local Gaussian approximation centered at the MLE, with covariance from the inverse observed Fisher information, represents trajectory uncertainty. For payload transmissions selected for feedback, TS samples a trajectory hypothesis and maps it to a payload beam toward the sampled state, while the remaining transmissions use the MLE-predicted beam. To handle nonstationary mobility, an asymptotic chi-square characterization of in-window estimation risk motivates joint adaptation of observation-window length and polynomial degree, while online residual statistics adjust the update interval and feedback ratio. The framework is also extended to uniform planar arrays with elevation tracking. Simulations under smooth and sharp-turn trajectories show high payload-accounted mean normalized beamforming gain, low normalized-gain variance, and high effective-symbol reliability, while the adaptive mechanism provides substantial robustness under nonstationary sharp-turn mobility.

eess.SP

The number of limit cycles of piecewise linear Li\'enard systems

For the planar Li\'{e}nard differential system $\dot{x}=F(x)-y$, $\dot{y}=x$, where $F(x)$ is a piecewise linear function, Tonnelier (SIAM J. Appl. Math., 2002) conjectured that the maximum number of limit cycles of the system is $n$ when $F(x)$ has $n$ fold points and no jump points, and $2n$ when $F(x)$ has $n$ jump points and no fold points. This conjecture was confirmed by Llibre et al. (J. Nonlinear Sci., 2015) (resp. Chen et al. (J. London Math. Soc., 2026a)) when $F(x)$ has one fold point and no jump points (resp. two fold points). More recently, Chen et al. (J. London Math. Soc., 2026b) proved that the conjecture is correct when $F(x)$ has no fold points and one jump point. All other cases remain open. Here we verify that the lower bound for the maximum number of limit cycles of the system can be $n$ when $F(x)$ has only $n$ fold points, and $2n$ when $F(x)$ has only $n$ jump points, thereby confirming the lower bound part of Tonnelier's conjecture. Moreover, when $F(x)$ has $m$ jump points and $n-m$ fold points, $0\le m\le n$, we also show that the system can have $n+m=(n-m)+2m$ limit cycles. In addition, a complete classification of the {dynamics} near infinity for this class of systems is provided.

math.DS

Universal CKM for Environment-Aware Wireless Networks: Enabling Cross-Device and Cross-Task Channel Knowledge Transfer

Channel knowledge map (CKM) is a promising technology for environment-aware sixth-generation (6G) wireless networks. However, most existing CKMs are tightly coupled with wireless devices and downstream tasks, which limit their scalability and reusability in wireless networks. To address these limitations, this article proposes the concept of universal CKM (uCKM) as a foundational wireless environment prior, which aims to enable cross-device and cross-task channel knowledge transfer for environment-aware wireless networks. We first revisit the representative CKMs and discuss their limitations. Then, the uCKM-enabled new paradigm for environment-aware wireless networks is introduced, and its benefits are highlighted from the perspectives of uCKM construction and utilization phases, for which we propose the visions of ``All for uCKM'' and ``uCKM for All'', i.e., the data acquired by all devices and tasks should contribute to the construction of uCKM, and vice versa. Subsequently, we discuss the main challenges of uCKM and propose potential solutions. Last, we provide simulation results to demonstrate the feasibility and performance gains brought by uCKM and outline future research directions.

cs.IT

Reduced-Order Physics-Informed Neural Network with Adaptive Basis Refinement for Structural Identification

Physics-informed neural networks (PINNs) provide a flexible framework for solving forward and inverse problems. However, their direct application to structural dynamics remains limited by high system dimensionality and model-form errors arising from incomplete physics. Reduced-order models (ROMs) can alleviate the dimensionality bottleneck, yet existing PINN-ROM couplings typically rely on fixed reduced subspaces, target forward simulations, or assume complete physics, restricting their use for inverse identification under parametric variability or incomplete system knowledge. To address these limitations, this work proposes a Reduced-Order Physics-Informed Neural Network (RO-PINN) framework with adaptive basis refinement for structural identification under known and incomplete physics. Via projection, reduced governing equations are embedded directly into the PINN loss, facilitating learning in a low-dimensional latent space. An adaptive scheme updates the projection basis during training so that the latent space is progressively realigned with evolving structural parameters or learned residual restoring forces. This realignment reduces basis-mismatch errors and limits their influence on the inferred residual force. The method is validated on a four-story steel frame with nonlinear hysteretic braces under sparse and noisy measurements. Results show parameter identification comparable to or more accurate than Bayesian model updating with lower computational cost in the considered cases, recovery of unmodeled nonlinear restoring forces under incomplete physics, and joint identification of residual restoring forces and structural parameters within the same framework. Overall, RO-PINN provides a unified framework for structural identification by integrating reduced-order modeling, adaptive basis refinement, and physics-informed learning within a single formulation.

cs.CE

Tensor Decomposition-Based Wireless Sensing for MIMO-OFDM ISAC via Flexible Spatial-Temporal-Spectral Optimization

Integrated sensing and communication (ISAC) is regarded as a key enabling technique in future 6th-generation (6G) mobile communication systems. However, existing multi-input multi-output (MIMO) orthogonal frequency division multiplexing (OFDM) ISAC designs generally rely on the fixed-position antennas and fixed allocation of time-frequency resources, thereby limiting the degrees of freedom of wireless sensing along the spatial-temporal-spectral dimensions. In this paper, we propose a novel wireless sensing framework for MIMO-OFDM ISAC systems with flexible spatial-temporal-spectral optimization and propose a tensor decomposition-based approach to estimate target parameters, including azimuth/elevation angles, ranges, and velocities. Specifically, we first establish a monostatic wireless sensing model for MIMO-OFDM ISAC systems, where the positions of antenna elements, the allocation of OFDM symbols and subcarriers can be flexibly configured. Then, we formulate the problem of estimating target parameters as a tensor decomposition problem admitting to the canonical polyadic format, which enables the parallel target parameters estimation process from corresponding factor matrices along the spatial, temporal, and spectral dimensions, respectively. Based on the decomposed factor matrices, we derive the Cramer-Rao Bound (CRB) for the unknown target parameters and reveal that the estimation accuracy of azimuth/elevation angles, velocities and ranges is fundamentally determined by the array geometry, the distribution of OFDM symbols and subcarriers. Building on this insight, we obtain an optimized solution for the positions of antenna elements, and optimal solutions for the subcarrier allocation and OFDM symbol allocation to minimize the CRB, as well as the mean square error of target parameters estimation.

eess.SP

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $\pi_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.

cs.RO

LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset

Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making. Existing optimization-based, Machine Learning (ML), and Reinforcement Learning (RL) approaches often rely on predefined models or task-specific training, limiting their generalization and adaptability in uncertain scenarios. Recent Large Language Model (LLM)-assisted approaches offer promising reasoning capabilities but remain constrained by limited agentic functionality, including insufficient memory, planning, and tool interaction mechanisms.This paper proposes an LLM-Agent-Based Path Finder (LAPF) framework for autonomous UAV navigation in town-scale outdoor environments. LAPF extends LLM-assisted navigation by integrating perception, memory, planning, and action modules into a closed-loop cognitive architecture. The proposed agent leverages prior navigation experiences, performs Chain-of-Thought (CoT) reasoning, couples each detected hazard to a bounded corrective action, and dynamically refines waypoint decisions based on environmental feedback.The three independent trials per method demonstrate that LAPF achieves mean path lengths of 512.83 m and 506.37 m, compared to the straight-line optimum of 497.33 m, corresponding to path length reductions of 17.2% and 15.6% relative to CoT prompting and absolute path efficiencies of 97.1% and 98.1% in open-field and obstacle-injected scenarios, respectively. Furthermore, LAPF is the only evaluated approach that couples every detected hazard to a bounded, metric-neutral corrective action while maintaining near-goal stability, with zero clamp events in both scenarios, whereas CoT prompting increases from 9.7 to 14.0 events.

cs.RO

Trigger the Straggler: Load Hijack on Mixture-of-Experts LLMs

Expert parallelism (EP) is a common strategy for serving large Mixture-of-Experts (MoE) models across multiple GPUs by distributing experts among devices. Router decisions then determine both which experts process each token and which GPUs execute the resulting work. This procedure exposes a supply-chain attack surface in the serving schedule. We introduce Load Hijack, in which a malicious model provider modifies only a checkpoint's router weights, distributes the poisoned checkpoint, and retains a private trigger. When the trigger appears, the poisoned router concentrates token-to-expert assignments on experts co-located on one GPU. The resulting load makes that GPU a straggler and forces peer devices to wait, while routing on ordinary inputs remains near the clean reference. We find this conditional behavior difficult to achieve because an objective that rewards target-expert use on triggered inputs can also bias ordinary-input routing toward the same experts. To resolve this conflict, Load Hijack employs a three-stage optimization procedure that produces strong trigger-dependent concentration while keeping ordinary-input routing close to the clean reference. Across three MoE families and four corpora, Load Hijack directs 92.3% to 95.6% of triggered token assignments to the target experts. In live EP serving, triggered traffic produces 1.43x the time-to-first-token and 0.86x the throughput measured under ordinary traffic. These results show that poisoned routers can act as trigger-controlled device schedulers and motivate checkpoint audits of routing and runtime load.

cs.CR

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and hypothesis generation. Traditional symbolic systems such as SemRep provide high precision but limited recall, while large language models (LLMs) offer stronger contextual reasoning but remain prone to false-positive predictions. We developed ANCHOR-RE, a framework that integrates ontology-guided reasoning, external knowledge grounding, and data-driven verification rules into LLM inference. We evaluated it on three BioRE benchmarks (SemRepGS, DDI, and ChemProt) using both proprietary and open-weight LLMs. To assess generalizability beyond benchmark datasets while reducing potential evaluation bias from LLM pretraining contamination, we conducted a temporal evaluation using 100 biomedical articles published in 2026. With the proprietary backbone, ANCHOR-RE outperformed direct LLM prompting, improving micro-F1 from 0.654 to 0.676 on SemRepGS, from 0.769 to 0.872 on DDI, and from 0.939 to 0.941 on ChemProt. On DDI and ChemProt, it also outperformed previously reported inference-only methods and approached fine-tuned or instruction-tuned systems without parameter updates. Similar performance gains observed with open-weight LLMs indicate that the benefits were not limited to the proprietary backbone. On the post-cutoff set, manual assessment of 500 randomly sampled predictions yielded a precision of 69%, maintaining consistent precision on previously unseen biomedical literature. Neuro-symbolic reasoning can improve the reliability of LLM-based BioRE without fine-tuning. Results across multiple benchmarks, model families, and post-cutoff literature support ANCHOR-RE as a practical training-free approach to biomedical literature mining.

cs.CL

Large language models for partial differential equation workflows

Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows that connect modelling assumptions, governing equations, numerical solvers, diagnostics, and decisions. Large language models (LLMs) are beginning to support such workflows by linking natural language, symbolic mathematics, code, solver outputs, and feedback. Here we examine recent advances in LLM-assisted PDE research across three stages: the discovery and formulation of governing models, the generation and revision of executable numerical solvers, and the use of simulation feedback to support control, design, and optimization. Across these stages, current systems act primarily as workflow-level interfaces. Despite this progress, the field remains limited by the scarcity of high-quality datasets and benchmarks, especially for knowledge discovery and real-world applications, where expert annotation, executable problem construction, and task-level feedback require substantial domain effort. A further challenge is the persistent gap between simulation-based results and real-world scientific and engineering systems, which limits the direct transfer of numerical simulations, control policies, and optimized designs to practical settings. These challenges make LLM-assisted PDE workflows a critical testbed for developing scientific AI systems that can connect language, computation, physical constraints, and real-world decision-making.

cs.AI