SearcharxivSearch

arXiv subjects

Zhenyu Ma

Publications and source records attributed to Zhenyu Ma.

12 recordsLinked to original sources

Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. Existing benchmarks largely measure whether manipulated content is retrieved or endorsed, but do not track whether an agent verifies suspicious evidence, revises adopted claims, or recovers before producing its final recommendation. We introduce HAE-GEO, a benchmark that tracks the full trajectory from exposure to recovery under progressively more persuasive Web poisoning. Agents interact via a multi-turn Search-Scrape interface across three attack levels (L1 direct assertion, L2 contextual camouflage, and L3 apparent corroboration), supported by a controlled corpus of 72,039 clean pages and 770 poisoned pages per level spanning 8 product categories and 154 brands. Evaluation combines deterministic behavioral measures with six semantic rubric dimensions. Evaluating 10 agents, we find three recurring patterns: evidence recognition degrades under the corroboration trap; agentic search improves final resistance without improving evidence recognition or utility; and defense prompting increases verification, yet rarely converts verification into recovery.

cs.CR

UI-Venus-2 Technical Report

Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.

cs.AI

PRAXIS: Case-distilled and code-verified AI agents for biological research

Large language models are moving scientific research from text assistance toward agentic workflows, yet biological research requires strong object validation, methodological suitability, reproducibility, and auditability. Prompt engineering, general RAG, or tool use alone cannot reliably produce domain-specific scientific judgment. Here, we present PRAXIS, a verifiable biological research agent framework driven by literature learning and case distillation. PRAXIS converts research experience, failure boundaries, domain rules, and executable procedures into structured long-term memory. By coordinating successful cases, negative cases, rules, and skills, PRAXIS supports problem definition, object validation, method selection, workflow execution, result interpretation, and review feedback across diverse biocomputational tasks. We instantiated PRAXIS as an agent suite for biomedical computing and evaluated it through object validation, case retrieval, memory ablation, public benchmarks, and cross-agent workflows. The results show that case-based learning improves method selection, error suppression, and workflow organization in complex biological research tasks. Rather than replacing scientists, PRAXIS provides a general pathway for transforming research experience into executable, auditable, and transferable agent capabilities.

q-bio.QM

MDAgent: A Multi-Agent Framework for End-to-End Molecular Dynamics Research

Molecular dynamics (MD) simulation is a powerful tool for studying biomolecular structural changes, molecular recognition, transmembrane transport, and functional mechanisms. However, its practical bottleneck lies not only in software operation or parameter setup, but in translating experimental questions into executable, interpretable, and reviewable computational workflows. Here, we present MDAgent, a multi-agent system for end-to-end molecular dynamics research. The system integrates problem understanding, literature-guided strategy design, simulation execution, trajectory analysis, mechanistic interpretation, and quality supervision into a unified workflow, enabling agents not only to run simulations but also to generate research-oriented computational plans and analytical reports. We further introduce a case-based learning mechanism based on Skill and Memory, which stores reusable knowledge from prior tasks, including parameter choices, operational rules, analytical logic, and problem-solving pathways, thereby supporting cross-task transfer without retraining the underlying model. Across multiple representative molecular simulation tasks, MDAgent achieved stable end-to-end performance with improved strategic adaptability, interpretability, and generalization. In an independent complex task involving conformational transitions of TMEM16F and XKR8, the system successfully completed system design, simulation, and mechanistic analysis for large membrane proteins. These results show that combining multi-agent collaboration with case-based learning can transform MD agents from workflow automation tools into scientific question-oriented computational research systems, providing a scalable framework for AI-driven automated research.

q-bio.QM

Transferable Expertise for Autonomous Agents via Real-World Case-Based Learning

LLM-based autonomous agents perform well on general reasoning tasks but still struggle to reliably use task structure, key constraints, and prior experience in complex real-world settings. We propose a case-based learning framework that converts experience from past tasks into reusable knowledge assets, allowing agents to transfer prior case experience to new tasks and perform more structured analysis. Unlike methods based mainly on pretrained knowledge or static prompts, our framework emphasizes extracting and reusing task-relevant knowledge, analytical prompts, and operational skills from real cases. We evaluate the method on a unified benchmark of six complex task categories and compare it with Zero-Shot, Few-Shot, Checklist Prompt, and Rule Memory baselines. Results show that our method achieves consistently strong performance across all tasks and matches or outperforms the best baseline in every case, with especially clear gains on more complex tasks. Further analysis shows that the advantage of case-based learning increases with task complexity, and that practical knowledge acquired by one agent can be reused by others. These findings suggest that case-based learning offers a promising path for building professional agents for real-world work.

cs.AI

Direct imaging of quantum interference and Non-Abelian entanglement in Hopfion: an magnetic soliton possess loop-like anyonic properties

This work provides the first experimental elucidation of quantum topological effects in individual hopfions, establishing their potential as building blocks for three-dimensional topological quantum spintronics. The observed Non-Abelian characteristics suggest pathways toward fault-tolerant quantum operations through controlled hopfion braiding in engineered magnetic metamaterials.

cond-mat.mes-hall

A Safety-Oriented Self-Learning Algorithm for Autonomous Driving: Evolution Starting from a Basic Model

Autonomous driving vehicles with self-learning capabilities are expected to evolve in complex environments to improve their ability to cope with different scenarios. However, most self-learning algorithms suffer from low learning efficiency and lacking safety, which limits their applications. This paper proposes a safety-oriented self-learning algorithm for autonomous driving, which focuses on how to achieve evolution from a basic model. Specifically, a basic model based on the transformer encoder is designed to extract and output policy features from a small number of demonstration trajectories. To improve the learning efficiency, a policy mixed approach is developed. The basic model provides initial values to improve exploration efficiency, and the self-learning algorithm enhances the adaptability and generalization of the model, enabling continuous improvement without external intervention. Finally, an actor approximator based on receding horizon optimization is designed considering the constraints of the environmental input to ensure safety. The proposed method is verified in a challenging mixed traffic environment with pedestrians and vehicles. Simulation and real-vehicle test results show that the proposed method can safely and efficiently learn appropriate autonomous driving behaviors. Compared reinforcement learning and behavior cloning methods, it can achieve comprehensive improvement in learning efficiency and performance under the premise of ensuring safety.

cs.RO

Rapid FRD determination for multiplexed fibre systems -- I. The quasi-near field model and its uncertainties

Focal Ratio Degradation (FRD) in fibres is a crucial factor to control in astronomical instruments in order to minimize light loss. As astronomical instrumentation has advanced, the integration of large populations of fibres has become common. However, determining FRD in multiplexed fibre systems has become a challenging and time-consuming task. The Integral Field Unit for the Fiber Arrayed Solar Optical Telescope (FASOT-IFU) represents the most densely arranged fibre-based IFU in a single unit. Due to the close packing of fibres in the V-groove of the slit end, measuring FRD is particularly challenging as the output spots are prone to overlapping with adjacent fibres. In this paper, a novel method based on the quasi-near field model is proposed to enable rapid FRD measurement in highly multiplexed fibre systems like IFUs and multi-object observation systems. The principle and uncertainties associated with the method are investigated. The method's validity is demonstrated by applying it to determine the FRD in FASOT-IFU, with the achieved FRD performance meeting the acceptable requirements of FASOT-IFU, where the output focal ratio primarily falls within the range of 5.0-7.0. The results indicate that the proposed method offers several advantages, including the simultaneous and rapid measurement of FRD in multiple fibres with high accuracy (error smaller than 0.35 in F-ratio). Furthermore, besides FRD, the method exhibits potential for extensive measurements of throughput, scrambling, and spectral analysis.

astro-ph.IM

Preferential bond formation and interstitial/vacancy annihilation rate drive atomic clustering in gallium ion sputtered compound materials

The investigation of chemical reactions during the ion irradiation is a frontier for the study of the ion-material interaction. In order to derive the contribution of bond formation to chemistry of ion produced nanoclusters, the valence electron energy loss spectroscopy (VEELS) was exploited to investigate the Ga$^+$ ion damage in Al$_2$O$_3$, InP and InGaAs, where each target material has been shown to yield different process for altering the clustering of recoil atoms: metallic Ga, metallic In and InGaP clusters in Al$_2$O$_3$, InP and InGaAs respectively. Supporting simulations based on Monte Carlo and crystal orbital Hamiltonianindicate that the chemical constitution of cascade induced nano-precipitates is a result of a competition between interstitial/vacancy consumption rate and preferential bond formation.

cond-mat.mtrl-sci

A modified method for determining the FRD and length properties of optical fibres in astronomy

Focal ratio degradation (FRD) is a major contributor to throughput and light loss in a fibre spectroscopic telescope system. We combine the guided mode theory in geometric optics and a well-known model, power distribution model (PDM), to predict and explain the FRD dependence properties. We present a robust method by modifying the energy distribution method (EDM) with \emph{f-intercept} to control the input condition. This method provides a way to determine the proper position of the fibre end on the focal plane to improve energy utilization and FRD performance, which lifts the relative throughput up to 95\% with variation of output focal ratio less than 2\%. And this method can also help to optimize the arrangement of the position of focal-plane plate to enhance the coupling efficiency in a telescope. To investigate length properties, we modified PDM by introducing a new parameter, focal distance \emph{f}, into the original model to make it available for multi-position measurement system. The results show that the modified model is robust and feasible for measuring the key parameter \emph{d}$_0$ to simulate the transmission characteristics. The output focal ratio in the experiment does not follow the prediction trend but shows an interesting phenomenon that the output focal ratio increases at first to the peak, then decreases and remains stable finally with increasing fibre length longer than 15m, which provides a reference for choosing appropriate length of fibre to improve the FRD performance for the design of the fibre system in a telescope.

astro-ph.IM

DEEM, a versatile platform of FRD measurement for highly multiplexed fibre systems in astronomy

We present a new method of DEEM, the direct energy encircling method, for characterising the performance of fibres in most astronomical spectroscopic applications. It's a versatile platform to measure focal ratio degradation (FRD), throughput, and point spread function (PSF). The principle of DEEM and the relation between the encircled energy (EE) and the spot size were derived and simulated based on the power distribution model (PDM). We analysed the errors of DEEM and pointed out the major error source for better understanding and optimisation. The validation of DEEM has been confirmed by comparing the results with conventional method which shows that DEEM has good robustness with high accuracy in both stable and complex experiment environments. Applications on the integral field unit (IFU) show that the FRD of 50$μ$m core fibre is substandard for the requirement which requires the output focal ratio to be slower than 4.5. The homogeneity of throughput is acceptable and higher than 85 per cent. The prototype IFU of the first generation helps to find out the imperfections to optimise the new design of the next generation based on the staggered structure with 35$μ$m core fibres of $N.A.$=0.12, which can improve the FRD performance. The FRD dependence on wavelength and core size is revealed that higher output focal ratio occurs at shorter wavelengths for large core fibres, which is in agreement with the prediction of PDM. But the dependence of the observed data is weaker than the prediction.

astro-ph.IM

SINAP surface preparation processing for 500MHz superconducting cavity

This paper illustrates the design, fabrication and experiment results of surface preparation system for 500MHz superconducting cavity at Shanghai Institute of Applied Physic (SINAP). The SINAP established a set of clean room, buffered chemical polishing equipment, and high pressure ultra-pure water rinsing facility. The whole surface preparation procedure has been operated successfully and verified by the successful vertical tests of 500MHz single cell superconducting cavity.

physics.acc-ph