SearcharxivSearch

arXiv subjects

Zirui Zhao

Publications and source records attributed to Zirui Zhao.

At least 19 recordsLinked to original sources

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the same pixels, while code can inspect and modify that state directly. StateAct is a code-first, multi-agent harness built around this distinction. Its main agent works directly with program state by using code, while a dedicated GUI subagent handles screenshot-and-click interaction on the few subgoals that need it, just 28 of 108 tasks and 1.1% of main-agent steps. The same direct access to program state also supports verification: an independent finish gate double-checks the saved result for structural failures, e.g., output that is missing, unsaved, or written to the wrong path. To stay on track over hundreds of steps, the main agent hands subgoals to fresh subagents, keeping its own context focused. On OSWorld 2.0, StateAct lifts Claude Opus 4.8 from 20.6% to 26.9% on binary success, and from 54.8% to 61.6% on partial success, at ~ 9x lower cost per task than the same model driven by screenshots alone; a code-only variant with no GUI subagent reaches only 45.9% partial, below that screenshot-based baseline's 54.8%. In general, grounding action, verification, and memory in state, what we call state-grounding, shifts the main bottleneck from perception toward reasoning: failures depend more on what the agent thinks than on what it sees.

cs.SE

GPA: Learning GUI Process Automation from Demonstrations

GUI Process Automation (GPA) is a lightweight but general vision-based Robotic Process Automation (RPA), which enables fast and stable process replay with only a single demo. Addressing the fragility of traditional RPA and the non-deterministic risks of current vision language model-based GUI agents, GPA introduces three core benefits: (1) Robustness via Sequential Monte Carlo-based localization to handle rescaling and detection uncertainty; (2) Deterministic and Reliability safeguarded by readiness calibration; and (3) Privacy through fast, fully local execution. This approach delivers the adaptability, robustness, and security required for enterprise workflows. It can also be used as an MCP/CLI tool by other agents with coding capabilities so that the agent only reasons and orchestrates while GPA handles the GUI execution. We conducted a pilot experiment to compare GPA with Gemini 3 Pro (with CUA tools) and found that GPA achieves higher success rate with 10 times faster execution speed in finishing long-horizon GUI tasks.

cs.CV

When Diffusion Breaks Constraints: Sequential Autoregressive Generation with RL and MCTS

Data-driven generative models excel in language and vision, but diffusion models often fail in constrained planning and design tasks, exhibiting severe constraint violations in engineering inverse design, molecular generation, multi-robot planning, and floorplan/scene synthesis even with projection or guidance. Such tasks combine hard-to-specify semantic goals with strict geometric or physical constraints (e.g., non-overlap, connectivity), yielding feasible solutions that lie on low-dimensional, small, and sometimes disconnected regions of the output space. This paper studies the failure mode through tangram generation from language, where seven fixed shapes must form a text-described silhouette while remaining connected and non-overlapping, and a simplified rectangle composition task with a learned bounding-box constraint. We find diffusion models struggle to satisfy constraints, consistent with difficulty generating samples near low-dimensional submanifolds. Motivated by locally feasible reparameterizations, we reformulate constrained generation as discrete autoregressive sequential generation. Reinforcement learning improves feasibility and task success, and Monte Carlo tree search quantifies the value of look-ahead when feasible regions shrink. Overall, the empirical, theoretical, and prior-work evidence points to a structural limitation of continuous density matching on this class of constrained-generation problems, and suggests sequential constraint-aware generation as a promising alternative.

cs.CV

Synergistic effects of rare-earth doping on the magnetic properties of orthochromates: A machine learning approach

Multiferroic materials, particularly rare-earth orthochromates (RECrO$_3$), have garnered significant interest due to their unique magnetic and electric-polar properties, making them promising candidates for multifunctional devices. Although extensive research has been conducted on their antiferromagnetic (AFM) transition temperature (N$\acute{\textrm{e}}$el temperature, $T_\textrm{N}$), ferroelectricity, and piezoelectricity, the effects of doping and substitution of rare-earth (RE) elements on these properties remain insufficiently explored. In this study, convolutional neural networks (CNNs) were employed to predict and analyze the physical properties of RECrO$_3$ compounds under various doping scenarios. Experimental and literature data were integrated to train machine learning models, enabling accurate predictions of $T_\textrm{N}$, besides remanent polarization ($P_\textrm{r}$) and piezoelectric coefficients ($d_{33}$). The results indicate that doping with specific RE elements significantly impacts $T_\textrm{N}$, with optimal doping levels identified for enhanced performance. Furthermore, high-entropy RECrO$_3$ compounds were systematically analyzed, demonstrating how the inclusion of multiple RE elements influences magnetic properties. This work establishes a robust framework for predicting and optimizing the properties of RECrO$_3$ materials, offering valuable insights into their potential applications in energy storage and sensor technologies.

cond-mat.mtrl-sci

Topologically-protected superluminal pair annihilation in photonic time crystals

Photonic time crystals (PTCs) - dielectric media whose permittivity is periodically modulated in time - map to a Dirac equation with an imaginary mass, opening a momentum gap (k-gap) where modes grow or decay exponentially. Here, we introduce a sequence of temporal Jackiw-Rebbi kinks that act as a programmable flip of the Dirac mass, exchanging the amplifying and decaying in-gap modes. By launching two seeded pulses with a controlled relative phase, we demonstrate topological pair annihilation in spacetime domain, the phase-selective cancellation of counter-propagating, k-gap-amplified modes. The resulting spatiotemporal cascade appears superluminal, yet causality is preserved because the cascaded pattern carries no net energy flux. To facilitate implementation, we construct a minimal time-varying non-Hermitian lattice model and reproduce the phase-selective pair annihilation behavior, establishing a direct continuum-lattice correspondence. Our results identify topological kinks as temporal gating to manipulate the growth and wave propagation of time-varying media.

physics.optics

Ultrafast Stern-Gerlach and Anomalous Bragg Diffraction Regimes of Low-energy Free Electron Interaction with Light

Recent advances in photon-induced near-field electron microscopy (PINEM) have significantly impacted allied disciplines such as laser-driven accelerators and free electron radiations, collectively fostering the emergence of free-electron quantum optics (FEQO). A central objective of FEQO is to achieve coherent optical control of free electrons, analogous to light manipulation of atoms in atom optics. Motivated by this analogy, we propose an ultrafast Stern-Gerlach (USG) regime for low-energy quantum electron wavepacket (QEW), which crucially incorporates the effects of second-order dispersion inherent to slow electrons. We demonstrate that the USG diffraction induces spectral splitting and shifting of the QEW via a longitudinal electric field gradient, with the two dominant truncated sidebands forming a pseudospin degree of freedom for an effective "two-level" electron. Furthermore, by examining the wave-particle duality of the QEW during light interaction, we identify a dispersion-induced anomalous Bragg diffraction regime. This regime exhibits a distinct spectral pattern, differentiating it from these reported PINEM (Raman-Nath), dielectric laser accelerators (DLA), anomalous PINEM, and Bragg diffraction regimes. Our study provides a comprehensive classification for light-induced diffraction regimes for both swift and slow electrons. These findings underscore the pivotal role of slow-electron dispersion and duality nature in free-electron optics, offering promising avenues for electron wavefunction quantum engineering ultrafast interferometers.

quant-ph

MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

The Model Context Protocol has emerged as a transformative standard for connecting large language models to external data sources and tools, rapidly gaining adoption across major AI providers and development platforms. However, existing benchmarks are overly simplistic and fail to capture real application challenges such as long-horizon reasoning and large, unfamiliar tool spaces. To address this critical gap, we introduce MCP-Universe, the first comprehensive benchmark specifically designed to evaluate LLMs in realistic and hard tasks through interaction with real-world MCP servers. Our benchmark encompasses 6 core domains spanning 11 different MCP servers: Location Navigation, Repository Management, Financial Analysis, 3D Design, Browser Automation, and Web Searching. To ensure rigorous evaluation, we implement execution-based evaluators, including format evaluators for agent format compliance, static evaluators for time-invariant content matching, and dynamic evaluators that automatically retrieve real-time ground truth for temporally sensitive tasks. Through extensive evaluation of leading LLMs, we find that even SOTA models such as GPT-5 (43.72%), Grok-4 (33.33%) and Claude-4.0-Sonnet (29.44%) exhibit significant performance limitations. In addition, our benchmark poses a significant long-context challenge for LLM agents, as the number of input tokens increases rapidly with the number of interaction steps. Moreover, it introduces an unknown-tools challenge, as LLM agents often lack familiarity with the precise usage of the MCP servers. Notably, enterprise-level agents like Cursor cannot achieve better performance than standard ReAct frameworks. Beyond evaluation, we open-source our extensible evaluation framework with UI support, enabling researchers and practitioners to seamlessly integrate new agents and MCP servers while fostering innovation in the rapidly evolving MCP ecosystem.

cs.AI

Generating and Weaving Topological Event Wavepackets in Photonic Spacetime Crystals with Fully Energy-Momentum Gapped

We propose a novel type of topological excitation topological event wavepackets (TEWs) emerging in photonic spacetime crystals (STCs) with spacetime modulated dielectric constants. These TEWs exhibit strong spatiotemporal localization and are topologically protected by a fully opened energy momentum ({\omega}k) gap, within which conventional steady states are absent. We further demonstrate that TEWs are spectrally confined within the {\omega}k-gap, providing a combined measurement for probing the emergence of TEW and the {\omega}k-gap size. Furthermore, we construct a spacetime winding number to elucidate the protection of these events. Unlike previously reported nolinearity-induced event solitons, TEWs originate from topological configuration for linear media, thereby more accessible and versatile for experimental realization. Moreover, we show that TEWs can be periodically woven to form an event lattice, enabling to suppress unwanted noise amplification. Our findings open a new pathway toward topological control in photonic spacetime-modulated systems, enabling the {\omega}k-gap band enginering for wave manipulation ranging from microwave to optical regimes.

physics.optics

GTA1: GUI Test-time Scaling Agent

Graphical user interface (GUI) agents autonomously complete tasks across platforms (\eg, Linux) by sequentially decomposing user instructions into action proposals that iteratively interact with visual elements in the evolving environment. However, two main challenges arise: i) planning (\ie, the action proposal sequence) under expansive action space, where selecting an appropriate plan is non-trivial, as many valid ones may exist; ii) accurately grounding actions in complex and high-resolution interfaces, \ie, precisely interacting with visual targets. This paper investigates the aforementioned challenges with our \textbf{G}UI \textbf{T}est-time Scaling \textbf{A}gent, namely GTA1. First, we conduct test-time scaling to select the most appropriate action proposal: at each step, multiple candidate proposals are sampled and evaluated and selected by a judge model. It trades off computation for better decision quality by concurrent sampling. Second, we propose a model that improves grounding of the selected action proposals to its corresponding visual elements. Our key insight is that reinforcement learning (RL) facilitates grounding through inherent objective alignments, rewarding successful clicks on interface elements. Experimentally, GTA1 achieves state-of-the-art performance on both grounding and agent task execution benchmarks. The code and models are released here.

cs.AI

Integrating Machine Learning with Triboelectric Nanogenerators: Optimizing Electrode Materials and Doping Strategies for Intelligent Energy Harves

The integration of machine learning techniques with triboelectric nanogenerators (TENGs) offers a transformative pathway for optimizing energy harvesting technologies. In this study, we propose a comprehensive framework that utilizes graph neural networks to predict and enhance the performance of TENG electrode materials and doping strategies. By leveraging an extensive dataset of experimental and computational results, the model effectively classifies electrode materials, predicts optimal doping ratios, and establishes robust structure-property relationships. Key findings include a 65.7% increase in energy density for aluminum-doped PTFE and an 85.7% improvement for fluorine-doped PTFE, highlighting the critical influence of doping materials and their concentrations. The model further identifies PTFE as a highly effective negative electrode material, achieving a maximum energy density of 1.12 J/cm$^2$ with 7% silver (Ag) doping when copper (Cu) is used as the positive electrode. This data-driven approach not only accelerates material discovery but also significantly reduces experimental costs, providing novel insights into the fundamental factors influencing TENG performance. The proposed methodology establishes a robust platform for intelligent material design, advancing the development of sustainable energy technologies and self-powered systems.

cond-mat.mtrl-sci

Insights into dendritic growth mechanisms in batteries: A combined machine learning and computational study

In recent years, researchers have increasingly sought batteries as an efficient and cost-effective solution for energy storage and supply, owing to their high energy density, low cost, and environmental resilience. However, the issue of dendrite growth has emerged as a significant obstacle in battery development. Excessive dendrite growth during charging and discharging processes can lead to battery short-circuiting, degradation of electrochemical performance, reduced cycle life, and abnormal exothermic events. Consequently, understanding the dendrite growth process has become a key challenge for researchers. In this study, we investigated dendrite growth mechanisms in batteries using a combined machine learning approach, specifically a two-dimensional artificial convolutional neural network (CNN) model, along with computational methods. We developed two distinct computer models to predict dendrite growth in batteries. The CNN-1 model employs standard convolutional neural network techniques for dendritic growth prediction, while CNN-2 integrates additional physical parameters to enhance model robustness. Our results demonstrate that CNN-2 significantly enhances prediction accuracy, offering deeper insights into the impact of physical factors on dendritic growth. This improved model effectively captures the dynamic nature of dendrite formation, exhibiting high accuracy and sensitivity. These findings contribute to the advancement of safer and more reliable energy storage systems.

physics.comp-ph

Grounding Emotional Descriptions to Electrovibration Haptic Signals

Designing and displaying haptic signals with sensory and emotional attributes can improve the user experience in various applications. Free-form user language provides rich sensory and emotional information for haptic design (e.g., ``This signal feels smooth and exciting''), but little work exists on linking user descriptions to haptic signals (i.e., language grounding). To address this gap, we conducted a study where 12 users described the feel of 32 signals perceived on a surface haptics (i.e., electrovibration) display. We developed a computational pipeline using natural language processing (NLP) techniques, such as GPT-3.5 Turbo and word embedding methods, to extract sensory and emotional keywords and group them into semantic clusters (i.e., concepts). We linked the keyword clusters to haptic signal features (e.g., pulse count) using correlation analysis. The proposed pipeline demonstrates the viability of a computational approach to analyzing haptic experiences. We discuss our future plans for creating a predictive model of haptic experience.

cs.HC

Automatic Curriculum Expert Iteration for Reliable LLM Reasoning

Hallucinations (i.e., generating plausible but inaccurate content) and laziness (i.e. excessive refusals or defaulting to "I don't know") persist as major challenges in LLM reasoning. Current efforts to reduce hallucinations primarily focus on factual errors in knowledge-grounded tasks, often neglecting hallucinations related to faulty reasoning. Meanwhile, some approaches render LLMs overly conservative, limiting their problem-solving capabilities. To mitigate hallucination and laziness in reasoning tasks, we propose Automatic Curriculum Expert Iteration (Auto-CEI) to enhance LLM reasoning and align responses to the model's capabilities--assertively answering within its limits and declining when tasks exceed them. In our method, Expert Iteration explores the reasoning trajectories near the LLM policy, guiding incorrect paths back on track to reduce compounding errors and improve robustness; it also promotes appropriate "I don't know" responses after sufficient reasoning attempts. The curriculum automatically adjusts rewards, incentivizing extended reasoning before acknowledging incapability, thereby pushing the limits of LLM reasoning and aligning its behaviour with these limits. We compare Auto-CEI with various SOTA baselines across logical reasoning, mathematics, and planning tasks, where Auto-CEI achieves superior alignment by effectively balancing assertiveness and conservativeness. The code is available at https://github.com/SalesforceAIResearch/Auto-CEI .

cs.LG

Investigating Material Interface Diffusion Phenomena through Graph Neural Networks in Applied Materials

Understanding and predicting interface diffusion phenomena in materials is crucial for various industrial applications, including semiconductor manufacturing, battery technology, and catalysis. In this study, we propose a novel approach utilizing Graph Neural Networks (GNNs) to investigate and model material interface diffusion. We begin by collecting experimental and simulated data on diffusion coefficients, concentration gradients, and other relevant parameters from diverse material systems. The data are preprocessed, and key features influencing interface diffusion are extracted. Subsequently, we construct a GNN model tailored to the diffusion problem, with a graph representation capturing the atomic structure of materials. The model architecture includes multiple graph convolutional layers for feature aggregation and update, as well as optional graph attention layers to capture complex relationships between atoms. We train and validate the GNN model using the preprocessed data, achieving accurate predictions of diffusion coefficients, diffusion rates, concentration profiles, and potential diffusion pathways. Our approach offers insights into the underlying mechanisms of interface diffusion and provides a valuable tool for optimizing material design and engineering. Additionally, our method offers possible strategies to solve the longstanding problems related to materials interface diffusion.

cond-mat.mtrl-sci

Deep learning-driven evaluation and prediction of ion-doped NASICON materials for enhanced solid-state battery performance

We developed a convolutional neural network (CNN) model capable of predicting the performance of various ion-doped NASICON compounds by leveraging extensive datasets from prior experimental investigation.The model demonstrated high accuracy and efficiency in predicting ionic conductivity and electrochemical properties. Key findings include the successful synthesis and validation of three NASICON materials predicted by the model, with experimental results closely matching the model predictions. This research not only enhances the understanding of ion-doping effects in NASICON materials but also establishes a robust framework for material design and practical applications. It bridges the gap between theoretical predictions and experimental validations.

cond-mat.mtrl-sci

Predicting doping strategies for ternary nickel-cobalt-manganese cathode materials to enhance battery performance using graph neural networks

The exceptional electrochemical performance of lithium-ion batteries has spurred considerable interest in advanced battery technologies, particularly those utilizing ternary nickel-cobalt-manganese (NCM) cathode materials, which are renowned for their robust electrochemical performance and structural stability. Building upon this research, investigators have explored doping additional elements into NCM cathode materials to further enhance their electrochemical performance and structural integrity. However, the multitude of doping strategies available for NCM battery systems presents a challenge in determining the most effective approach. In this study, we elucidate the potential of ternary NCM systems as cathode materials for lithium-ion batteries. We compile a comprehensive database of lithium-ion batteries employing NCM systems from various sources of prior research and develop a corresponding data-driven model utilizing graph neural networks to predict optimal doping strategies. Our aim is to provide insights into the NCM-based battery systems for both fundamental understanding and practical applications.

cond-mat.mtrl-sci

On the Empirical Complexity of Reasoning and Planning in LLMs

Chain-of-thought (CoT), tree-of-thought (ToT), and related techniques work surprisingly well in practice for some complex reasoning tasks with Large Language Models (LLMs), but why? This work seeks the underlying reasons by conducting experimental case studies and linking the performance benefits to well-established sample and computational complexity principles in machine learning. We experimented with 6 reasoning tasks, ranging from grade school math, air travel planning, ..., to Blocksworld. The results suggest that (i) both CoT and ToT benefit significantly from task decomposition, which breaks a complex reasoning task into a sequence of steps with low sample complexity and explicitly outlines the reasoning structure, and (ii) for computationally hard reasoning tasks, the more sophisticated tree structure of ToT outperforms the linear structure of CoT. These findings provide useful guidelines for the use of LLM in solving reasoning tasks in practice.

cs.AI

Seamless Virtual Reality with Integrated Synchronizer and Synthesizer for Autonomous Driving

Virtual reality (VR) is a promising data engine for autonomous driving (AD). However, data fidelity in this paradigm is often degraded by VR inconsistency, for which the existing VR approaches become ineffective, as they ignore the inter-dependency between low-level VR synchronizer designs (i.e., data collector) and high-level VR synthesizer designs (i.e., data processor). This paper presents a seamless virtual reality SVR platform for AD, which mitigates such inconsistency, enabling VR agents to interact with each other in a shared symbiotic world. The crux to SVR is an integrated synchronizer and synthesizer IS2 design, which consists of a drift-aware lidar-inertial synchronizer for VR colocation and a motion-aware deep visual synthesis network for augmented reality image generation. We implement SVR on car-like robots in two sandbox platforms, achieving a cm-level VR colocalization accuracy and 3.2% VR image deviation, thereby avoiding missed collisions or model clippings. Experiments show that the proposed SVR reduces the intervention times, missed turns, and failure rates compared to other benchmarks. The SVR-trained neural network can handle unseen situations in real-world environments, by leveraging its knowledge learnt from the VR space.

cs.RO