SearcharxivSearch

arXiv subjects

Zuo Wang

Publications and source records attributed to Zuo Wang.

At least 19 recordsLinked to original sources

Aspire: Can Models Self-Evolve from Vague Goals?

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.

cs.CL

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

cs.AI

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads. To separate compliance from coincidence we introduce Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1-85.9% and AP-Acc 66.1-78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.

cs.AI

DeepPropNet: an operator learning-based predictor for thermal plasma properties

Thermal plasma properties play a critical role in plasma simulations and plasma-related applications. However, their strong nonlinear dependence on temperature, pressure, and gas composition makes accurate and efficient evaluation challenging. In this work, an operator learning-based model, termed DeepPropNet, is proposed for fast prediction of thermodynamic and transport properties of thermal plasmas. Two architectures are developed, including a single-property model (S-DeepPropNet) and a Mixture of Experts (MoE)-based multi-property model (MoE-DeepPropNet). The proposed models learn the nonlinear mapping from plasma operating conditions to physical properties based on high-fidelity datasets. The MoE architecture enables efficient multi-property prediction within a unified framework. Predictions are performed for binary SF6-N2 and ternary C4F7N-CO2-O2 mixtures. The results show that the proposed models achieve high accuracy, with relative L2 errors on the order of 10-3 to 10-2, while maintaining strong generalization capability under unseen conditions. The applicability of DeepPropNet is further demonstrated by coupling with finite volume method (FVM) and physics-informed neural networks (PINNs). The results indicate that DeepPropNet provides an efficient and scalable approach for plasma property prediction and plasma simulations.

physics.plasm-ph

Emergent Macroscopic Nonreciprocity from Identical Active Particles via Spontaneous Symmetry Breaking

Nonreciprocity is known to generate a wide range of exotic phenomena in multi-species many-body systems, where different species influence one another through couplings that violate Newton's third law. In contrast, in the absence of explicitly imposed macroscopic nonreciprocal processes, single-species nonreciprocity -- another distinct form of nonreciprocity -- typically plays only a limited role in shaping macroscopic physics. Here, using a single-species Vicsek model with a vision cone and extrinsic noise, we show that spontaneous symmetry breaking (SSB) can dramatically enhance the macroscopic consequences of microscopic single-species nonreciprocity. In the ordered phase, this enhancement gives rise to an emergent macroscopic nonreciprocity that induces the system of identical active particles to admit an effective description with a "two-species" non-Hermitian structure. The resulting SSB-enhanced nonreciprocity substantially promotes traveling-band formation and, more strikingly, drives a novel real-space condensation of identical active particles, characterized by a "traveling line" with vanishing longitudinal width. Our findings uncover a fundamental mechanism by which microscopic single-species nonreciprocity can exert strong macroscopic influences in complex systems.

cond-mat.stat-mech

Black-Hole Signatures in the Finite-Temperature Critical Ising Chain

We demonstrate that the finite-temperature critical transverse-field Ising chain exhibits quantitative signatures of black-hole physics in its dual gravitational description within the AdS/CFT correspondence. Its finite-temperature dynamics and thermodynamics are consistently captured by a mixed thermal-AdS/BTZ black hole saddle, leading to three mutually compatible observations. First, antipodal excitation transport collapses onto a universal temperature-dependent curve determined by the relative AdS and BTZ contributions to the gravitational partition function, reflecting horizon absorption. Second, in the high-temperature regime, the retarded response exhibits exponential relaxation governed by the lowest quasi-normal mode of the dual black hole. Third, the temperature derivative of the von Neumann entropy develops a pronounced minimum at a temperature consistent with the Hawking-Page transition. These results identify critical quantum spin chains as minimal and experimentally accessible platforms for probing dynamical and thermodynamic aspects of quantum black holes in controllable many-body systems.

quant-ph

Suppression of Decoherence at Exceptional Transitions

Decoherence is strongly influenced by environmental criticality, with conventional Hermitian critical points typically enhancing the loss of quantum coherence. Here, we show that this paradigm is fundamentally altered in non-Hermitian environments. Focusing on qubits coupled to non-Hermitian spin chains and interacting ultracold Fermi gases, we find that approaching exceptional points can either enhance or strongly suppress decoherence, depending on the balance between Hermitian and non-Hermitian system-environment couplings. In particular, when these couplings are comparable, decoherence is dramatically suppressed at exceptional transitions. We trace this behavior to the distinct response of the environmental ground state near non-Hermitian degeneracies and demonstrate the robustness of this effect across multiple models. Finally, we show that the predicted suppression of decoherence is directly observable on current digital quantum simulation platforms. Our results establish exceptional points as a concrete mechanism for suppressing decoherence and identify non-Hermitian criticality as a new avenue for coherence control in open quantum systems and quantum technologies.

quant-ph

Jensen-Shannon Divergence Message-Passing for Rich-Text Graph Representation Learning

In this paper, we investigate how the widely existing contextual and structural divergence may influence the representation learning in rich-text graphs. To this end, we propose Jensen-Shannon Divergence Message-Passing (JSDMP), a new learning paradigm for rich-text graph representation learning. Besides considering similarity regarding structure and text, JSDMP further captures their corresponding dissimilarity by Jensen-Shannon divergence. Similarity and dissimilarity are then jointly used to compute new message weights among text nodes, thus enabling representations to learn with contextual and structural information from truly correlated text nodes. With JSDMP, we propose two novel graph neural networks, namely Divergent message-passing graph convolutional network (DMPGCN) and Divergent message-passing Page-Rank graph neural networks (DMPPRG), for learning representations in rich-text graphs. DMPGCN and DMPPRG have been extensively texted on well-established rich-text datasets and compared with several state-of-the-art baselines. The experimental results show that DMPGCN and DMPPRG can outperform other baselines, demonstrating the effectiveness of the proposed Jensen-Shannon Divergence Message-Passing paradigm

cs.LG

A Novel Graph-Sequence Learning Model for Inductive Text Classification

Text classification plays an important role in various downstream text-related tasks, such as sentiment analysis, fake news detection, and public opinion analysis. Recently, text classification based on Graph Neural Networks (GNNs) has made significant progress due to their strong capabilities of structural relationship learning. However, these approaches still face two major limitations. First, these approaches fail to fully consider the diverse structural information across word pairs, e.g., co-occurrence, syntax, and semantics. Furthermore, they neglect sequence information in the text graph structure information learning module and can not classify texts with new words and relations. In this paper, we propose a Novel Graph-Sequence Learning Model for Inductive Text Classification (TextGSL) to address the previously mentioned issues. More specifically, we construct a single text-level graph for all words in each text and establish different edge types based on the diverse relationships between word pairs. Building upon this, we design an adaptive multi-edge message-passing paradigm to aggregate diverse structural information between word pairs. Additionally, sequential information among text data can be captured by the proposed TextGSL through the incorporation of Transformer layers. Therefore, TextGSL can learn more discriminative text representations. TextGSL has been comprehensively compared with several strong baselines. The experimental results on diverse benchmarking datasets demonstrate that TextGSL outperforms these baselines in terms of accuracy.

cs.CL

Causal Heterogeneous Graph Learning Method for Chronic Obstructive Pulmonary Disease Prediction

Due to the insufficient diagnosis and treatment capabilities at the grassroots level, there are still deficiencies in the early identification and early warning of acute exacerbation of Chronic obstructive pulmonary disease (COPD), often resulting in a high prevalence rate and high burden, but the screening rate is relatively low. In order to gradually improve this situation. In this paper, this study develop a Causal Heterogeneous Graph Representation Learning (CHGRL) method for COPD comorbidity risk prediction method that: a) constructing a heterogeneous Our dataset includes the interaction between patients and diseases; b) A cause-aware heterogeneous graph learning architecture has been constructed, combining causal inference mechanisms with heterogeneous graph learning, which can support heterogeneous graph causal learning for different types of relationships; and c) Incorporate the causal loss function in the model design, and add counterfactual reasoning learning loss and causal regularization loss on the basis of the cross-entropy classification loss. We evaluate our method and compare its performance with strong GNN baselines. Following experimental evaluation, the proposed model demonstrates high detection accuracy.

cs.LG

DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle

Real-world enterprise data intelligence workflows encompass data engineering that turns raw sources into analytical-ready tables and data analysis that convert those tables into decision-oriented insights. We introduce DAComp, a benchmark of 210 tasks that mirrors these complex workflows. Data engineering (DE) tasks require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data analysis (DA) tasks pose open-ended business problems that demand strategic planning, exploratory analysis through iterative coding, interpretation of intermediate results, and the synthesis of actionable recommendations. Engineering tasks are scored through execution-based, multi-metric evaluation. Open-ended tasks are assessed by a reliable, experimentally validated LLM-judge, which is guided by hierarchical, meticulously crafted rubrics. Our experiments reveal that even state-of-the-art agents falter on DAComp. Performance on DE tasks is particularly low, with success rates under 20%, exposing a critical bottleneck in holistic pipeline orchestration, not merely code generation. Scores on DA tasks also average below 40%, highlighting profound deficiencies in open-ended reasoning and demonstrating that engineering and analysis are distinct capabilities. By clearly diagnosing these limitations, DAComp provides a rigorous and realistic testbed to drive the development of truly capable autonomous data agents for enterprise settings. Our data and code are available at https://da-comp.github.io

cs.CL

WideSearch: Benchmarking Agentic Broad Info-Seeking

From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid development of Large Language Models (LLMs), automated search agents powered by LLMs offer a promising solution to liberate humans from this tedious work. However, the capability of these agents to perform such "wide-context" collection reliably and completely remains largely unevaluated due to a lack of suitable benchmarks. To bridge this gap, we introduce WideSearch, a new benchmark engineered to evaluate agent reliability on these large-scale collection tasks. The benchmark features 200 manually curated questions (100 in English, 100 in Chinese) from over 15 diverse domains, grounded in real user queries. Each task requires agents to collect large-scale atomic information, which could be verified one by one objectively, and arrange it into a well-organized output. A rigorous five-stage quality control pipeline ensures the difficulty, completeness, and verifiability of the dataset. We benchmark over 10 state-of-the-art agentic search systems, including single-agent, multi-agent frameworks, and end-to-end commercial systems. Most systems achieve overall success rates near 0\%, with the best performer reaching just 5\%. However, given sufficient time, cross-validation by multiple human testers can achieve a near 100\% success rate. These results demonstrate that present search agents have critical deficiencies in large-scale information seeking, underscoring urgent areas for future research and development in agentic search. Our dataset, evaluation pipeline, and benchmark results have been publicly released at https://widesearch-seed.github.io/

cs.CL

A Conjoint Graph Representation Learning Framework for Hypertension Comorbidity Risk Prediction

The comorbidities of hypertension impose a heavy burden on patients and society. Early identification is necessary to prompt intervention, but it remains a challenging task. This study aims to address this challenge by combining joint graph learning with network analysis. Motivated by this discovery, we develop a Conjoint Graph Representation Learning (CGRL) framework that: a) constructs two networks based on disease coding, including the patient network and the disease difference network. Three comorbidity network features were generated based on the basic difference network to capture the potential relationship between comorbidities and risk diseases; b) incorporates computational structure intervention and learning feature representation, CGRL was developed to predict the risks of diabetes and coronary heart disease in patients; and c) analysis the comorbidity patterns and exploring the pathways of disease progression, the pathological pathogenesis of diabetes and coronary heart disease may be revealed. The results show that the network features extracted based on the difference network are important, and the framework we proposed provides more accurate predictions than other strong models in terms of accuracy.

cs.LG

Stable excitations and holographic transportation in tensor networks of critical spin chains

The AdS/CFT correspondence conjectures a duality between quantum gravity theories in anti-de Sitter (AdS) spacetime and conformal field theories (CFTs) on the boundary. One intriguing aspect of this correspondence is that it offers a pathway to explore quantum gravity through tabletop experiments. Recently, a multi-scale entanglement renormalization ansatz (MERA) model of AdS/CFT that can be implemented using contemporary quantum simulators has been proposed [R. Sahay, M. D. Lukin, and J. Cotler, arXiv:2401.13595 (2024)]. Particularly, local bulk excitations (entitled "hologrons") manifesting attractive interactions given by AdS gravity were found. However, the fundamental question concerning the stability of these identified hologrons is still left open. Here, we address this question and find that hologrons are unstable during dynamic evolution. In searching for stable bulk excitations with attractive interactions, we find they can be constructed by the local primary operators in the boundary CFT. Furthermore, we identify a class of boundary excitations that exhibit the bizarre behavior of "holographic transportation", which can be directly observed on the boundary system implemented in experiments.

quant-ph

Graph Clustering with Cross-View Feature Propagation

Graph clustering is a fundamental and challenging learning task, which is conventionally approached by grouping similar vertices based on edge structure and feature similarity.In contrast to previous methods, in this paper, we investigate how multi-view feature propagation can influence cluster discovery in graph data.To this end, we present Graph Clustering With Cross-View Feature Propagation (GCCFP), a novel method that leverages multi-view feature propagation to enhance cluster identification in graph data.GCCFP employs a unified objective function that utilizes graph topology and multi-view vertex features to determine vertex cluster membership, regularized by a module that supports key latent feature propagation. We derive an iterative algorithm to optimize this function, prove model convergence within a finite number of iterations, and analyze its computational complexity. Our experiments on various real-world graphs demonstrate the superior clustering performance of GCCFP compared to well-established methods, manifesting its effectiveness across different scenarios.

cs.LG

Emergent non-Hermitian conservation laws at exceptional points

Non-Hermitian systems can manifest rich static and dynamical properties at their exceptional points (EPs). Here, we identify yet another class of distinct phenomena that is hinged on EPs, namely, the emergence of a series of non-Hermitian conservation laws. We demonstrate these distinct phenomena concretely in the non-Hermitian Heisenberg chain and formulate a general theory for identifying these emergent non-Hermitian conservation laws at EPs. By establishing a one-to-one correspondence between the constant of motions at EPs and those in corresponding auxiliary Hermitian systems, we trace their physical origin back to the presence of emergent symmetries in the auxiliary systems. Concrete simulations on quantum circuits show that these emergent conserved dynamics can be readily observed in current digital quantum computing systems.

quant-ph

Measurement-induced integer families of critical dynamical scaling in quantum many-body systems

A quantum many-body system can undergo transitions in the presence of continuous measurement. In this work, we find that a generic class of critical dynamical scaling behavior can emerge at these measurement-induced transitions. Remarkably, depending on the symmetry that can be respected by the system, different integer families of dynamical scaling can emerge. The origin of these scaling families can be traced back to the presence of hierarchies of high order exceptional points in the effective non-Hermitian descriptions of the systems. Direct experimental observation of this class of dynamical scaling behavior can be readily achieved using ultracold atoms in optical lattices or through intermediate-scale quantum computing systems.

cond-mat.quant-gas

Real-space condensation of reciprocal active particles driven by spontaneous symmetry breaking induced nonreciprocity

We investigate the steady-state and dynamical properties of a reciprocal many-body system consisting of self-propelled active particles with local alignment interactions that exists within a fan-shaped neighborhood of each particle. We find that the nonreciprocity can emerge in this reciprocal system once the spontaneous symmetry breaking is present, and the effective description of the system assumes a non-Hermitian structure that directly originates from the emergent nonreciprocity. This emergent nonreciprocity can impose strong influences on the properties the system. In particular, it can even drive a real-space condensation of active particles. Our findings pave the way for identifying a new class of physics in reciprocal systems that is driven by the emergent nonreciprocity.

cond-mat.stat-mech