Searcharxiv⌕ Search

arXiv subjects

Jingzhe Xu

Publications and source records attributed to Jingzhe Xu.

9 recordsLinked to original sources

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design training and test tasks such that test-time gains can be attributed to training experience, and remain vulnerable to data contamination. We present GDPevo, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it. Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable. GDPevo spans CRM, ERP, finance, healthcare, legal, and data-centric workflows. Its V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group. Full automation enables the pipeline to expand the suite to 240 tasks in 24 groups (V2) within two days, providing a practical response to contamination. Using GDPevo, we evaluate four agents, each comprising a harness and a model, under four supervision types. Self-evolution consistently improves held-out accuracy by up to 16.44 percentage points. But the best evolved agents remain far below the fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized. We publicly release the pipeline, benchmark, and full evaluation results at https://github.com/Prism-Shadow/GDPevo.

cs.AI↗

AgenticData: An Agentic Data Analytics System for Heterogeneous Data

Existing unstructured data analytics systems rely on experts to write code and manage complex analysis workflows, making them both expensive and time-consuming. To address these challenges, we introduce AgenticData, an innovative agentic data analytics system that allows users to simply pose natural language (NL) questions while autonomously analyzing data sources across multiple domains, including both unstructured and structured data. First, AgenticData employs a feedback-driven planning technique that automatically converts an NL query into a semantic plan composed of relational and semantic operators. We propose a multi-agent collaboration strategy by utilizing a data profiling agent for discovering relevant data, a semantic cross-validation agent for iterative optimization based on feedback, and a smart memory agent for maintaining short-term context and long-term knowledge. Second, we propose a semantic optimization model to refine and execute semantic plans effectively. Our system, AgenticData, has been tested using three benchmarks. Experimental results showed that AgenticData achieved superior accuracy on both easy and difficult tasks, significantly outperforming state-of-the-art methods.

cs.DB↗

PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?

Data preparation is a central and time-consuming stage in data analysis workflows. Traditionally, commercial tools have relied on graphical user interfaces (GUIs) to simplify data preparation, allowing users to define transformations through visual operators and workflows. Recent advances in large language models (LLMs) raise the possibility of a paradigm shift toward natural language (NL)-driven data preparation, in which users can specify preparation intents in NL directly. However, it remains unclear how far current LLM-based agents are from this paradigm shift in practice. Existing code generation benchmarks do not capture key characteristics of data preparation, including ambiguous user intents, imperfect real-world data, and the need to translate code into interpretable workflows for validation. To bridge this gap, we present PrepBench, a benchmark designed to evaluate NL-driven data preparation along three core capabilities: interactive disambiguation, prep-code generation, and code-to-workflow translation. We crawl data from the Preppin' Data Challenges, and then extend it into a systematically designed benchmark. The benchmark covers diverse domains, and each task involves 3 to 18 data preparation steps. Nearly half of the tasks require over 100 lines of Python code, and the longest solutions approach 300 lines. Our evaluation shows that, despite recent progress, realizing this paradigm shift remains challenging for state-of-the-art LLMs. PrepBench provides a principled benchmark for measuring this gap and helps identify key challenges toward realizing NL-driven data preparation.

cs.DB↗

Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization

Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noisy labels, facing limitations such as high computational costs, heavy hyperparameter tuning process, and coarse-grained optimization. To address these challenges, we propose a novel two-stage noisy learning framework that enables instance-level optimization through a dynamically weighted loss function, avoiding hyperparameter tuning. To obtain stable and accurate information about noise modeling, we introduce a simple yet effective metric, termed wrong event, which dynamically models the cleanliness and difficulty of individual samples while maintaining computational costs. Our framework first collects wrong event information and builds a strong base model. Then we perform noise-robust training on the base model, using a probabilistic model to handle the wrong event information of samples. Experiments on five synthetic and real-world LNL benchmarks demonstrate our method surpasses state-of-the-art methods in performance, achieves a nearly 75% reduction in computational time and improves model scalability.

cs.LG↗

A Complex-Coefficient Voltage Control for Virtual Synchronous Generators for Dynamic Enhancement and Power-Voltage Decoupling

As electric power systems evolve towards decarbonization, the transition to inverter-based resources (IBRs) presents challenges to grid stability, necessitating innovative control solutions. Virtual synchronous generator (VSG) emerges as a prominent solution. However, conventional VSGs are prone to instability in strong grids, slow voltage regulation, and coupled power-voltage response. To address these issues, this paper introduces an advanced VSG control strategy. A novel analysis of the VSG control dynamics is presented through a second-order closed-loop complex single-input single-output system, employing a vectorized geometrical pole analysis technique for enhanced voltage stability and dynamics. The proposed comprehensive controller design mitigates issues related to control interacted subsynchronous resonance and $dq \leftrightarrow 3ϕ$ transformation-induced voltage-coupled power transients, achieving improved system robustness and simplified control tuning. Key contributions include a two-fold design: optimized voltage transition characteristics through direct pole placement and transient power overshoot correction via a compensator. Validated by simulation and experiments, the findings offer a pragmatic solution for integrating VSG technology into decarbonizing power systems, ensuring reliability and efficiency.

eess.SY↗

The spherical transform of a Schwartz function on the free two step nilpotent lie group

Let $F(n)$ be a connected and simply connected free 2-step nilpotent lie group and $K$ be a compact subgroup of Aut($F(n)$). We say that $(K,F(n))$ is a Gelfand pair when the set of integrable $K$-invariant functions on $F(n)$ forms an abelian algebra under convolution. In this paper, we consider the case when $K=O(n)$. In this case, the Gelfand space $(O(n),F(n)$ is equipped with the Godement-Plancherel measure, and the spherical transform $\land:L_{O(n)}^{2}(F(n))\rightarrow L^{2}(Δ(O(n),F(n)))$ is an isometry. I will prove the Gelfand space $Δ(O(n),F(n))$ is equipped with the Godement-Plancherel measure and the inversion formula. Both of which have something related to its correspond Heisenberg group. The main result in this paper provides a complete characterization of the set $φ_{O(n)} (F(n))^{\wedge}$=$\{\widehat{f}\mid f\in φ_{O(n)} (F(n))\}$ of spherical transforms of $O(n)$-invariant Schwartz functions on $F(n)$. I show that a function $F$ on $Δ(O(n),F(n))$ belongs to $φ_{O(n)} (F(n))^{\wedge}$ if and only if the functions obtained from $F$ via application of certain derivatives and difference operators satisfy decay conditions.

math.RT↗

Spectra for Gelfand pairs associated with the free two step nilpotent lie group

Let $F(n)$ be a connected and simply connected free 2-step nilpotent lie group and $K$ be a compact subgroup of Aut($F(n)$). We say that $(K,F(n))$ is a Gelfand pair when the set of integrable $K$-invariant functions on $F(n)$ forms an abelian algebra under convolution. In this paper, we consider the case when $K=O(n)$. In [1], we know the only possible Galfand pairs for $(K,F(n))$ is $(O(n),F(n))$, $(SO(n),F(n))$. So we just consider the case $(O(n),F(n))$, the other case can be obtained in the similar way.We study the natural topology on $Δ(O(n),F(n))$ given by uniform convergence on compact subsets in $F(n)$. We show $Δ(O(n),F(n))$ is a complete metric space. Our main result gives a necessary and sufficient result for the sequence of the "type 1" bounded $O(n)$-spherical functions uniform convergence to the "type 1" bounded $O(n)$-spherical function on compact sets in $F(n)$. What's more, the "type 1" bounded $O(n)$-spherical functions are dense in $Δ(O(n),F(n))$. Further, we define the Fourier transform according to the "type 2" bounded $O(n)$-spherical functions and gives some basic properties of it.

math.RT↗

A new method to prove the irreducibility of the eigenspace representations for Rn semidirect with a finite pseudo-reflection group

We show that the Eigenspace Representations for $\mathbb{R}^{n}$ semidirect with a finite pseudo-reflection group $K$, which satisfy some generic property are equivalent to the induced representations from $\mathbb{R}^{n}$ to $\mathbb{R}^{n} \rtimes K$, which satisfy the same property by Mackey little group method.And the proof of the equivalence is by using matrix coefficients and invariant theory.As a consequence, these eigenspace representations are irreducible.

math.RT↗