SearcharxivSearch

arXiv subjects

Maosen Zhou

Publications and source records attributed to Maosen Zhou.

3 recordsLinked to original sources

ContextWeave: A Real-World Workflow Benchmark

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.

cs.AI

Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction

The evolution of Large Language Models (LLMs) from passive responders to autonomous agents necessitates a fundamental shift in learning paradigms -- from static imitation to incentive-driven decision making. However, this transition is significantly impeded by the lack of scalable infrastructure capable of constructing high-quality interaction signals for effective policy learning. To address this, we introduce a comprehensive method designed to systematically scale the diversity and complexity of interactive environments. Our method realizes this scaling by addressing three orthogonal dimensions: (1) Complexity: NexAU, a flexible agent framework that supports building complex agent hierarchies via simple configurations; (2) Diversity: NexA4A automatically generates diverse agent hierarchies from natural language to cover infinite domains; and (3) Fidelity: NexGAP bridges the simulation-reality gap by integrating dynamic real-world environment for grounded trajectories synthesis. We train Nex-N1 upon the diverse and complex interactive environments established by our infrastructure. Empirical results on benchmarks such as SWE-bench and tau2 demonstrate that Nex-N1 consistently outperforms SOTA open-source models and achieves competitive performance against frontier proprietary models on complex agentic tasks. We open-source the Nex ecosystem and model weights to facilitate further research.

cs.CL

LHC Search of New Higgs Boson via Resonant Di-Higgs Production with Decays into 4W

Searching for new Higgs particle beyond the observed light Higgs boson h(125GeV) will unambiguously point to new physics beyond the standard model. We study the resonant production of a CP-even heavy Higgs state $H^0$ in the di-Higgs channel via, $gg\to H^0\to h^0h^0\to WW^*WW^*$, at the LHC Run-2 and the high luminosity LHC (HL-LHC). We analyze two types of the $4W$ decay modes, one with the same-sign di-leptons ($4W\to\ell^\pmν\ell^\pmν4q$) and the other with tri-leptons ($4W\to\ell^\pmν\ell^\mpν\ell^\pmν2q$). We perform a full simulation for the signals and backgrounds, and estimate the discovery potential of the heavy Higgs state at the LHC Run-2 and the HL-LHC, in the context of generical two-Higgs-doublet models (2HDM). We determine the viable parameter space of the 2HDM as allowed by the theoretical constraints and the current experimental limits. We systematically analyze the allowed parameter space of the 2HDM which can be effectively probed by the heavy Higgs searches of the LHC, and further compare this with the viable parameter region under the current theoretical and experimental bounds.

hep-ph