SearcharxivSearch

arXiv subjects

Xiaowen Zhang

Publications and source records attributed to Xiaowen Zhang.

At least 19 recordsLinked to original sources

On Perfect Divisibility of Bull-Free Graphs Without Long Paths

A graph $G$ is {\em perfectly divisible} if, for every induced subgraph $H$ of $G$, $V(H)$ can be partitioned into $A$ and $B$ such that $H[A]$ is perfect and $ω(H[B])<ω(H)$. Chudnovsky and Sivaraman [J. Graph Theory \textbf{90} (2019) 54-60] proved that every ($P_5$, bull)-free graph is perfectly divisible, while Chen and Xu [Discrete Appl. Math. \textbf{372} (2025) 298-307] proved the same for ($P_7,C_5$, bull)-free graphs. We extend these results by proving that every ($P_8,C_5$, bull)-free graph is perfectly divisible and that, letting $F$ denote the Grötzsch graph, a ($P_6$, bull)-free graph is perfectly divisible if and only if it is $F$-free.

math.CO

Bidirectional Code Reuse in Software Redesign: An Action Research Study of Static Analyzers

Software redesign preserves functionality while improving quality attributes, but manual reuse of code and tests is costly and error-prone, especially in cross-repository redesigns. Focusing on static analyzers where cross-repository redesign needs often arise, we conduct a bidirectional study of the ongoing Soot/SootUp redesign using an action research methodology that combines empirical investigation with validated open-source contributions. Our study reveals: (1) non-linear migration that necessitates bidirectional reuse, (2) deferred reuse via TODOs, (3) neglected test porting, and (4) residual bug propagation during migrations. We identify tracking corresponding code and tests as the key challenge and address it by retrofitting clone detection to derive code mappings between original and redesigned projects. Guided by semantic reuse patterns derived from our study, we propose the Semantic Alignment Score (SAS), which incorporates semantic cues from preserved identifiers, API documentation, and comments. Evaluations on three redesigned project pairs (Soot/SootUp, FindBugs/SpotBugs, and ANTLR3/ANTLR4) show that SAS improves average F1 for code mapping detection by up to 0.34 on our manually labeled benchmark of 1,805 method pairs, while strong traditional detectors approach or surpass LLM detectors. The code mapping analysis also uncovered ongoing maintenance needs, leading to five issues and 10 pull requests, of which eight have been merged.

cs.SE

OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving

LLM serving systems increasingly disaggregate inference into finer-grained stages, with recent approaches separating attention from FFN or MoE execution during decode. This operator-level disaggregated serving (ODS) can improve hardware matching and enable independent scaling, particularly across heterogeneous devices. However, existing systems fix operator boundaries and lack a unified characterization of when disaggregation reduces serving cost. We present OpWeave, an end-to-end framework for heterogeneous ODS. OpWeave provides an analytical cost model that bounds the gains of homogeneous and heterogeneous ODS over colocated serving. It jointly optimizes operator partitioning and deployment configuration through a regularity-aware planner that keeps the search tractable even for hybrid-attention models. A vLLM-based runtime executes the synthesized plans with flexible operator stages across heterogeneous device groups. In our evaluation, OpWeave reduces serving cost by up to $1.78\times$ on homogeneous and $1.89\times$ on heterogeneous GPU clusters relative to the best feasible baseline, while meeting latency SLOs.

cs.DC

Structure, Coloring, and Perfect Divisibility of $(P_2\cup P_4, C_3)$-Free Graphs

Goedgebeur and Schaudt [J. Graph Theory 87 (2018), 188-207] conjectured that every $4$-vertex-critical $(P_7,C_3)$-free graph belongs to a family of seven explicitly defined graphs. In this paper, we establish a structural theorem for connected $(P_2\cup P_4,C_3)$-free graphs. As a consequence, we prove that the Mycielski-Grötzsch graph is the unique $4$-vertex-critical graph in this class, thereby confirming the conjecture of Goedgebeur and Schaudt for $(P_2\cup P_4,C_3)$-free graphs. Our structural theorem also yields a characterization of the chromatic number of these graphs and an $O(n^4)$-time algorithm for deciding whether an $n$-vertex $(P_2\cup P_4,C_3)$-free graph is $3$-colorable. We further study perfect divisibility in the larger class of $(P_2\cup P_4,\text{bull})$-free graphs. We prove that a $(P_2\cup P_4,\text{bull})$-free graph is perfectly divisible if and only if it is Mycielski-Grötzsch graph-free. This result generalizes the main theorem of Deng and Chang [Graphs Combin. 41 (2025), 63].

math.CO

Understanding Bugs in Modern Agentic Frameworks: A Study of Symptoms, Root Causes, and Triggering Conditions

Modern agentic frameworks such as CrewAI and AutoGen have evolved into complex, autonomous multi-agent systems, introducing reliability challenges that go beyond earlier pipeline-based LLM libraries. However, existing empirical studies focus on earlier LLM libraries or task-level bugs, leaving the unique complexities of these agentic frameworks unexplored. We present a comprehensive study of 409 fixed bugs across five representative agentic frameworks, proposing a five-layer architectural abstraction. Our taxonomy identifies previously unreported symptom categories---Unexpected Execution Sequence, User Configuration Ignored, and Incomplete/Incorrect Trace---and isolates agent-specific root causes including Model-Related Fault, Cognitive Context Mismanagement, and Orchestration Fault. Notably, the model integration layer is the most bug-prone yet receives disproportionately low test inclusion rate during bug fixing (47%), revealing a critical validation gap. Despite varying design paradigms, bug symptoms, root causes, and bug-prone components show substantial cross-framework consistency (JS similarity 0.62--0.88). Finally, we present the first systematic study of bug-triggering conditions, identifying error-prone factor combinations across element configurations, input patterns, and operations, and demonstrate their transferability across frameworks, providing a foundation for test oracle design and cross-framework benchmark.

cs.SE

The ASTRID Simulation at z=0: From Massive Black Holes to Large-scale Structure

We present the $z=0$ results for the cosmological simulation ASTRID. Hosting $2\times 5500^3\approx$ 0.33 trillion particles in a box of $370\, {\rm Mpc}$ per side, ASTRID is one of the largest cosmological hydrodynamic simulations evolved to $z=0$. ASTRID features a large population of massive black holes (MBHs), covering a wide mass range $4\times10^{4}\sim 2\times 10^{11}\ M_{\odot}$. The adopted dynamical friction model provides a relatively accurate description of MBH dynamics, making ASTRID a powerful tool to study MBH growth and mergers in a cosmological context. ASTRID successfully captures the co-evolution of MBHs and their host galaxies, producing $M_{\rm BH}-M_{\star}$ and $M_{\rm BH}-σ$ relations in good agreement with observations. Notably, ASTRID generates scatter in these relations that is more consistent with observations than previous simulations, indicating a more realistic MBH diversity. The galaxy stellar mass function at $z=0$ is generally consistent with observational constraints. When dust attenuation is applied, the galaxy luminosity function also agrees well with observations, and the bimodality in galaxy colors is reproduced as well. ASTRID hosts a large population of massive galaxy groups and clusters: 7 halos have $M_{\rm 200c}>10^{15}\ M_{\odot}$, and 9709 halos have $M_{\rm 200c}>10^{13}\ M_{\odot}$. We quantify the stellar mass content in these halos, and find that the correlations between the stellar and halo mass match well with observational constraints. Finally, we present the $z=0$ power spectra of MBH and galaxies, as well as their bias with respect to the matter power spectrum. We find that MBHs with $M_{\rm BH}\geq 10^{8}\ M_{\odot}$ and galaxies with $M_{\star}\geq 10^{10.5}\ M_{\odot}$ serve as good tracers of large-scale structure.

astro-ph.GA

Skill-Guided Continuation Distillation for GUI Agents

Improving GUI agents typically relies on behavior cloning on expert trajectories. However, as the current policy deviates from the expert policy, it inevitably encounters policy-induced off-trajectory states during closed-loop execution, i.e., states that fall outside the expert trajectories. Since expert trajectories provide no demonstrations for these unseen states, such states receive no effective supervision, leaving the policy unable to select the correct action. To close this supervision gap, we propose Skill-Guided Continuation Distillation (SGCD), an iterative self-improvement framework. SGCD first runs the plain policy without skill guidance for a few steps to reach realistic off-trajectory states. From these states, a skill-guided policy then completes the task and produces successful continuations, which are mixed with expert trajectories to supply supervision over policy-induced off-trajectory states. The skills are extracted from both successful and failed rollouts, consisting of Continuation Plans, Critical Targets, Failure Traps, and Success Criteria. On OSWorld-Verified, SGCD improves the success rate of three base models from the low-30\% range to over 50\%, demonstrating its effectiveness and generality.

cs.AI

WebChallenger: A Reliable and Efficient Generalist Web Agent

Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the repetitive tasks where such agents would be most useful. We argue this gap stems not from insufficient model capability but from agent architectures that fail to replicate three human cognitive advantages: selective attention to relevant page regions, persistent memory of website structure, and procedural fluency with common interaction patterns. We introduce WebChallenger, a web agent framework that addresses each gap through architecture design rather than model scale, built around PageMem: a structured page representation deterministically constructed from the DOM that exposes each page as a hierarchy of semantic sections with short summaries. On this shared substrate we build three mechanisms that mirror the three cognitive advantages: a divide-and-conquer observation pipeline that lets the agent skim section summaries and extract details only from task-relevant regions; a lightweight exploration and memory system that traverses each website once to build a reusable map of pages and element behaviors; and compound action workflows that collapse common multi-step interactions into single agent actions, handling partial state changes automatically. Because all three operate over PageMem, the framework generalizes across websites without site-specific adapters. Using off-the-shelf open-weight models without fine-tuning, our system achieves 56.3% on WebArena, 48.7% on VisualWebArena, 51.0% on Online-Mind2Web, and 70.9% on WorkArena, approaching frontier proprietary systems at a fraction of the cost. Our code is released at https://github.com/jayoohwang1/webchallenger

cs.CL

Three-Dimensional Atomic-Scale Structural Transformation in a SrTiO3 Grain Boundary

Grain boundaries (GBs) in complex oxides play critical roles in governing their functional properties, which are intrinsically linked to their three-dimensional (3D) atomic configurations and local chemical environments that can deviate markedly from those of the bulk. However, the 3D atomic structures of GBs remain poorly understood due to the projection limitations of conventional (S)TEM. Here, using multislice electron ptychography, we resolve the 3D atomic structure of a Σ13(510)/[001] tilt GB in SrTiO3 with simultaneous visualization of both cation and oxygen columns. Depth-resolved reconstruction reveals pronounced structural inhomogeneity along the GB, uncovering a transition from the canonical symmetric configuration (STR1) to an asymmetric configuration (STR2) that is hidden in conventional projection imaging. Quantitative analysis of atomic-column intensities demonstrates that these two GB configurations possess distinct local chemical and vacancy distributions. By further mapping the atomic displacement fields, we reveal that the transformation between STR1 and STR2 proceeds via local atomic shuffling at the GB core and collective shear displacement in the adjoining grains, mediated by the step and dislocation character of the junction, respectively. Moreover, analysis of oxygen octahedral rotations reveals a strong dependence on the local atomic structure with pronounced asymmetry around the STR2 region. These findings establish a direct link among the 3D atomic structure, local chemical composition, and lattice order parameters at the GB, underscoring the critical importance of depth-resolved characterization in understanding and engineering GB-mediated functionalities in complex oxides.

cond-mat.mtrl-sci

Physical design of cold neutron direct geometry inelastic spectrometer at China Spallation Neutron Source

The Cold-Neutron Inelastic Spectrometer (CNIS) is a direct-geometry, time-of-flight instrument designed for China Spallation Neutron Source (CSNS) and optimized to probe low-energy lattice and magnetic excitations. The instrument integrates a long flight path with bent supermirror guides and an elliptical-focusing geometry to suppress high-energy background while improving cold-neutron delivery to the sample. A flexible multi-disk chopper suite provides pulse shaping, band selection and monochromatization, enabling multi-$E_\textrm{i}$ operation. Modular features, including an interchangeable high-focusing guide insert, radial collimation and a vacuum ``airbox'' for simplified sample-environment integration, enhance signal-to-noise and operational versatility. Through combined flight-path and chopper optimization, CNIS achieves excellent routine-mode energy resolution and can reach approximately $\sim 1\%$ in a dedicated high-resolution configuration. CNIS is planned to commence user operation in 2029, offering a highly flexible platform for cold-neutron inelastic scattering studies.

physics.app-ph

HyMem: Hybrid Memory Architecture with Dynamic Retrieval Scheduling

Large language model (LLM) agents demonstrate strong performance in short-text contexts but often underperform in extended dialogues due to inefficient memory management. Existing approaches face a fundamental trade-off between efficiency and effectiveness: memory compression risks losing critical details required for complex reasoning, while retaining raw text introduces unnecessary computational overhead for simple queries. The crux lies in the limitations of monolithic memory representations and static retrieval mechanisms, which fail to emulate the flexible and proactive memory scheduling capabilities observed in humans, thus struggling to adapt to diverse problem scenarios. Inspired by the principle of cognitive economy, we propose HyMem, a hybrid memory architecture that enables dynamic on-demand scheduling through multi-granular memory representations. HyMem adopts a dual-granular storage scheme paired with a dynamic two-tier retrieval system: a lightweight module constructs summary-level context for efficient response generation, while an LLM-based deep module is selectively activated only for complex queries, augmented by a reflection mechanism for iterative reasoning refinement. Experiments show that HyMem achieves strong performance on both the LOCOMO and LongMemEval benchmarks, outperforming full-context while reducing computational cost by 92.6\%, establishing a state-of-the-art balance between efficiency and performance in long-term memory management.

cs.AI

FlowForge: A Staged Local Rollout Engine for Flow-Field Prediction

Deep learning surrogates for CFD flow-field prediction often rely on large, complex models, which can be slow and fragile when data are noisy or incomplete. We introduce FlowForge, a staged local rollout engine that predicts future flow fields by compiling a locality-preserving update schedule and executing it with a shared lightweight local predictor. Rather than producing the next frame in a single global pass, FlowForge rewrites spatial sites stage by stage so that each update conditions only on bounded local context exposed by earlier stages. This compile-execute design aligns inference with short-range physical dependence, keeps latency predictable, and limits error amplification from global mixing. Across PDEBench, CFDBench, and BubbleML, FlowForge matches or improves upon strong baselines in pointwise accuracy, delivers consistently better robustness to noise and missing observations, and maintains stable multi-step rollout behavior while reducing per-step latency.

cs.LG

S3CDM: A secret-sharing-scheme-based cyberattack detection model and its simulation implementation

We design and develop a secret-sharing-scheme-based cyberattack detection model(S3CDM)that can detect unauthorized or illegal activities (especially insider attacks) and protect sensitive information within complex network infrastructures of large organizations. The model splits a secret among a group of legitimate participants or components for authentication, integration and detection of unauthorized activities. Traditional Shamir's polynomial interpolation based and our own hash function based schemes are utilized in the model, they both are practical and efficient to make sure the communications between different components are secure and any unauthorized activities can be detected. The model offers a flexible multi-factor authentication method to enhance the overall system security. Probability analysis [3] shows that multiple component model is more resistant against cyberattacks than the single component one. To demonstrate the feasibility, we implement the S3CDM in three parts on Google Cloud Platform, i.e., the front end UI (User Interface) running on an HTTP server, the back end individual services written in Python, and a PostgreSQL database. Docker is used to manage the start and stop of individual services and their URLs. We demonstrate how to use the UI with a use case of simulation of broken path in details.

cs.CR

AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process

Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text reasoning to enhance both effectiveness and efficiency. However, existing evaluations rely on static difficulty labels and simplistic metrics, which fail to capture the dynamic nature of difficulty relative to varying model capacities. Consequently, they obscure the distinction between adaptive mode selection and general performance while neglecting fine-grained process analyses. In this paper, we propose AdaptMMBench, a comprehensive benchmark for adaptive multimodal reasoning across five domains: real-world, OCR, GUI, knowledge, and math, encompassing both direct perception and complex reasoning tasks. AdaptMMBench utilizes a Matthews Correlation Coefficient (MCC) metric to evaluate the selection rationality of different reasoning modes, isolating this meta-cognition ability by dynamically identifying task difficulties based on models' capability boundaries. Moreover, AdaptMMBench facilitates multi-dimensional process evaluation across key step coverage, tool effectiveness, and computational efficiency. Our evaluation reveals that while adaptive mode selection scales with model capacity, it notably decouples from final accuracy. Conversely, key step coverage aligns with performance, though tool effectiveness remains highly inconsistent across model architectures.

cs.CV

Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting

Object counting is a fundamental task in computer vision, with broad applicability in many real-world scenarios. Fully-supervised counting methods require costly point-level annotations per object. Few weakly-supervised methods leverage only image-level object counts as supervision and achieve fairly promising results. They are, however, often limited to counting a single category, e.g. person. In this paper, we propose WS-COC, the first MLLM-driven weakly-supervised framework for class-agnostic object counting. Instead of directly fine-tuning MLLMs to predict object counts, which can be challenging due to the modality gap, we incorporate three simple yet effective strategies to bootstrap the counting paradigm in both training and testing: First, a divide-and-discern dialogue tuning strategy is proposed to guide the MLLM to determine whether the object count falls within a specific range and progressively break down the range through multi-round dialogue. Second, a compare-and-rank count optimization strategy is introduced to train the MLLM to optimize the relative ranking of multiple images according to their object counts. Third, a global-and-local counting enhancement strategy aggregates and fuses local and global count predictions to improve counting performance in dense scenes. Extensive experiments on FSC-147, CARPK, PUCPR+, and ShanghaiTech show that WS-COC matches or even surpasses many state-of-art fully-supervised methods while significantly reducing annotation costs. Code is available at https://github.com/viscom-tongji/WS-COC.

cs.CV

STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning

In vision-language models (VLMs), misalignment between textual descriptions and visual coordinates often induces hallucinations. This issue becomes particularly severe in dense prediction tasks such as spatial-temporal video grounding (STVG). Prior approaches typically focus on enhancing visual-textual alignment or attaching auxiliary decoders. However, these strategies inevitably introduce additional trainable modules, leading to significant annotation costs and computational overhead. In this work, we propose a novel visual prompting paradigm that avoids the difficult problem of aligning coordinates across modalities. Specifically, we reformulate per-frame coordinate prediction as a compact instance-level identification problem by assigning each object a unique, temporally consistent ID. These IDs are embedded into the video as visual prompts, providing explicit and interpretable inputs to the VLMs. Furthermore, we introduce STVG-R1, the first reinforcement learning framework for STVG, which employs a task-driven reward to jointly optimize temporal accuracy, spatial consistency, and structural format regularization. Extensive experiments on six benchmarks demonstrate the effectiveness of our approach. STVG-R1 surpasses the baseline Qwen2.5-VL-7B by a remarkable margin of 20.9% on m_IoU on the HCSTVG-v2 benchmark, establishing a new state of the art (SOTA). Surprisingly, STVG-R1 also exhibits strong zero-shot generalization to multi-object referring video object segmentation tasks, achieving a SOTA 47.3% J&F on MeViS.

cs.CV

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

Vision language models (VLMs) have achieved impressive performance across a variety of computer vision tasks. However, the multimodal reasoning capability has not been fully explored in existing models. In this paper, we propose a Chain-of-Focus (CoF) method that allows VLMs to perform adaptive focusing and zooming in on key image regions based on obtained visual cues and the given questions, achieving efficient multimodal reasoning. To enable this CoF capability, we present a two-stage training pipeline, including supervised fine-tuning (SFT) and reinforcement learning (RL). In the SFT stage, we construct the MM-CoF dataset, comprising 3K samples derived from a visual agent designed to adaptively identify key regions to solve visual tasks with different image resolutions and questions. We use MM-CoF to fine-tune the Qwen2.5-VL model for cold start. In the RL stage, we leverage the outcome accuracies and formats as rewards to update the Qwen2.5-VL model, enabling further refining the search and reasoning strategy of models without human priors. Our model achieves significant improvements on multiple benchmarks. On the V* benchmark that requires strong visual reasoning capability, our model outperforms existing VLMs by 5% among 8 image resolutions ranging from 224 to 4K, demonstrating the effectiveness of the proposed CoF method and facilitating the more efficient deployment of VLMs in practical applications.

cs.CV

Longwave-transparent low-emissivity material

Low emissivity (low-e) materials are crucial for conserving thermal energy in buildings, cold chain logistics and transportation by minimizing unwanted radiative heat loss or gain. However, their metallic nature intrinsically causes severe longwave attenuation, hindering their broad applications. Here, we introduce, for the first time, an all-dielectric longwave-transparent low-emissivity material (LLM) with ultra-broadband, high transmittance spanning 9 orders of magnitude, from terahertz to kilohertz frequencies. This meter-scale LLM not only achieves energy savings of up to 41.1% over commercial white paint and 10.2% over traditional low-e materials, but also unlocks various fundamentally new capabilities including high-speed wireless communication in energy-efficient buildings, wireless energy transfer with radiative thermal insulation, as well as non-invasive terahertz security screening and radio frequency identification in cold chain logistics. Our approach represents a new photonic solution towards carbon neutrality and smart city development, paving the way for a more sustainable and interconnected future.

physics.optics