SearcharxivSearch

arXiv subjects

Zhou Xu

Publications and source records attributed to Zhou Xu.

At least 19 recordsLinked to original sources

RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning

The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterogeneous capabilities and inference costs make efficiently routing queries a significant challenge. Existing paradigms are inflexible: one-shot routers commit before observing responses, whereas conventional cascades stop adaptively but follow a fixed model order. Cascade routing removes both restrictions by reconsidering whether to stop or invoke another model after each response. Current methods use a predict-then-optimize pipeline estimating response quality and future model utility. However, prediction loss for quality or utility is not equivalent to routing-decision loss. A lower prediction error does not necessarily yield a better action; a small boundary-crossing error can reverse a ``stop'' or model-selection decision. Therefore, we propose RLCascadeRouter, a quality-estimator-free framework that formulates cascade routing as a Markov decision process with actions comprising ``stop'' and model selection. It uses trajectory returns and advantages to directly optimize the performance-cost objective. Its Cascade Policy Network models candidate complementarity for model selection and remaining-action value for stopping, eliminating independent post-hoc response-quality estimators. Evaluated across ten LLMRouterBench benchmarks with thirteen LLMs, RLCascadeRouter outperforms strong baselines and achieves superior performance-cost trade-offs. It incorporates unseen models without retraining, and ablation studies validate both policy components.

cs.AI

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these interfaces as separable capabilities, leaving long-horizon cross-interface orchestration under-tested. Thus, we introduce WeaveBench, a long-horizon hybrid-interface benchmark with 114 tasks across 8 real-world work domains, grounded in real user requests and publicly verifiable artifacts. Each task requires agents to combine GUI observations/actions with CLI/code operations within a single trajectory. We evaluate these tasks on a real Ubuntu desktop inside deployed CLI-agent runtimes, augmented with a minimal desktop-control plugin. We also propose a companion trajectory-aware judge that inspects deliverables, files, screenshots, logs, and action traces, while detecting shortcut behaviors such as fabricated visual evidence or hard-coded metrics. Across frontier model-runtime pairings, the best PassRate reaches only 41.2%, showing the benchmark remains far from saturated. The trajectory-aware judge further reveals that outcome-only grading substantially overestimates agent performance. Overall, WeaveBench exposes a critical gap in CUA evaluation and provides an effective testbed to measure whether agents can orchestrate GUI, CLI, and code operations across long-horizon real-world tasks.

cs.AI

Atomic Structure of Grain Boundaries, Dislocations and Associated Strain in Templated Co-evaporated Photoactive Halide Perovskites

Structural defects, particularly grain boundaries, play a crucial role in governing charge transport and the optoelectronic properties of metal halide perovskites, thereby limiting the performance of devices. Solar cells incorporating templated FA0.9Cs0.1PbI3-xClx show significant improvements in grain orientation and steady-state power conversion efficiency; however, the underlying mechanisms remain unclear. In this study, we address this gap by employing a suite of tailored low-dose electron microscopy techniques to investigate the templated FA0.9Cs0.1PbI3-xClx film, revealing that it exhibits a preferred crystallographic orientation along the <001> zone axis, with arbitrary grain rotations about that axis, indicative of a Volmer-Weber growth mechanism. We determine the atomic structure of the resulting high-angle and low-angle grain boundaries. We also reveal the presence of edge dislocations and their associated strain fields, demonstrating the compressive strain on one side of the dislocation core and tensile strain on the opposite side. Furthermore, we find dislocations associated with stacking faults. These atomic-level insights uncover which grain boundaries and intra-grain defects are likely to act as recombination centres or modify band gaps, crucial for understanding which defects influence the performance of perovskite solar cell devices.

cond-mat.mtrl-sci

ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents

Training-free KV cache compression is essential for deploying vision-language GUI agents under memory and latency constraints, yet existing methods are designed for generic language workloads and ignore the distinctive structure of GUI interaction traces. We characterize three GUI-specific workload properties--high inter-frame visual redundancy, extremely small UI-element spatial footprints, and near-uniform cross-layer attention sparsity--that cause existing schemes to retain as few as 39% of oracle-important KV pairs at the standard 20% budget. To address this, we propose ST-Lite, a training-free compression scheme whose three components each target one property: Trajectory-aware Semantic Gating (TSG) filters redundant historical frames, Component-centric Spatial Saliency (CSS) preserves fine-grained element boundaries, and a flat per-layer budget avoids hierarchical misallocation. Across seven GUI benchmarks and two backbones in the deployment-relevant 10%-40% window, ST-Lite consistently outperforms all existing compression baselines, matching or exceeding Full Cache task accuracy on the primary backbone at the 20% budget while delivering up to 2.35x decoding speedup at fivefold compression. The implementation is available at https://github.com/94wen94/ST-Lite.

cs.CV

Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents

Pure-vision GUI agents provide universal interaction capabilities but suffer from severe efficiency bottlenecks due to the massive spatiotemporal redundancy inherent in high-resolution screenshots and historical trajectories. We identify two critical misalignments in existing compression paradigms: the temporal mismatch, where uniform history encoding diverges from the agent's "fading memory" attention pattern, and the spatial topology conflict, where unstructured pruning compromises the grid integrity required for precise coordinate grounding, inducing spatial hallucinations. To address these challenges, we introduce GUIPruner, a training-free framework tailored for high-resolution GUI navigation. It synergizes Temporal-Adaptive Resolution (TAR), which eliminates historical redundancy via decay-based resizing, and Stratified Structure-aware Pruning (SSP), which prioritizes interactive foregrounds and semantic anchors while safeguarding global layout. Extensive evaluations across diverse benchmarks demonstrate that GUIPruner consistently achieves state-of-the-art performance, effectively preventing the collapse observed in large-scale models under high compression. Notably, on Qwen2-VL-2B, our method delivers a 3.4x reduction in FLOPs and a 3.3x speedup in vision encoding latency while retaining over 94% of the original performance, enabling real-time, high-precision navigation with minimal resource consumption.

cs.CV

Target-Balanced Score Distillation

Score Distillation Sampling (SDS) enables 3D asset generation by distilling priors from pretrained 2D text-to-image diffusion models, but vanilla SDS suffers from over-saturation and over-smoothing. To mitigate this issue, recent variants have incorporated negative prompts. However, these methods face a critical trade-off: limited texture optimization, or significant texture gains with shape distortion. In this work, we first conduct a systematic analysis and reveal that this trade-off is fundamentally governed by the utilization of the negative prompts, where Target Negative Prompts (TNP) that embed target information in the negative prompts dramatically enhancing texture realism and fidelity but inducing shape distortions. Informed by this key insight, we introduce the Target-Balanced Score Distillation (TBSD). It formulates generation as a multi-objective optimization problem and introduces an adaptive strategy that effectively resolves the aforementioned trade-off. Extensive experiments demonstrate that TBSD significantly outperforms existing state-of-the-art methods, yielding 3D assets with high-fidelity textures and geometrically accurate shape.

cs.CV

Collusion-Driven Impersonation Attack on Channel-Resistant RF Fingerprinting

Radio frequency fingerprint (RFF) is a promising device identification technology, with recent research shifting from robustness to security due to growing concerns over vulnerabilities. To date, while the security of RFF against basic spoofing such as MAC address tampering has been validated, its resilience to advanced mimicry remains unknown. To address this gap, we propose a collusion-driven impersonation attack that achieves RF-level mimicry, successfully breaking RFF identification systems across diverse environments. Specifically, the attacker synchronizes with a colluding receiver to match the centralized logarithmic power spectrum (CLPS) of the legitimate transmitter; once the colluder deems the CLPS identical, the victim receiver will also accept the forged fingerprint, completing RF-level spoofing. Given that the distribution of CLPS features is relatively concentrated and has a clear underlying structure, we design a spoofed signal generation network that integrates a variational autoencoder (VAE) with a multi-objective loss function to enhance the similarity and deceptive capability of the generated samples. We carry out extensive simulations, validating cross-channel attacks in environments that incorporate standard channel variations including additive white Gaussian noise (AWGN), multipath fading, and Doppler shift. The results indicate that the proposed attack scheme essentially maintains a success rate of over 95% under different channel conditions, revealing the effectiveness of this attack.

cs.CR

A Hierarchical Constructive Heuristic for Large-Scale Survivable Traffic Grooming Problem under Double-Link Failures

This paper studies a survivable traffic grooming problem in large-scale optical transport networks under double-link failures (STG2). Each communication demand must be assigned a route for every possible scenario involving zero, one, or two failed fiber links. Protection against double-link failures is crucial for ensuring reliable telecommunications services while minimizing equipment costs, making it essential for telecommunications companies today. However, this significantly complicates the problem and is rarely addressed in existing studies. Furthermore, current research typically examines networks with fewer than 300 nodes, much smaller than some emerging networks containing thousands of nodes. To address these challenges, we propose a novel hierarchical constructive heuristic for STG2. This heuristic constructs and assigns routes to communication demands across different scenarios by following a hierarchical sequence. It incorporates several innovative optimization techniques and utilizes parallel computing to enhance efficiency. Extensive experiments have been conducted on large-scale STG2 instances provided by our industry partner, encompassing networks with 1,000 to 2,600 nodes. Results demonstrate that within a one-hour time limit and a 16 GB memory limit set by the industry partner, our heuristic improves the objective values of the best-known solutions by 18.5\% on average, highlighting its significant potential for practical applications.

math.CO

LP Relaxations for Routing and Wavelength Assignment with Partial Path Protection: Formulations and Computations

As a variant of the routing and wavelength assignment problem (RWAP), the RWAP with partial path protection (RWAP-PPP) designs a reliable optical-fiber network for telecommunications. It assigns paths and wavelengths to meet communication requests, not only in normal working situations but also in potential failure cases where an optical link fails. The literature lacks efficient relaxations to produce tight lower bounds on the optimal objective value of the RWAP-PPP. Consequently, the solution quality for the RWAP-PPP cannot be properly assessed, which is critical for telecommunication providers in customer bidding and service improvement. Due to numerous failure scenarios, developing effective lower bounds for the RWAP-PPP is challenging. To address this, we formulate and analyze various linear programming (LP) relaxations of the RWAP-PPP. Among them, we propose a novel LP relaxation yielding promising lower bounds. To solve it, we develop a Benders decomposition algorithm with valid inequalities to enhance performance. Computational results on practical networks, including large ones with hundreds of nodes and edges, demonstrate the effectiveness of the LP relaxation and efficiency of its algorithm. The obtained lower bounds achieve average optimality gaps of 8.6%. Compared with a direct LP relaxation of the RWAP, which has average gaps of 36.7%, significant improvements are observed. Consequently, our LP relaxation and algorithm effectively assess RWAP-PPP solution quality, offering significant research and practical value.

math.OC

Co-Learning: Code Learning for Multi-Agent Reinforcement Collaborative Framework with Conversational Natural Language Interfaces

Online question-and-answer (Q\&A) systems based on the Large Language Model (LLM) have progressively diverged from recreational to professional use. This paper proposed a Multi-Agent framework with environmentally reinforcement learning (E-RL) for code correction called Code Learning (Co-Learning) community, assisting beginners to correct code errors independently. It evaluates the performance of multiple LLMs from an original dataset with 702 error codes, uses it as a reward or punishment criterion for E-RL; Analyzes input error codes by the current agent; selects the appropriate LLM-based agent to achieve optimal error correction accuracy and reduce correction time. Experiment results showed that 3\% improvement in Precision score and 15\% improvement in time cost as compared with no E-RL method respectively. Our source code is available at: https://github.com/yuqian2003/Co_Learning

cs.SE

DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition

Neuromorphic sensors, specifically event cameras, revolutionize visual data acquisition by capturing pixel intensity changes with exceptional dynamic range, minimal latency, and energy efficiency, setting them apart from conventional frame-based cameras. The distinctive capabilities of event cameras have ignited significant interest in the domain of event-based action recognition, recognizing their vast potential for advancement. However, the development in this field is currently slowed by the lack of comprehensive, large-scale datasets, which are critical for developing robust recognition frameworks. To bridge this gap, we introduces DailyDVS-200, a meticulously curated benchmark dataset tailored for the event-based action recognition community. DailyDVS-200 is extensive, covering 200 action categories across real-world scenarios, recorded by 47 participants, and comprises more than 22,000 event sequences. This dataset is designed to reflect a broad spectrum of action types, scene complexities, and data acquisition diversity. Each sequence in the dataset is annotated with 14 attributes, ensuring a detailed characterization of the recorded actions. Moreover, DailyDVS-200 is structured to facilitate a wide range of research paths, offering a solid foundation for both validating existing approaches and inspiring novel methodologies. By setting a new benchmark in the field, we challenge the current limitations of neuromorphic data processing and invite a surge of new approaches in event-based action recognition techniques, which paves the way for future explorations in neuromorphic computing and beyond. The dataset and source code are available at https://github.com/QiWang233/DailyDVS-200.

cs.CV

A real-time hole depth diagnostic based on coherent imaging with plasma amendment during femtosecondlaser hole-drilling

An in-process coherent imaging diagnostic has been developed to real-time measure the hole depth during air-film hole drilling by a femtosecond laser. A super-luminescent diode with a wavelength of 830~13 nm is chosen as the coherent light source which determines a depth resolution of 12 μm. The drilled hole is coupled as a part of the sample arm and the depth variation can be extracted from the length variation of the optical path. Interference is realized in the detection part and a code has been written to discriminate the interference fringes. Density of plasma in the hole is diagnosed to evaluate its amendment to the optical path length and the depth measurement error induced by plasma is non-ignorable when drilling deep holes.

physics.ins-det

Non-negative Sparse and Collaborative Representation for Pattern Classification

Sparse representation (SR) and collaborative representation (CR) have been successfully applied in many pattern classification tasks such as face recognition. In this paper, we propose a novel Non-negative Sparse and Collaborative Representation (NSCR) for pattern classification. The NSCR representation of each test sample is obtained by seeking a non-negative sparse and collaborative representation vector that represents the test sample as a linear combination of training samples. We observe that the non-negativity can make the SR and CR more discriminative and effective for pattern classification. Based on the proposed NSCR, we propose a NSCR based classifier for pattern classification. Extensive experiments on benchmark datasets demonstrate that the proposed NSCR based classifier outperforms the previous SR or CR based approach, as well as state-of-the-art deep approaches, on diverse challenging pattern classification tasks.

cs.CV

Predicting Crash Fault Residence via Simplified Deep Forest Based on A Reduced Feature Set

The software inevitably encounters the crash, which will take developers a large amount of effort to find the fault causing the crash (short for crashing fault). Developing automatic methods to identify the residence of the crashing fault is a crucial activity for software quality assurance. Researchers have proposed methods to predict whether the crashing fault resides in the stack trace based on the features collected from the stack trace and faulty code, aiming at saving the debugging effort for developers. However, previous work usually neglected the feature preprocessing operation towards the crash data and only used traditional classification models. In this paper, we propose a novel crashing fault residence prediction framework, called ConDF, which consists of a consistency based feature subset selection method and a state-of-the-art deep forest model. More specifically, first, the feature selection method is used to obtain an optimal feature subset and reduce the feature dimension by reserving the representative features. Then, a simplified deep forest model is employed to build the classification model on the reduced feature set. The experiments on seven open source software projects show that our ConDF method performs significantly better than 17 baseline methods on three performance indicators.

cs.SE

MIMO Radar Waveform-Filter Design for Extended Target Detection from a View of Games

This paper studies the Two-Person Zero Sum(TPZS) game between a Multiple-Input Multiple-Output(MIMO) radar and an extended target with payoff function being the output Signal-to-Interference-pulse-Noise Ratio(SINR) at the radar receiver. The radar player wants to maximize SINR by adjusting its transmit waveform and receive filter. Conversely, the target player wants to minimize SINR by changing its Target Impulse Response(TIR) from a scaled sphere centered around a certain TIR. The interaction between them forms a Stackelberg game where the radar player acts as a leader. The Stackelberg equilibrium strategy of radar, namely robust or minimax waveform-filter pair, for three different cases are taken into consideration. In the first case, Energy Constraint(EC) on transmit waveform is introduced, where we theoretically prove that the Stackelberg equilibrium is also the Nash equilibrium of the game, and propose Algorithm 1 to solve the optimal waveform-filter pair through convex optimization. Note that the EC can't meet the demands of radar transmitter due to high Peak Average to power Ratio(PAR) of the transmit waveform, thus Constant Modulus and Similarity Constraint(CM-SC) on waveform is considered in the second case, and Algorithm 2 is proposed to solve this problem, where we theoretically prove the existence of Nash equilibrium for its Semi-Definite Programming(SDP) relaxation form. And the optimal waveform-filter pair is solved by calculating the Nash equilibrium followed by the randomization schemes. In the third case,...

eess.SP

Measuring Female Representation and Impact in Films over Time

Women have always been underrepresented in movies and not until recently has the representation of women in movies improved. To investigate the improvement of female representation and its relationship with a movie's success, we propose a new measure, the female cast ratio, and compare it to the commonly used Bechdel test result. We employ generalized linear regression with $L_1$ penalty and a Random Forest model to identify the predictors that influence female representation, and evaluate the relationship between female representation and a movie's success in three aspects: revenue/budget ratio, rating, and popularity. Three important findings in our study have highlighted the difficulties women in the film industry face both upstream and downstream. First, female filmmakers, especially female screenplay writers, are instrumental for movies to have better female representation, but the percentage of female filmmakers has been very low. Second, movies that have the potential to tell insightful stories about women are often provided with lower budgets, and this usually causes the films to in turn receive more criticism. Finally, the demand for better female representation from moviegoers has also not been strong enough to compel the film industry to change, as movies that have poor female representation can still be very popular and successful in the box office.

cs.CY

The monochromatic X-rays facilities at NIM

Space scientific exploration is becoming the main battlefield for mankind to explore the universe. Countries around the world have successively launched various space exploration satellites. Accurate calibration on the ground is a key factor for space science satellites to obtain observational results. In order to provide calibration for various satellite-borne detectors, several monochromatic X-rays facilities has been built at National Institute of Metrology, P.R. China (NIM). These facilities are mainly based on grating diffraction and Bragg diffraction, the energy range of produced monochromatic X-rays is (0.218-301) keV. The facilities have a good performance on energy stability, monochromaticity and flux stability. Monochromaticity of all facilities is better than 3.0%, the stability of energy is better than 1.0% over 8 hours, and the stability of flux is better than 2.0% over 8 hours. The calibration experiments of satellite-borne detectors, such as energy linearity, energy resolution, detection efficiency and temperature response can be carried out on the facilities. So far we have completed the calibration of two satellites, and there are still three satellites in progress. This work will contribute to the development of X-ray astronomy, and contribute to the development of Chinese space science.

astro-ph.IM

Robust MIMO Radar Waveform-Filter Design for Extended Target Detection in the Presence of Multipath

The existence of multipath brings extra "looks" of targets. This paper considers the extended target detection problem with a narrow band Multiple-Input Multiple-Output(MIMO) radar in the presence of multipath from the view of waveform-filter design. The goal is to maximize the worst-case Signal-to-Interference-pulse-Noise Ratio(SINR) at the receiver against the uncertainties of the target and multipath reflection coefficients. Moreover, a Constant Modulus Constraint(CMC) is imposed on the transmit waveform to meet the actual demands of radar. Two types of uncertainty sets are taken into consideration. One is the spherical uncertainty set. In this case, the max-min waveform-filter design problem belongs to the non-convex concave minimax problems, and the inner minimization problem is converted to a maximization problem based on Lagrange duality with the strong duality property. Then the optimal waveform is optimized with Semi-Definite Relaxation(SDR) and randomization schemes. Therefore, we call the optimization algorithm Duality Maximization Semi-Definite Relaxation(DMSDR). Additionally, we further study the case of annular uncertainty set which belongs to non-convex non-concave minimax problems. In order to address it, the SDR is utilized to approximate the inner minimization problem with a convex problem, then the inner minimization problem is reformulated as a maximization problem based on Lagrange duality. We resort to a sequential optimization procedure alternating between two SDR problems to optimize the covariance matrix of transmit waveform and receive filter, so we call the algorithm Duality Maximization Double Semi-Definite Relaxation(DMDSDR). The convergences of DMDSDR are proved theoretically. Finally, numerical results highlight the effectiveness and competitiveness of the proposed algorithms as well as the optimized waveform-filter pair.

eess.SP