SearcharxivSearch

arXiv subjects

Andrew Xu

Publications and source records attributed to Andrew Xu.

8 recordsLinked to original sources

Accelerating the Conjugate Gradient Method by Solving Multiple GPU-Parallelized Duplicate Systems

We propose a method to accelerate the conjugate gradient method (CG) on parallel computing hardware by performing block conjugate gradient on one linear system. We aim to reduce the number of iterations until convergence by solving several copies of the system to explore multiple regions of the solution space at once. Although solving multiple systems requires more floating point operations than solving one, these computations are highly parallelizable, even without additional hardware resources. Our method does not replace preconditioning and can be used with any preconditioner. We developed linear regression models to predict iteration reduction and speedup as functions of the coefficient matrix's size, its number of nonzeros, and the solve time with CG. Our experiments show that our method can reduce solve time by up to 6 times in dense systems and up to 5 times in sparse systems that require long solve times with CG relative to their size.

math.NA

AllocBench: Measuring Online Tool Allocation Capability in LLM Agents

Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer an agent that creates a small number of highly reusable tools, rather than many one-offs. We introduce a paired benchmark that tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in two contexts: an abstract text-based formulation and a code-construction task. We find that every frontier model we test---Claude Haiku, Claude Opus, GPT-5.4-mini, and GPT-5.6 Sol---acts near-optimally in the abstract framing but fails to transfer this ability to script-writing. Through further experiments, we identify the particular failure modes for each model. Notably, the first three models fail even when the scripts are not evaluated, while GPT-5.6 Sol stays selective under that weaker manipulation and collapses only at full construction. Furthermore, an open-source Qwen model policy-trained for abstract allocation generalizes this ability across held-out lexical variations, but sees no improvement at script allocation. Together, these results establish online tool allocation as a significant capability boundary, even for modern frontier models.

cs.LG

Lomekwi: Resource-Bounded Tool Discovery in LLM Agents

Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distinguish tool use from tool discovery and decompose the latter into curiosity (the model's ability to discover the parts needed to build the tool), recognition (the model's ability to discover the process of creating the tool), and efficiency (the model's use of the tool after creation). We show that this framework can be applied to existing discovery tasks, such as Voyager. In addition, we provide evidence that recognition inversely scales with model size, and we introduce and analyze a class of combinatorial games that demonstrates this. We further observe inverse scaling in a separate environment designed to emulate real-world tasks.

cs.AI

Compositional Adversarial Training for Robust Visual Watermarking

Robust watermarking is typically trained with random post-processing augmentation, but random sampling under-covers the combinatorial space of realistic attack pipelines and rarely encounters the rare compositions that actually break detection. This leads to unstable training and poor sample efficiency. We instead formulate watermark robustness as a min-max problem over a structured space of compositional transformations. We propose Compositional Adversarial Training (CAT), a plug-in framework that learns a sequential differentiable adversary that observes the current watermarked image and selects an attack family at each step to maximally disrupt message recovery. CAT combines a straight-through Gumbel-Softmax attack selection with entropy regularization, allowing the backward pass to be end-to-end differentiable and aggregate gradient information across attack families, yielding faster, smoother convergence without collapsing to a single attack mode. We evaluate CAT on post-generation watermarks VideoSeal 0.0, VideoSeal 1.0, and PixelSeal and in-generation WMAR under both single-step and two-step attack suites, on in-distribution and multiple out-of-distribution image and video benchmarks. CAT consistently outperforms random-augmentation baselines trained with the same augmentation budget, with the largest gains on hard composed attacks and OOD evaluations; improving overall watermark capacity by up to $63.5\%$ in the single-step attack setting and $13.0\%$ in the compositional setting. In the autoregressive setting, CAT improves the TPR@FPR$=1\%$ by $12\%$ on average on difficult geometric transformations. These results show that robust visual watermarking benefits from training against adaptive compositional adversaries rather than independent random corruptions.

cs.CV

RankLLM: A Python Package for Reranking with LLMs

The adoption of large language models (LLMs) as rerankers in multi-stage retrieval systems has gained significant traction in academia and industry. These models refine a candidate list of retrieved documents, often through carefully designed prompts, and are typically used in applications built on retrieval-augmented generation (RAG). This paper introduces RankLLM, an open-source Python package for reranking that is modular, highly configurable, and supports both proprietary and open-source LLMs in customized reranking workflows. To improve usability, RankLLM features optional integration with Pyserini for retrieval and provides integrated evaluation for multi-stage pipelines. Additionally, RankLLM includes a module for detailed analysis of input prompts and LLM responses, addressing reliability concerns with LLM APIs and non-deterministic behavior in Mixture-of-Experts (MoE) models. This paper presents the architecture of RankLLM, along with a detailed step-by-step guide and sample code. We reproduce results from RankGPT, LRL, RankVicuna, RankZephyr, and other recent models. RankLLM integrates with common inference frameworks and a wide range of LLMs. This compatibility allows for quick reproduction of reported results, helping to speed up both research and real-world applications. The complete repository is available at rankllm.ai, and the package can be installed via PyPI.

cs.IR

Centralized Selection with Preferences in the Presence of Biases

This paper considers the scenario in which there are multiple institutions, each with a limited capacity for candidates, and candidates, each with preferences over the institutions. A central entity evaluates the utility of each candidate to the institutions, and the goal is to select candidates for each institution in a way that maximizes utility while also considering the candidates' preferences. The paper focuses on the setting in which candidates are divided into multiple groups and the observed utilities of candidates in some groups are biased--systematically lower than their true utilities. The first result is that, in these biased settings, prior algorithms can lead to selections with sub-optimal true utility and significant discrepancies in the fraction of candidates from each group that get their preferred choices. Subsequently, an algorithm is presented along with proof that it produces selections that achieve near-optimal group fairness with respect to preferences while also nearly maximizing the true utility under distributional assumptions. Further, extensive empirical validation of these results in real-world and synthetic settings, in which the distributional assumptions may not hold, are presented.

cs.DS

Dressed-state control of effective dipolar interaction between strongly-coupled solid-state spins

Strong interactions between spins in many-body solid-state quantum system is a crucial resource for exploring and applying non-classical states. In particular, electronic spins associated with defects in diamond system are a leading platform for the study of collective quantum phenomena and for quantum technology applications. While such solid-state quantum defect systems have the advantage of scalability and operation under ambient conditions, they face the key challenge of controlling interactions between the defects spins, since the defects are spatially fixed inside the host lattice with relative positions that cannot be well controlled during fabrication. In this work, we present a dressed-state approach to control the effective dipolar coupling between solid-state spins; and then demonstrate this scheme experimentally using two strongly-coupled nitrogen vacancy (NV) centers in diamond. Including Rabi driving terms between the m$_s$ = 0 and $\pm$1 states in the NV spin Hamiltonian allows us to turn on and off or tune the effective dipolar coupling between two NV spins. Through Ramsey spectroscopy, we detect the change of the effective dipolar field generated by the control NV spin prepared in different dressed states. To observe the change of interaction dynamics, we then deploy spin-lock-based polarization transfer measurements via a Hartmann-Hahn matching condition between two NV spins in different dressed states. We perform simulations that indicate the promise for this robust scheme to control the distribution of interaction strengths in strongly-interacting spin systems, including interaction strength homogenization in a spin ensemble, which can be a valuable tool for studying non-equilibrium quantum phases and generating high fidelity multi-spin correlated states for quantum-enhanced sensing.

quant-ph

Comparison of signal detectors for time domain radio SETI

The radio Search for Extra Terrestrial Intelligence (SETI) aims at identifying intelligent and communicative civilizations in the Universe through the detection of engineered transmissions. In the absence of prior knowledge concerning the expected signal, SETI detection pipelines necessitate high sensitivity, versatility, and limited computational complexity to maximize the search parameter space and minimize the probability of misses. This paper addresses the SETI detection problem as a binary hypothesis testing problem, and compares four detection schemes exploiting artificial features of the data collected by a single receiver radio telescope. After a theoretical comparison, those detectors are applied to real data collected with the Green Bank Telescope in West Virginia (USA).

eess.SP