SearcharxivSearch

arXiv subjects

Kai Xu

Publications and source records attributed to Kai Xu.

At least 19 recordsLinked to original sources

Event-triggered Implicit Perturbation for Zeroth-Order Fine-Tuning of Spiking Transformers

Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for fine-tuning non-differentiable, event-driven spiking neural networks (SNNs). However, its deployment on in-memory computing (IMC) accelerators is constrained by the repeated read-modify-write (RMW) operations arising from explicit weight perturbation and the prohibitive hardware footprint of random number generators (RNGs) for statistically independent per-weight perturbations. To address these challenges, we propose an implicit-perturbation ZO (IPZO) architecture in which perturbation sums computed by an event-triggered perturbation generation unit (PGU) are combined with the weighted sums produced by the IMC array, eliminating perturbation-induced RMW operations while preserving weight-stationary execution of IMC. By exploiting spike sparsity, the PGU generates and accumulates perturbation contributions only for spike-activated weight rows, reducing the required row dimension of the RNG array. An address-driven XOR recombination scheme (PGU-XOR) is further introduced to mitigate the spatial correlations caused by direct RNG reuse (PGU-Reuse). The results show that (1) PGU-XOR matches software RNGs in accuracy on Spikingformer/CIFAR-10 (76.41% vs. 76.53%) and perplexity (PPL) on SpikeGPT/WikiText-2 (54.20 vs. 53.23), whereas PGU-Reuse degrades accuracy by 9.56 percentage points and increases PPL by 11.8; (2) implemented in a TSMC 16-nm CMOS technology, PGU-XOR incurs 40.3%-46.0% area and 15.2%-48.9% energy overhead per matrix-vector multiplication relative to PGU-Reuse, yet its faster convergence reduces the total perturbation energy to 0.51x that of PGU-Reuse at iso-accuracy; (3) IPZO reduces the perturbation energy to 0.46x-0.83x that of conventional explicit weight perturbation for a batch size of B=64 and T=4 time steps, with the advantage growing as BT decreases.

cs.AR

Intersection Bounds for BPS Strings in Six-Dimensional Supergravity

In six-dimensional $\mathcal{N}=(1,0)$ supergravity, the structure of tensor moduli space is governed by primitive BPS string charges known as BPS generators and their intersection pairing. We derive bounds on the intersection numbers of these generators from a purely effective field theory (EFT) perspective. Although gauge anomaly cancellation constrains intersections between generators supporting gauge algebras, bounds for E-strings intersecting generators with self-intersection numbers $-2$ and $-3$ have previously remained incomplete. We show that the Zariski decomposition, interpreted as the charge lattice counterpart of the attractor mechanism, together with current algebra embeddings on the E-string worldsheet theory, yields strong universal bounds on these intersection numbers. These results establish the finiteness of tensor charge intersection numbers up to duality. The underlying structure was identified through AI-guided investigation and is proven here analytically using EFT arguments.

hep-th

From Blind Search to Memory-Aware Evolution: Efficient DBMS Tuning via Collaborative Diagnosis and Utility-Aware Retrieval

Modern DBMSs expose multiple configurable components (e.g., knobs, query hints, and indexes) that jointly determine query performance. Multi-component tuning is challenging due to the large combinatorial search space and the difficulty of learning effective tuning policies under limited feedback. Existing approaches still rely on blind search over the configuration space and interaction-heavy policy learning, leading to high tuning overhead and limited performance gains. Recent advances in large language models (LLMs) enable knowledge-driven tuning, but existing LLM-based methods fail to effectively exploit online feedback and historical observations, often converging prematurely to suboptimal configurations. In this paper, we present EvoTune, a memory-aware evolution framework for multi-component DBMS tuning. EvoTune first localizes a query-specific high-impact subspace via collaborative diagnosis, which combines lightweight pattern learning with LLM-based reasoning. It further introduces a utility-aware retrieval policy that selects informative observations based on their resulting long-term performance improvement, instead of similarity-based retrieval. To support continual improvement, EvoTune organizes tuning feedback into a hierarchical memory and incrementally refines both subspace localization and tuning policies without requiring LLM fine-tuning. Extensive experiments show that EvoTune consistently outperforms state-of-the-art baselines, achieving up to 44.5% performance improvement under the same tuning budget and reaching the best competing baseline's final performance up to 3.9X faster.

cs.DB

Connectivity-induced surface-loss penalty in superconducting qubit-coupler lattices

Recent advances in design and fabrication have increased the energy-relaxation times of isolated superconducting transmon qubits to the hundreds-of-microseconds regime, with reported values exceeding 500 $\mu$s. However, the same progress has not automatically translated to multiqubit processors, where qubits are embedded in connected qubit-coupler lattices and often exhibit much shorter lifetimes than isolated qubits. To identify possible sources of this discrepancy, here we use finite-element simulation to investigate how surface participation ratios and the resulting surface dielectric loss change when a qubit is embedded in a flip-chip qubit-coupler lattice. Controlled comparisons show that higher connectivity can indeed lead to larger surface loss: in the simulated lattice, connecting a qubit to two and four couplers increases the surface loss by factors of 1.3 and 1.8, respectively. We attribute this change to the combined effects of added edge fields from coupling claws, field redistribution over the larger connected metal network, and hybridization with coupler modes. We further examine how this connectivity-induced surface-loss penalty depends on the geometric design parameters of both the qubit electrodes and the coupling claws, and derive guidelines for designing low-loss multiqubit processors.

quant-ph

HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations

Standard pipelines for physics-ready 3D reconstruction rely on a decoupled two-stage paradigm: extracting surface geometry followed by an error-prone tetrahedralization process. While recent Lagrangian methods like TetSphere Splatting attempt to bypass this by directly optimizing volumetric primitives, their homeomorphic constraints prevent topology-adaptive optimization. Consequently, they produce disjoint tetrahedra rather than a single connected mesh, rendering the structures unsuitable for further physical simulations. To address this, we propose a topology-adaptive framework for holistic tetrahedral mesh reconstruction through end-to-end topological and geometric optimization. First, by coupling Gaussian spheres to tetrahedral elements and leveraging edge connections, we estimate a continuous opacity field for differentiable element pruning. Next, jointly minimizing mesh smoothing energy and multi-view Gaussian rendering error drives alternating geometric refinement while preserving topological adaptivity. Consequently, our approach effectively constructs a unified and topologically coherent tetrahedral mesh. Extensive experiments demonstrate that our method outperforms state-of-the-art techniques by achieving superior geometric accuracy and producing coherent, single-connected tetrahedral meshes, thereby effectively bypassing the error-prone conventional tetrahedralization step for reconstructed surface meshes and streamlining downstream physical simulation.

cs.GR

Infinity-harmonic functions and inverse mean curvature flow clusters

An $\infty$-harmonic function is a viscosity solution of $\nabla^2 u(\nabla u,\nabla u)=0$, or equivalently, an absolute minimizer of $\|\nabla u\|_{L^\infty}$. We prove a variety of new structural and regularity results in two dimensions, including: 1. $\infty$-harmonic functions in domains of $\mathbb{R}^2$ are $C^{1,1/3}$. 2. Critical points are isolated, and at each critical point, the solution has a unique quasiradial blow-up. 3. Entire solutions with polynomial growth have unique quasiradial blow-downs, and are determined by their Fourier modes at infinity. These results are consequences of a new theory relating $\infty$-harmonic functions to inverse mean curvature flow (IMCF) clusters -- which are piecewise weak solutions of IMCF with common obstacle-type boundary conditions on the interfaces (a simple example is an embedded family of cuspidal curves evolving by inverse curvature). This connection arises as the $p\to\infty$ limit of the classical duality between $p$-harmonic and $q$-harmonic functions in $\mathbb{R}^2$, where $\frac1p+\frac1q=1$.

math.AP

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers

Fully homomorphic encryption (FHE) enables computation on encrypted data, but practical encrypted Transformer inference is bottlenecked by the sequential composition of many nonlinear blocks. We study whether Structured Newton Layer Parallelism (SNLP) can make this inter-layer composition more FHE-friendly: each Transformer block still requires polynomial approximations for operations such as softmax and RMSNorm, but SNLP reduces the layerwise sequential nonlinear depth from L stages to a small number of solver iterations plus linear structured corrections. Using a simulation framework based on Chebyshev polynomial approximations, we measure error accumulation under sequential versus SNLP inference across 8 models and 4 architecture families. On a 0.5B IDN-trained model, SNLP reduces symbolic bootstraps from 53 to 20 (2.65x) with only +1.2% perplexity degradation, while lowering error amplification (1.36x vs. 1.42x). Across all tested models, SNLP has lower amplification than sequential inference. Ablations show that softmax approximation dominates the error budget and CKKS arithmetic noise is negligible in our setting, suggesting that SNLP is complementary to block-level FHE-friendly operator design rather than a replacement for it.

cs.LG

A superconducting qutrit link beyond the qubit limit

Superconducting microwave links have enabled deterministic state transfer and remote entanglement between qubits, but deterministic links have so far operated with an effectively two-dimensional transmitted Hilbert space. Here we demonstrate a superconducting qutrit link between two independently packaged nodes connected by a microwave channel. Each node combines a transmon qutrit, a transmission resonator, and a tunable Purcell-filter interface, allowing the two remote microwave-photon interfaces to be matched in both frequency and bandwidth. We implement two transition-selective photon-mediated operations that transfer the $|e\rangle$ and $|f\rangle$ qutrit components in distinct temporal modes of the same channel. We tomographically characterize arbitrary qutrit-state transfer, obtaining a mean transferred-state fidelity of 83.68% and a qutrit process fidelity of 77.12%, exceeding both the classical qutrit-transfer benchmark and the best possible average fidelity of an effective qubit channel used to transmit an arbitrary qutrit. Using partial-transfer operations, we reconstruct a remote two-qutrit state with negativity 0.730, a tomography-inferred dense-coding capacity of 2.273 bits, and a tomography-inferred Collins-Gisin-Linden-Massar-Popescu (CGLMP) parameter $I_3=2.332$, all beyond the corresponding qubit or local bounds. These results demonstrate a superconducting microwave link that uses the native three-level structure of transmons as a genuine high-dimensional communication resource.

quant-ph

DynaMOMA: Instantaneous Prediction of Grasp Poses for Mobile Manipulation of Dynamic Objects

Mobile manipulation is a fundamental robotics task and has advanced rapidly in recent years, enabling robots to navigate, reach, and interact with objects in complex environments. However, mobile manipulation of dynamic objects remains highly challenging, as robots must coordinate the mobile base and arm while adapting to continuously evolving target poses. A key challenge lies in predicting temporally consistent short-horizon grasp trajectories from dynamic observations. In this work, we propose \ours{}, a dynamic mobile manipulation framework that couples instantaneous grasp trajectory prediction with whole-body control policy. Our predictor uses an anchor-based diffusion model to generate temporally consistent short-horizon grasp trajectories conditioned on historical observations. The predicted trajectories are then encoded as compact features and fed to a whole-body reinforcement learning policy, which controls the mobile manipulator for dynamic grasping. We further introduce a anticipation-guided reward that equips the policy with an anticipatory grasping horizon by adaptively shifting the target from the current grasp observation to the instantaneously predicted grasp trajectory. Through extensive experiments in Isaac Gym simulation, we show that our method achieves strong performance in mobile manipulation of dynamic objects across diverse settings and grasping metrics. Furthermore, our predictor and policy demonstrate strong generalizability in real-world experiments.

cs.RO

Is Agent Code Less Maintainable Than Human Code?

Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when future agents build on top of it, potentially leading to compounding downstream effects. We investigate how agent code compares to human code in these maintenance settings, presenting CodeThread, a framework to construct controlled experiments from repository-level coding benchmarks. Applying CodeThread to four frontier coding agents and four benchmarks, we find that agents are less effective at resolving tasks when building on agent code compared to human code, with task resolve rate drops of up to 13.1%. Regression analysis reveals that many traditional software engineering maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty. These findings highlight the need to evaluate these systems not only by immediate task resolution but also by code maintainability, and point to potential sources of downstream errors introduced by agent code.

cs.SE

A symmetric relaxation method for entire two-dimensional cellular networks and its implications

To simulate the relaxation of an entire 2D cellular network, this study proposes a symmetric relaxation method for both inner and marginal vertices. The relaxations of these two types of vertices are determined by the central angle symmetry of associated cells and the angle symmetry at each vertex, but with different major considerations. Trimmed Voronoi networks with varying irregularity are used as initial networks for the relaxation simulation. In particular, we propose a regular hexagon disordering method to generate Voronoi networks and find that the inner cells of networks with an irregularity value of one exhibit a conserved edge number distribution, as found in other 2D cellular networks. Simulation results agree with the von Neumann-Mullins law for both inner and marginal cells, and a modified equation including a geometric correction term significantly improves prediction quality. The Aboav-Weaire law and Lewis law are also reproduced, with the latter showing that relaxed cells tend to approach the ellipses' maximum inscribed polygons. Analysis of edge length, interior angle, and shape index reveals that symmetric relaxation inhibits T1 (neighbour exchange) topological transitions by reducing short edges while increasing area disparity among neighbouring cells. The findings suggest that T1 events may be triggered when force disequilibrium overcomes the stabilising effect of symmetric relaxation, providing a possible mechanistic explanation for T1 in 2D foams.

physics.bio-ph

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on multiple cameras (egocentric, eye-to-hand, and wrist-mounted) for policy learning, current multi-view world models simply concatenate view tokens without explicit geometric reasoning. This causes cross-view object drift, depth inconsistency, and texture misalignment. We trace these failures to two deficiencies: the absence of an explicit inter-view communication mechanism and the lack of a 3D geometric prior. We argue that resolving both simultaneously is necessary and sufficient. To address this, we present PAIWorld, a framework that augments diffusion-transformer world models via three core components: (1) Geometry-Aware Cross-View Attention blocks that establish an explicit pathway across views, (2) Geometric Rotary Position Embedding that encodes camera ray directions and extrinsic poses into the attention mechanism, and (3) Latent 3D-REPA, which distills 3D-aware features from frozen 3D foundation models to ensure 3D consistency. Built upon a DiT-based world foundation model, PAIWorld achieves state-of-the-art multi-view 3D consistency on robotic manipulation benchmarks, ranking 1st on the WorldArena leaderboard and 2nd on the AgiBot-Challenge2026 leaderboard, while enabling downstream applications such as model-based planning, world action models, and multi-view policy post-training.

cs.RO

On BPS Branes

We study supersymmetric BPS branes (BPS-B) in supergravity theories. Some of these states are anticipated by BPS black-brane (BPS-BB) solutions of supergravity. In particular, we define and distinguish the cone generated by BPS branes from the subcone of charges that admit BPS black-brane attractor solutions in the infrared limit of the supergravity effective field theory. We denote these cones by $C_{\rm BPS-B}$ and $C_{\rm BPS-BB}$, respectively. We conjecture that, in any supersymmetric theory of quantum gravity, every integrally charged state lying in $C_{\rm BPS-BB}$ is realized by a BPS state in the spectrum. Furthermore, we conjecture and present evidence that when $C_{\rm BPS-B}$ is moduli independent, it can be determined as the cone dual to the $C_{\rm BPS-BB}$ under the electric-magnetic pairing.

hep-th

Intrinsic Selection and Particle Resampling for Inference-Time Scaling Beyond Domain Verifiability

Inference-Time Scaling (ITS) has largely succeeded in verifiable domains like math and coding, where cheap verification enables scalable output selection. However, extending ITS to tasks prone to systematic failure - driven by faulty initial assumptions or unmet multidimensional constraints - typically relies on costly external solvers or brittle, model-based verifiers. Our key insight is that the intrinsic statistics of parallel sample sets, specifically length-adjusted tail entropy, provide a robust discriminative signal for solution quality without access to ground truth. Crucially, these statistics serve as a difficulty gate for adaptive compute allocation, dynamically routing problems across scaling regimes. First, Intrinsic Selection (iS) ranks candidates post-hoc, matching consensus-based algorithms across three domains and improving engineering design selection by 20% over pass@1 baselines. Second, Intrinsic Particle Filtering (iPF) generalizes this to step-level resampling, guiding generation toward high-confidence reasoning trajectories to improve pass@1 by 6.1 points on average on hard math problems. Finally, Particle Distillation (dPF) injects privileged guidance via early logit blending and KL-guided resampling, steering generation past systematic reasoning errors to satisfy expert rubrics, yielding up to 26.5% gains on complex clinical responses. Our pipeline applies seamlessly across broad-purpose, domain-specialized, and multimodal architectures, successfully extending ITS to open-ended domains without requiring trained reward models or exact ground-truth verification.

cs.LG

sGPO: Trading Inference FLOPs for Training Efficiency in RLVR

Standard Reinforcement Learning with Verifiable Rewards (RLVR) training allocates a fixed rollout budget to every query, without regard for what each query's difficulty means for the current policy. This leads to two symmetric failure modes: easy queries produce near-zero advantage because the policy already solves them, while unsolvable queries produce no signal because the policy never solves them. Both regimes waste training FLOPs without contributing to a learning gradient. We introduce sorted Group Policy Optimization (sGPO), a compute-efficient strategy that trades a small budget of inference FLOPs for a large reduction in wasted training FLOPs. The key insight is that cheap inference compute can serve as a single offline proxy for query difficulty. By generating a small batch of parallel samples per query under the initial policy, we obtain a model-aware empirical success rate. This motivates setting the training rollout group size to the inverse of this success rate, a practical rule that maximizes sample efficiency by extracting the most advantage per generated rollout. This single profiling pass simultaneously drives data filtering (removing trivial queries and sub-sampling unsolvable ones), adaptive group size allocation, and curriculum construction (scheduling queries from easy to hard). sGPO matches or exceeds baseline performance while reducing total training compute by a factor of three, with the upfront inference profiling cost included.

cs.LG

Programmable spectral symmetries in an anisotropic quantum Rabi simulator

The quantum Rabi model captures fundamental aspects of light--matter interaction, where symmetry dictates both spectra and dynamics. Over the past years, experiments have explored many of its nonperturbative properties, but have mostly focused on the isotropic limit, where rotating and counterrotating processes are locked together, leaving the broader symmetry landscape largely unexplored. Here we realize a programmable anisotropic quantum Rabi model in a superconducting processor, with independent control of the rotating and counterrotating couplings $(g_1,g_2)$ and of a transverse bias $\varepsilon$. Continuous anisotropy tuning, combined with a duality mapping, gives access to the full parameter space from the Jaynes-Cummings to the anti-Jaynes-Cummings limits. In the deep-strong-coupling regime, we show that anisotropy reconstructs the spectrum and turns complete collapse-revival dynamics into incomplete revivals even near degeneracy. With adiabatic state preparation and joint tomography, we resolve an anisotropy-induced ground-state parity switch, a crossing that has no analogue in the isotropic model. We further observe selective tunnelling associated with hidden symmetry in biased Rabi models and track its anisotropic displacement within the same device. These results establish a controllable route to engineering nonperturbative light--matter Hamiltonians, where symmetry, spectrum, and dynamics can be programmed independently.

quant-ph

Role of Characteristic Length Scale in Interface Graphitization-Induced Wear Resistance of Diamond and Amorphous Carbon

The evolution of interfacial atomic structures critically influences the friction and wear behavior of carbon-based materials. However, how the characteristic length scale of friction-induced sp\textsuperscript{2} reconstruction governs macroscopic wear remains poorly understood, particularly for diamond and amorphous carbon where the interfacial graphitization modes differ fundamentally. In this work, we develop a machine learning potential for these carbon systems and investigate the structural evolution at interfaces in both diamond/diamond and amorphous/amorphous carbon systems using molecular dynamics simulations. Our results reveal distinct atomic-scale characteristics of graphitization at the two interfaces. Diamond interfaces develop a laterally continuous sp\textsuperscript{2} reconstruction layer with a characteristic length of 30--45~\AA, while amorphous carbon interfaces form only fully isolated sp\textsuperscript{2} patches of 8--12~\AA. This disparity in characteristic length scale determines the density of weakly bonded interfacial atoms left outside the reconstruction layer, thereby directly dictating the macroscopic wear rate. Based on these insights, we propose a strategy to regulate friction-induced graphitization in diamond coatings by protecting specific crystallographic orientations, such as the (111) close-packed planes. This work bridges the gap between atomic-scale interfacial structure and macroscopic tribological performance, offering mechanistic guidelines for the rational design of wear-resistant carbon-based coatings.

cond-mat.mtrl-sci

PaintBench: Deterministic Evaluation of Precise Visual Editing

While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To probe this challenge, we introduce PaintBench, a dynamically scalable benchmark targeting 20 fundamental precise visual editing operations across four categories: geometric transformation, structural manipulation, color change, and symbolic reasoning. Procedural generation with configurable complexity enables an effectively infinite, contamination-resistant evaluation suite, and deterministic pixel-level evaluation eliminates reliance on bias-prone judge models. Across 11 image editing models, we find overall low performance, with the current highest-performing industry leader scoring only 17.1% (mIoU). Task decomposition reveals especially challenging operation types (geometric transformation, most structural manipulation, formula-based color change) and model-specific specializations. Fine-grained benchmark diagnostics further show performance degradations induced by scene variations in object count, background complexity, color scheme, and edit-region size. To test generalization of PaintBench scores to applied task performance, we create a procedural, deterministic evaluation for data visualization editing (TinyGrafixBench) and find strong linear correlation with PaintBench scores ($R^2 = 0.91$, $p < 0.001$). Altogether, PaintBench provides a rigorous foundation for measuring and driving progress in precise multimodal visual editing.

cs.GR