SearcharxivSearch

arXiv subjects

Xudong Li

Publications and source records attributed to Xudong Li.

At least 19 recordsLinked to original sources

Remote epitaxy beyond polarity

Remote epitaxy through a monolayer two-dimensional material-covered substrate establishes a crystallographic registry across the van der Waals (vdW) surface that enables the epitaxial growth, lift-off and transfer of single-crystalline films. A central belief in remote epitaxy is that the substrate facilitating the phenomenon must be a material with strong ionicity, as the interatomic electrostatic potential fluctuation in covalent and metallic materials is substantially attenuated by two-dimensional materials. Here, we show remote epitaxy is possible when the substrate is a metallic or covalently bonded material and experimentally demonstrate non-polar remote homo- and heteroepitaxy across a wide range of material systems, including both metals and semiconductors. The achieved non-polar remote interactions are designed and engineered by harnessing substrate conductivity and vicinal surface step-edge density. These findings indicate that remote epitaxy is universal and applicable to ionic, metallic, and covalent materials, expanding its capabilities and stimulating a plethora of new fundamental scientific questions about the mechanism of remote epitaxy.

cond-mat.mtrl-sci

Exact low-dimensional reformulations for regularized spectral approximation

We study exact low-dimensional reformulations for a regularized spectral approximation problem and its weighted variant. In the unweighted setting, weak-majorization monotonicity of the fidelity term enables an exact reformulation over singular values. In the weighted setting, we establish such a reformulation under compatibility and simultaneous diagonalization conditions. We further show, via a counterexample, that this exact reformulation can fail in the absence of simultaneous diagonalization.

math.OC

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing methods often overfit to the evolution tasks, rely exclusively on trajectory-derived signals, and optimize harness components jointly, causing interference across components. We propose HarnessCompass, a novel automatic harness evolution framework built around constrained evolution, proactive feedback, and component-wise optimization. HarnessCompass first enforces global constraints on evolution, restricting modifications to task-agnostic harness changes that generalize beyond the evolution tasks. It then augments trajectory-derived evidence with proactive first-person feedback from the agent about harness usage, yielding richer signals for evolution. Finally, it decouples the optimization of different harness components before consolidating them into a unified harness, reducing cross-component interference while preserving component synergy. On SWE-bench Verified with GPT-5.4, HarnessCompass improves Pass@1 from 54\% to 66\% in only 5 evolution iterations, outperforming AHE in both effectiveness and evolution efficiency. In addition, the evolved harness transfers effectively to held-out tasks and other models, demonstrating substantially stronger generalization than prior automatic harness evolution methods.

cs.LG

Hybrid BaTiO3/TiO2 Metasurface for Efficient Gigahertz-Speed Free-Space Electro-Optic Modulation

Free-space electro-optic modulators are key to emerging photonic systems, yet their performance remains limited by trade-offs between modulation efficiency, bandwidth, and device aperture. Here we report a hybrid BaTiO3 (BTO)/TiO2 metasurface for large-aperture, efficient, gigahertz-speed free-space electro-optic modulation. Combining scalable BTO film growth by radio-frequency magnetron sputtering with mature TiO2 nanofabrication, we pattern the metasurface in TiO2 on an unetched BTO layer. The resulting devices support guided-mode resonances with quality factors exceeding 1300 and an optical confinement factor of ~0.8, while the continuous BTO layer makes efficient use of the applied voltage, together maximizing the overlap between the optical and driving fields within the BTO. A device with a 0.3 mm x 0.3 mm metasurface achieves a transmittance modulation efficiency of ~0.020 per volt and a -3 dB electro-optic bandwidth of ~0.8 GHz, with an effective Pockels coefficient of ~151 pm/V for the BTO. This establishes a scalable route to high-performance free-space electro-optic modulators for LiDAR, free-space optical communication, and reconfigurable optical computing.

physics.optics

Ultrafast programmable Bragg reflection in photonic integrated circuits

Distributed Bragg reflectors (DBRs) are foundational building blocks of classical and quantum photonic technologies. However, their optical responses are typically fixed upon fabrication, limiting circuit robustness, reconfigurability, and functionality in applications from high-speed communications to quantum computing. Here, we demonstrate photonic chip-based programmable DBRs at telecommunications wavelengths, which are formed by electro-optically inducing refractive index contrast between periodic ferroelectric domains in thin-film lithium niobate waveguides. We achieve voltage-controlled Bragg reflection from zero to near-unity, and gigahertz-speed reflectivity modulation. Our results bring DBRs into the ultrafast programmable regime, opening new opportunities in topological photonics, cavity quantum electrodynamics, integrated lasers, and optical interconnects. The interplay between nanoscale ferroelectric domain engineering and strong electro-optic nonlinearity establishes a new design strategy for nanophotonic devices, otherwise inaccessible in bulk media.

physics.optics

Semismooth Newton methods for degenerate polyhedral projection

In this paper, we study dual semismooth Newton (SSN) methods for degenerate polyhedral projection problems, where generalized Jacobians of the dual residual may remain singular even arbitrarily close to the solution set. Rather than regularizing these singular systems, we exploit the nonuniqueness of the dual representation. We introduce a primal--dual lifted projection-equivalent set that always possesses extreme points without additional structural assumptions on the polyhedron, and show that its extreme-point geometry identifies dual representatives at which nonsingular generalized Jacobians of the dual residual can be constructed. This geometry is further linked to a full-column-rank condition and a generalized weak strict Robinson constraint qualification, showing that the regularity required by the Newton step can be recovered rather than imposed \emph{a priori}. We also establish displacement bounds that connect representative selection throughout the algorithm with the local Newton mechanism. Building on this variational framework, we develop an inexact dual SSN method with local superlinear convergence and a globalized version combining monotone representative selection with a Wolfe line search. The resulting method is globally convergent and eventually recovers the fast local rate. Numerical experiments on regularized optimal transport, battery-scheduling feasibility restoration, and occupation-measure projection demonstrate its robustness in highly degenerate settings.

math.OC

Maximal monotonicity of piecewise polyhedral mappings

Maximal monotone mappings that are piecewise polyhedral arise from subdifferentials of convex functions and saddle functions that are piecewise linear-quadratic and enter into algorithmic constructions of importance in linear-quadratic optimization and associated splitting methods. The question of whether those constructions preserve maximal monotonicity is then crucial. The usual answers to that invoke constraint qualifications involving the nonemptiness of intersections of certain relative interiors, but it is shown here that the piecewise polyhedral structure allows the relative interiors to be bypassed.

math.OC

A Semismooth Newton Augmented Lagrangian Method for Sparse Spectral Risk Optimization

Empirical risk minimization is a standard and effective paradigm for learning predictive models by minimizing average loss. In high-stakes decision-making, however, an average-loss criterion may underrepresent rare but severe losses. Spectral risk measures (SRMs) provide a principled framework by incorporating weighted order statistics of losses, but the induced nonsmoothness and nonseparability from sorting make the resulting optimization problems challenging. We propose a relative inexact proximal augmented Lagrangian method with a semismooth Newton subproblem solver for solving SRM-based optimization problems. Exploiting a dual reformulation and properties of the Moreau envelope, we reduce the subproblems to structured dual-variable formulations, significantly simplifying computation. We provide explicit generalized Jacobian characterizations and tailor the pool adjacent violators algorithm for their efficient evaluation. Numerical results on synthetic and real-data instances show that the proposed method attains lower running times than the tested ADMM baseline while producing comparable stationarity residuals and sparse solutions.

math.OC

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognition. Existing methods typically build such representations through textual reasoning or 3D reconstruction. However, they often falter during multi-step reasoning, particularly when required to dynamically re-anchor evidence to the specific camera-, object-, or direction-centric reference frames demanded by complex queries. To address this, we propose OmniView-Space, a framework designed to maintain spatial consistency through multimodal egocentric evidence. Our approach consists of three core components: (1) Multi-Perspective Spatial Mapping (MPSM), which re-anchors reconstructed geometry into a query-aligned visual cognitive map and a textual spatial graph; (2) Tool-Guided Egocentric Reasoning, an interleaved policy trained to actively select the ego anchor required by the query and request the corresponding MPSM evidence; and (3) Cognitive-Map Distillation, which uses MPSM-generated trajectories and ego-frame rewards to train the model to reason with self-generated cognitive maps. Experiments on single- and multi-image spatial reasoning benchmarks show that OmniView-Space achieves state-of-the-art performance. Furthermore, the distilled model maintains this performance while reducing reliance on external geometry pipelines.

cs.CV

Peer-to-Peer Cloud Service Market for Data Centers Oriented to Computation-Electricity Coordination

Energy-intensive data centers (DCs) have emerged as substantial and flexible loads in modern power systems, underscoring the critical need for computation-electricity coordination. Harnessing the spatio-temporal flexibility of DC workloads is a promising approach to facilitate this coordination. However, existing studies overlook the collaborative potential of computational resource sharing among geo-distributed DCs, thereby failing to fully unlock this flexibility. In this paper, a bi-level computation-electricity coordination framework is proposed to explicitly capture the bidirectional interactions between DCs and power grid. Firstly, a peer-to-peer cloud service market (P2P-CSM) for geo-distributed DCs is proposed, which enables bilateral cloud service transactions to leverage regional heterogeneities (e.g., electricity prices, cooling efficiency). Secondly, locational marginal prices are embedded into the framework to reflect network congestion and nodal price disparities. Thirdly, a dual consensus alternating direction method of multipliers (ADMM)-based decentralized algorithm is developed as the P2P market clearing algorithm, and a bisection-assisted iterative algorithm is proposed to ensure rigorous convergence of the framework. Case studies conducted on modified IEEE 30-bus system validate that the P2P-CSM achieves a win-win computation-electricity coordination: it not only increases total DC operational profit by 22.8\%, but also effectively alleviates grid congestion and yields a 3.2\% reduction in total energy consumption.

eess.SY

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning

Tool-integrated reasoning (TIR) enables large language models (LLMs) to enhance their capabilities by interacting with external tools, such as code interpreters (CI). Most recent studies focus on exploring various methods to equip LLMs with the ability to use tools. However, how to further boost the reasoning ability of already tool-capable LLMs at inference time remains underexplored. Improving reasoning at inference time requires no additional training and can help LLMs better leverage tools to solve problems. We observe that, during tool-capable LLM inference, both the number and the proportion of erroneous tool calls are negatively correlated with answer correctness. Moreover, erroneous tool calls are typically resolved successfully within a few subsequent turns. If not, LLMs often struggle to resolve such errors even with many additional turns. Building on the above observations, we propose PruneTIR, a rather effective yet efficient framework that enhances the tool-integrated reasoning at inference time. During LLM inference, PruneTIR prunes trajectories, resamples tool calls, and suspends tool usage through three components: Success-Triggered Pruning, Stuck-Triggered Pruning and Resampling, and Retry-Triggered Tool Suspension. These three components enable PruneTIR to mitigate the negative impact of erroneous tool calls and prevent LLMs from getting stuck in repeated failed resolution attempts, thereby improving overall LLM performance. Extensive experimental results demonstrate the effectiveness of PruneTIR, which significantly improves Pass@1 and efficiency while reducing the working context length for tool-capable LLMs.

cs.CL

Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demands not only accurate scores but also actionable, localized feedback. However, current MLLM-based methods adopt a single-look, language-only paradigm, which departs from human evidence-seeking judgment and yields weakly grounded rationales, limiting their reliability for in-the-loop refinement. We propose Q-DeepSight, a think-with-image framework that emulates this human-like process. It performs interleaved Multimodal Chain-of-Thought (iMCoT) with tool-augmented evidence acquisition (e.g., crop-and-zoom) to explicitly determine where quality degrades and why. To train these long iMCoT trajectories via reinforcement learning, we introduce two techniques: Perceptual Curriculum Reward (PCR) to mitigate reward sparsity and Evidence Gradient Filtering (EGF) to improve credit assignment for visually-grounded reasoning. Q-DeepSight achieves state-of-the-art performance across diverse benchmarks, including natural, restored, and AI-generated content. Furthermore, we demonstrate its practical value with Perceptual-in-Generation (PiG), a training-free framework where Q-DeepSight's diagnoses guide iterative image enhancement, effectively closing the loop between assessment and refinement.

cs.CV

PLANING: A Loosely Coupled Triangle-Gaussian Framework for Streaming 3D Reconstruction

Streaming reconstruction from monocular image sequences remains challenging, as existing methods typically favor either high-quality rendering or accurate geometry, but rarely both. We present PLANING, an efficient on-the-fly reconstruction framework built on a hybrid representation that loosely couples explicit geometric primitives with neural Gaussians, enabling geometry and appearance to be modeled in a decoupled manner. This decoupling supports an online initialization and optimization strategy that separates geometry and appearance updates, yielding stable streaming reconstruction with substantially reduced structural redundancy. PLANING improves dense mesh Chamfer-L2 by 18.52% over PGSR, surpasses ARTDECO by 1.31 dB PSNR, and reconstructs ScanNetV2 scenes in under 100 seconds, over 5x faster than 2D Gaussian Splatting, while matching the quality of offline per-scene optimization. Beyond reconstruction quality, the structural clarity and computational efficiency of PLANING make it well suited for a broad range of downstream applications, such as enabling large-scale scene modeling and simulation-ready environments for embodied AI. Project page: https://city-super.github.io/PLANING/ .

cs.CV

ActiShade: Activating Overshadowed Knowledge to Guide Multi-Hop Reasoning in Large Language Models

In multi-hop reasoning, multi-round retrieval-augmented generation (RAG) methods typically rely on LLM-generated content as the retrieval query. However, these approaches are inherently vulnerable to knowledge overshadowing - a phenomenon where critical information is overshadowed during generation. As a result, the LLM-generated content may be incomplete or inaccurate, leading to irrelevant retrieval and causing error accumulation during the iteration process. To address this challenge, we propose ActiShade, which detects and activates overshadowed knowledge to guide large language models (LLMs) in multi-hop reasoning. Specifically, ActiShade iteratively detects the overshadowed keyphrase in the given query, retrieves documents relevant to both the query and the overshadowed keyphrase, and generates a new query based on the retrieved documents to guide the next-round iteration. By supplementing the overshadowed knowledge during the formulation of next-round queries while minimizing the introduction of irrelevant noise, ActiShade reduces the error accumulation caused by knowledge overshadowing. Extensive experiments show that ActiShade outperforms existing methods across multiple datasets and LLMs.

cs.CL

Scalable Kernel Quantile Regression: A Preconditioned Augmented Lagrangian Method

Kernel quantile regression (KQR) extends classical quantile regression to nonlinear settings using kernel methods, offering a powerful tool for modeling conditional distributions. However, its application to large-scale datasets remains challenging due to two intrinsic difficulties: the nonsmoothness of the quantile check loss and the computational burden imposed by the large, dense kernel matrix. Existing state-of-the-art solvers often struggle to handle both challenges simultaneously, leading to limited scalability and high computational cost. In this paper, we propose PALM-KQR, a highly efficient two-phase preconditioned augmented Lagrangian method for large-scale KQR. In the first phase, an inexact alternating direction method of multipliers (ADMM) is employed to compute a warm-start solution efficiently. The second phase refines this solution using an efficient semismooth Newton augmented Lagrangian method (ALM). Our key innovations include a dual semismooth Newton approach for handling the nonsmooth quantile check loss, and a specialized preconditioning strategy based on low-rank approximations of the kernel matrix, which exploits its structure and mitigates ill-conditioning in the linear systems arising in ALM, thereby significantly accelerating the iterative solvers. Extensive numerical experiments demonstrate that PALM-KQR substantially outperforms existing commercial and specialized KQR solvers in both efficiency and scalability.

math.OC

On the B-subdifferential of proximal operators of affine-constrained $\ell_1$ regularizer

In this work, we study the affine-constrained $\ell_1$ regularizers, which frequently arise in statistical and machine learning problems across a variety of applications, including microbiome compositional data analysis and sparse subspace clustering. With the aim of developing scalable second-order methods for solving optimization problems involving such regularizers, we analyze the associated proximal mapping and characterize its generalized differentiability, with a focus on its B-subdifferential. The revealed structured sparsity in the B-subdifferential enables us to design efficient algorithms within the proximal point framework. Extensive numerical experiments on real applications, including comparisons with state-of-the-art solvers, further demonstrate the superior performance of our approach. Our findings provide new insights into the sensitivity and stability properties of affine-constrained nonsmooth regularizers, and contribute to the development of fast second-order methods for a class of structured, constrained sparse learning problems.

math.OC

Spectral-temporal processing using integrated recursive electro-optic circuit

Advances in integrated photonics have enabled unprecedented level of control of light, powering a wide range of photonic technologies from communications and computing to precision metrology and quantum information. However, the conventional on-chip optical signal processing approaches based on optical waveguides and cavities suffer from their large physical footprint and narrow operating bandwidth, respectively. To address this, we propose and experimentally demonstrate, using the thin film lithium niobate (TFLN) photonics, a modular, recursive optical signal processing framework. In our approach, fast electro-optic (EO) switch is used to either keep an optical packet inside the loop, with the processing element embedded within, or to release it from the loop and direct it towards the output waveguide. By configuring the switch on a timescale shorter than a time the photon spends traversing the loop, this architecture can achieve different optical pathlengths in a compact footprint without sacrificing optical bandwidth. As an example, we embed a phase-modulator (PM) inside the loop and demonstrate a frequency shift of optical packets up to 420 GHz, using only a 3 GHz sinusoidal microwave signal. By replacing the PM with chirped Bragg gratings (CBG), we demonstrate recursive delay line featuring large group delays of 28 ps/nm over a 30 nm optical bandwidth. Finally, by introducing asymmetric Mach-Zehnder interferometer (AMZI) inside the loop, we demonstrate a reconfigurable differentiation of the optical packet in time, up to an unprecedented fifth-order. Our results establish a powerful and scalable platform for multifunctional photonic processing, setting the stage for next-generation integrated systems with ultrahigh reconfigurability and spectral-temporal versatility.

physics.optics

Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs

Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model. However, much less is known about the constrained average-reward MDP (CAMDP), where policies must satisfy long-run average constraints. In this work, we address this gap by studying the sample complexity of learning an $\epsilon$-optimal policy in CAMDPs under a generative model. We propose a model-based algorithm that operates under two settings: (i) relaxed feasibility, which allows small constraint violations, and (ii) strict feasibility, where the output policy satisfies the constraint. We show that our algorithm achieves sample complexities of $\tilde{O}\left(\frac{S A (B+H)}{ \epsilon^2}\right)$ and $\tilde{O} \left(\frac{S A (B+H)}{\epsilon^2 \zeta^2} \right)$ under the relaxed and strict feasibility settings, respectively. Here, $\zeta$ is the Slater constant indicating the size of the feasible region, $H$ is the span bound of the bias function, and $B$ is the transient time bound. Moreover, a matching lower bound of $\tilde{\Omega}\left(\frac{S A (B+H)}{ \epsilon^2\zeta^2}\right)$ for the strict feasibility case is established, thus providing the first minimax-optimal bounds for CAMDPs. Our results close the theoretical gap in understanding the complexity of constrained average-reward MDPs.

cs.LG