SearcharxivSearch

arXiv subjects

Xi Huang

Publications and source records attributed to Xi Huang.

At least 19 recordsLinked to original sources

Well-posedness of stochastic time-nonlocal telegraph equations with H\"{o}lder diffusion coefficient: hereditary phase-space lifting and novel generalized coupling method

We consider the initial-boundary value problem for the stochastic time-nonlocal telegraph equation with $(\mathcal{PC}_\varepsilon)$-type kernel $a$: \begin{align*} \gamma \partial_t \left( a \ast \partial_t (a \ast v)\right) =\Delta v-\partial_t (a \ast v)+ \Psi(v)+ \Phi(v) \frac{\mathrm{d}W(t)}{\mathrm{d}t}, \end{align*} where $W$ is a space-time Gaussian white noise, $\Psi$ satisfies a linear growth condition, and $\Phi$ is H\"{o}lder continuous and uniformly nondegenerate. This model characterizes high-frequency signal propagation in small-scale systems under stochastic fluctuations. We develop a new hereditary phase-space lifting framework for time-nonlocal telegraph equations. In addition, we propose a novel generalized coupling framework, which features a new construction of the damping control term for the velocity. Based on these analytic tools, we prove the first results on weak existence and uniqueness in law for mild solutions, valued in $L_{loc}^2(\mathbb R_+; H^{\delta})$, to the stochastic nonlocal telegraph equation. The regularity index $\delta$ can be arbitrarily close to $\min\{\frac{1}{2},\frac{\varepsilon}{2+\varepsilon} \}$ from below, and the admissible lower bound of the H\"{o}lder exponent $\kappa$ is quantitatively determined by the integrability exponent $\varepsilon$. For the corresponding IBVP of the stochastic damped wave equation, obtained by replacing $a$ with the Dirac measure $\delta_0$, $\kappa$ can be improved to any value in $(\frac{3}{5}, 1]$. More significantly, the generalized coupling framework also handles low-regularity nonlinearities depending on both displacement and velocity.

math.AP

Non-local evolution equations with L\'{e}vy diffusion: Well-posedness and limiting behavior

In this note we focus our attention on a class of nonlocal-in-time evolution equations with L\'{e}vy diffusion, they arise as models of unidirectional viscoelastic fluid flow and physical phenomena with memory effect.We first consider the existence of the classical solution to a nonlocal linear evolution problem under conditions on the involved memory kernels which allows complete positivity. Then we investigate the limit of this model to a generalized Rayleigh-Stokes equation, as the index of L\'{e}vy diffusion gets concentrated near two, we prove that the solution of nonlocal-in-time problem with L\'{e}vy diffusion uniformly converges to that of the generalized Rayleigh-Stokes equation and reveal the convergence rate.Finally, the existence and limiting behavior of the mild solution to a nonlocal evolution problem with nonlinearity are established. The proofs are based on subordination principle and relaxation function theory.

math.AP

The existence and nonexistence for time nonlocal evolution equations with superlinear sources

We consider the solvable behaviour for the initial value problem of time nonlocal evolution equations with the kernels of type $(\mathcal{PC})$. Our aim is to analyze some sufficient conditions ensuring local existence, integrability of mild solutions when the nonlinear term exhibits rapid growth. Moreover, a sufficient condition of the unsolvable result is also established. It turns out that the solvable behaviour is closely connected with the index on the initial value, which occurs a critical dimension phenomenon. The proofs rely on subordination and monotone iterative method. Finally, several examples are given to illustrate the wide applicability of the results.

math.AP

2026 Roadmap on Artificial Intelligence and Machine Learning for Smart Manufacturing

The evolution of artificial intelligence (AI) and machine learning (ML) is reshaping smart manufacturing by providing new capabilities for efficiency, adaptability, and autonomy across industrial value chains. However, the deployment of AI and ML in industrial settings still faces critical challenges, including the complexity of industrial big data, effective data management, integration with heterogeneous sensing and control systems, and the demand for trustworthy, explainable, and reliable operation in high-stakes industrial environments. In this roadmap, we present a comprehensive perspective on the foundations, applications, and emerging directions of AI and ML in smart manufacturing. It is structured in three parts. The first highlights the foundations and trends that frame the evolution of AI in smart manufacturing. The second focuses on key topics where AI is already enabling advances, including industrial big data analytics, advanced sensing and perception, autonomous systems, additive and laser-based manufacturing, digital twins, robotics, supply chain and logistics optimization, and sustainable manufacturing. The third section explores non-traditional ML approaches that are opening new frontiers, such as physics-informed AI, generative AI, semantic AI, advanced digital twins, explainable AI, RAMS, data-centric metrology, LLMs, and foundation models for highly connected and complex manufacturing systems. By identifying both opportunities and remaining barriers across these areas, this roadmap outlines the advances needed in methods, integration strategies, and industrial adoption. We hope this roadmap will serve as a guide for researchers, engineers, and practitioners to accelerate innovation, align academic and industrial priorities, and ensure that AI-driven smart manufacturing delivers reliable, sustainable, and scalable impact for the future of manufacturing ecosystems.

cs.AI

Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions

Collaborative learning has emerged as a key paradigm in large-scale intelligent systems, enabling distributed agents to cooperatively train their models while addressing their privacy concerns. Central to this paradigm is knowledge distillation (KD), a technique that facilitates efficient knowledge transfer among agents. However, the underlying mechanisms by which KD leverages memory and knowledge across agents remain underexplored. This paper aims to bridge this gap by offering a comprehensive review of KD in collaborative learning, with a focus on the roles of memory and knowledge. We define and categorize memory and knowledge within the KD process and explore their interrelationships, providing a clear understanding of how knowledge is extracted, stored, and shared in collaborative settings. We examine various collaborative learning patterns, including distributed, hierarchical, and decentralized structures, and provide insights into how memory and knowledge dynamics shape the effectiveness of KD in collaborative learning. Particularly, we emphasize task heterogeneity in distributed learning pattern covering federated learning (FL), multi-agent domain adaptation (MADA), federated multi-modal learning (FML), federated continual learning (FCL), federated multi-task learning (FMTL), and federated graph knowledge embedding (FKGE). Additionally, we highlight model heterogeneity, data heterogeneity, resource heterogeneity, and privacy concerns of these tasks. Our analysis categorizes existing work based on how they handle memory and knowledge. Finally, we discuss existing challenges and propose future directions for advancing KD techniques in the context of collaborative learning.

cs.DC

PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

Robotic manipulation systems benefit from complementary sensing modalities, where each provides unique environmental information. Point clouds capture detailed geometric structure, while RGB images provide rich semantic context. Current point cloud methods struggle to capture fine-grained detail, especially for complex tasks, which RGB methods lack geometric awareness, which hinders their precision and generalization. We introduce PointMapPolicy, a novel approach that conditions diffusion policies on structured grids of points without downsampling. The resulting data type makes it easier to extract shape and spatial relationships from observations, and can be transformed between reference frames. Yet due to their structure in a regular grid, we enable the use of established computer vision techniques directly to 3D data. Using xLSTM as a backbone, our model efficiently fuses the point maps with RGB data for enhanced multi-modal perception. Through extensive experiments on the RoboCasa and CALVIN benchmarks and real robot evaluations, we demonstrate that our method achieves state-of-the-art performance across diverse manipulation tasks. The overview and demos are available on our project page: https://point-map.github.io/Point-Map/

cs.RO

Continuous Variable Hamiltonian Learning at Heisenberg Limit via Displacement-Random Unitary Transformation

Characterizing continuous-variable (CV) Hamiltonians can be formulated as Hamiltonian learning under quantum measurement constraints: finite operator coefficients are inferred from noisy measurement outcomes obtained by probing an infinite-dimensional system. Existing Heisenberg-limited CV protocols are often limited to low-order structures, vulnerable to noise, or unresolved for generic multi-mode settings. We introduce Displacement-Random Unitary Transformation (D-RUT), an active data acquisition protocol with pre-specified probes and number-preserving transformations that reduce finite-order bosonic Hamiltonian learning to polynomial recovery. We prove Heisenberg-limited total evolution time with robustness to state preparation and measurement (SPAM) errors, and develop hierarchical multi-mode coefficient recovery with better statistical efficiency than simultaneous estimation. We also extend D-RUT to first-quantized Hamiltonian coefficient learning, and numerical experiments on single- and multi-mode nonlinear systems validate the predicted Heisenberg scaling.

quant-ph

Random data Cauchy theory for fully nonlocal telegraph equations

We consider the random Cauchy problem for the fully nonlocal telegraph equation of power type with the general $(\mathcal{PC}^{\ast})$ type kernel $(a,b)$. This equation can effectively characterize high-frequency signal transmission in small-scale systems. We establish a new completely positive kernel induced by $b$ (see Appendix \refeq{app b}) and derive two novel solution operators by using the relaxation functions associated with the new kernel,which are closely related to the operators $\cos(\theta(-\Delta)^{\frac{\beta}{4}} )$ and $(-\Delta)^{-\frac{\beta}{4} }\sin(\theta(-\Delta)^{\frac{\beta}{4}} )$ for $\beta\in(1,2]$. These operators enable, for the first time, the derivation of mixed-norm $L_t^qL_x^{p'}$ estimates for the novel solution operators. Next, utilizing probabilistic randomization methods, we establish the average effects, the local existence and uniqueness for a large set of initial data $u^\omega \in L^{2}(\Omega, H^{s,p}(\mathbb R^3))$ ($p\in (1,2)$) while also obtaining probabilistic estimates for local existence under randomized initial conditions. The results reveal a critical phenomenon in the temporal regularity of the solution regarding the regularity index $s$ of the initial data $u^\omega$.

math.AP

Fourier transform-based linear combination of Hamiltonian simulation

Linear combination of Hamiltonian simulation (LCHS) connects the general linear non-unitary dynamics with unitary operators and serves as the mathematical backbone of designing near-optimal quantum linear differential equation algorithms. However, the existing LCHS formalism needs to find a kernel function subject to complicated technical conditions on a half complex plane. In this work, we establish an alternative formalism of LCHS based on the Fourier transform. Our new formalism completely removes the technical requirements beyond the real axis, providing a simple and flexible way of constructing LCHS kernel functions. Specifically, we construct a different family of the LCHS kernel function, providing a $1.81$ times reduction in the quantum differential equation algorithms based on LCHS, and an $8.27$ times reduction in its quantum circuit depth at a truncation error of $\epsilon \le 10^{-8}$. Additionally, we extend the scope of the LCHS formula to the scenario of simulating linear unstable dynamics for a short or intermediate time period.

quant-ph

4D-PreNet: A Unified Preprocessing Framework for 4D-STEM Data Analysis

Automated experimentation with real time data analysis in scanning transmission electron microscopy (STEM) often require end-to-end framework. The four-dimensional scanning transmission electron microscopy (4D-STEM) with high-throughput data acquisition has been constrained by the critical bottleneck results from data preprocessing. Pervasive noise, beam center drift, and elliptical distortions during high-throughput acquisition inevitably corrupt diffraction patterns, systematically biasing quantitative measurements. Yet, conventional correction algorithms are often material-specific and fail to provide a robust, generalizable solution. In this work, we present 4D-PreNet, an end-to-end deep-learning pipeline that integrates attention-enhanced U-Net and ResNet architectures to simultaneously perform denoising, center correction, and elliptical distortion calibration. The network is trained on large, simulated datasets encompassing a wide range of noise levels, drift magnitudes, and distortion types, enabling it to generalize effectively to experimental data acquired under varying conditions. Quantitative evaluations demonstrate that our pipeline reduces mean squared error by up to 50% during denoising and achieves sub-pixel center localization in the center detection task, with average errors below 0.04 pixels. The outputs are bench-marked against traditional algorithms, highlighting improvements in both noise suppression and restoration of diffraction patterns, thereby facilitating high-throughput, reliable 4D-STEM real-time analysis for automated characterization.

cs.CV

MoRe-ERL: Learning Motion Residuals using Episodic Reinforcement Learning

We propose MoRe-ERL, a framework that combines Episodic Reinforcement Learning (ERL) and residual learning, which refines preplanned reference trajectories into safe, feasible, and efficient task-specific trajectories. This framework is general enough to incorporate into arbitrary ERL methods and motion generators seamlessly. MoRe-ERL identifies trajectory segments requiring modification while preserving critical task-related maneuvers. Then it generates smooth residual adjustments using B-Spline-based movement primitives to ensure adaptability to dynamic task contexts and smoothness in trajectory refinement. Experimental results demonstrate that residual learning significantly outperforms training from scratch using ERL methods, achieving superior sample efficiency and task performance. Hardware evaluations further validate the framework, showing that policies trained in simulation can be directly deployed in real-world systems, exhibiting a minimal sim-to-real gap.

cs.RO

4D-MISR: A unified model for low-dose super-resolution imaging via feature fusion

While electron microscopy offers crucial atomic-resolution insights into structure-property relationships, radiation damage severely limits its use on beam-sensitive materials like proteins and 2D materials. To overcome this challenge, we push beyond the electron dose limits of conventional electron microscopy by adapting principles from multi-image super-resolution (MISR) that have been widely used in remote sensing. Our method fuses multiple low-resolution, sub-pixel-shifted views and enhances the reconstruction with a convolutional neural network (CNN) that integrates features from synthetic, multi-angle observations. We developed a dual-path, attention-guided network for 4D-STEM that achieves atomic-scale super-resolution from ultra-low-dose data. This provides robust atomic-scale visualization across amorphous, semi-crystalline, and crystalline beam-sensitive specimens. Systematic evaluations on representative materials demonstrate comparable spatial resolution to conventional ptychography under ultra-low-dose conditions. Our work expands the capabilities of 4D-STEM, offering a new and generalizable method for the structural analysis of radiation-vulnerable materials.

cs.CV

BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning

We present the B-spline Encoded Action Sequence Tokenizer (BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separate tokenizer training and consistently produces tokens of uniform length, enabling fast action sequence generation via parallel decoding. Leveraging our B-spline formulation, BEAST inherently ensures generating smooth trajectories without discontinuities between adjacent segments. We extensively evaluate BEAST by integrating it with three distinct model architectures: a Variational Autoencoder (VAE) with continuous tokens, a decoder-only Transformer with discrete tokens, and Florence-2, a pretrained Vision-Language Model with an encoder-decoder architecture, demonstrating BEAST's compatibility and scalability with large pretrained models. We evaluate BEAST across three established benchmarks consisting of 166 simulated tasks and on three distinct robot settings with a total of 8 real-world tasks. Experimental results demonstrate that BEAST (i) significantly reduces both training and inference computational costs, and (ii) consistently generates smooth, high-frequency control signals suitable for continuous control tasks while (iii) reliably achieves competitive task success rates compared to state-of-the-art methods.

cs.RO

Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)

The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with adaptive, time-sensitive video retrieval. This paper introduces a novel framework that combines vector similarity search with graph-based data structures. By leveraging VLM embeddings for initial retrieval and modeling contextual relationships among video segments, our approach enables adaptive query refinement and improves retrieval accuracy. Experiments demonstrate its precision, scalability, and robustness, offering an effective solution for interactive video retrieval in dynamic environments.

cs.CV

Towards Fusing Point Cloud and Visual Representations for Imitation Learning

Learning for manipulation requires using policies that have access to rich sensory information such as point clouds or RGB images. Point clouds efficiently capture geometric structures, making them essential for manipulation tasks in imitation learning. In contrast, RGB images provide rich texture and semantic information that can be crucial for certain tasks. Existing approaches for fusing both modalities assign 2D image features to point clouds. However, such approaches often lose global contextual information from the original images. In this work, we propose FPV-Net, a novel imitation learning method that effectively combines the strengths of both point cloud and RGB modalities. Our method conditions the point-cloud encoder on global and local image tokens using adaptive layer norm conditioning, leveraging the beneficial properties of both modalities. Through extensive experiments on the challenging RoboCasa benchmark, we demonstrate the limitations of relying on either modality alone and show that our method achieves state-of-the-art performance across all tasks.

cs.RO

X-IL: Exploring the Design Space of Imitation Learning Policies

Designing modern imitation learning (IL) policies requires making numerous decisions, including the selection of feature encoding, architecture, policy representation, and more. As the field rapidly advances, the range of available options continues to grow, creating a vast and largely unexplored design space for IL policies. In this work, we present X-IL, an accessible open-source framework designed to systematically explore this design space. The framework's modular design enables seamless swapping of policy components, such as backbones (e.g., Transformer, Mamba, xLSTM) and policy optimization techniques (e.g., Score-matching, Flow-matching). This flexibility facilitates comprehensive experimentation and has led to the discovery of novel policy configurations that outperform existing methods on recent robot learning benchmarks. Our experiments demonstrate not only significant performance gains but also provide valuable insights into the strengths and weaknesses of various design choices. This study serves as both a practical reference for practitioners and a foundation for guiding future research in imitation learning.

cs.RO

ETA-IK: Execution-Time-Aware Inverse Kinematics for Dual-Arm Systems

This paper presents ETA-IK, a novel Execution-Time-Aware Inverse Kinematics method tailored for dual-arm robotic systems. The primary goal is to optimize motion execution time by leveraging the redundancy of both arms, specifically in tasks where only the relative pose of the robots is constrained, such as dual-arm scanning of unknown objects. Unlike traditional inverse kinematics methods that use surrogate metrics such as joint configuration distance, our method incorporates direct motion execution time and implicit collisions into the optimization process, thereby finding target joints that allow subsequent trajectory generation to get more efficient and collision-free motion. A neural network based execution time approximator is employed to predict time-efficient joint configurations while accounting for potential collisions. Through experimental evaluation on a system composed of a UR5 and a KUKA iiwa robot, we demonstrate significant reductions in execution time. The proposed method outperforms conventional approaches, showing improved motion efficiency without sacrificing positioning accuracy. These results highlight the potential of ETA-IK to improve the performance of dual-arm systems in applications, where efficiency and safety are paramount.

cs.RO