SearcharxivSearch

arXiv subjects

Lai Jiang

Publications and source records attributed to Lai Jiang.

At least 19 recordsLinked to original sources

Video2DoorTraversal: Push Door Traversal via Simulated Door Twins

Door opening and traversal is a long-horizon loco-manipulation task that requires precise handle interaction and coordinated base-arm control. We present Video2DoorTraversal, a single-video real-to-sim-to-real framework for wheel-legged mobile manipulators. Given one RGB video of a real door, DoorTwin reconstructs an instance-aligned, articulated, and simulation-ready door twin with realistic geometry and appearance. A simulation-in-the-loop agent converts the recovered articulation into a parameterized skill program and iteratively refines failed rollouts to generate physically executable demonstrations. These demonstrations are used to train ArticuACT, a dual-depth policy that predicts coordinated base, arm, and gripper commands using robot-centric camera conditioning and interaction-aware supervision. With all perception and policy inference running onboard, the system achieves a 96.57% average success rate across five real doors and an 80.95% zero-shot success rate on structurally similar unseen doors, while completing the full approach, opening, and traversal sequence in approximately 13s on average. Project Page: https://video2doortraversal.github.io/.

cs.RO

A broadband, individually addressing two- and three-dimensional photonic integrated circuit for trapped-ion qubit control

Trapped ions provide a high-fidelity platform for quantum information processing, yet delivery of multiple, distinct wavelengths across large networks of interaction zones remains a bottleneck. Conventional free-space light delivery lacks scalability, while on-chip grating couplers suffer from narrow operational bandwidth that increases circuit footprint and optical interfacing complexity. Here we show a broadband photonic integrated circuit capable of addressing individual ions. The circuit combines a planar waveguide lens with a micromirror fabricated using two-photon polymerization at wafer scale. This implementation can address three individual ions from $\lambda$ = 405 - 880 nm with -27 dB average intensity crosstalk at $5\,\mu\mathrm{m}$ pitch. We trap $^{40}\mathrm{Ca}^{+}$ and $^{138}\mathrm{Ba}^{+}$ ions above such devices, characterize optical crosstalk with barium ions, and demonstrate individual repumping of calcium ions. This monolithic photonic architecture brings broadband addressing in an on-chip modality to trapped-ion technology. More generally, integrating additive manufacturing into quantum devices is poised to unlock expanded design space for implementing novel quantum architectures.

quant-ph

Dimension drop for intersections of Cantor sets

Let $E\subset \mathbb{R}$ be a self-similar set generated by a homogeneous iterated function system $\Phi$ with contraction ratio $\rho\in (0,1)$. Assume that $\Phi$ satisfies the open set condition and $\dim_{\rm H}E<1$. Let $f$ be a $C^1$-diffeomorphism on $\mathbb{R}$. We prove that if $\log|f'(x)|/\log\rho\not\in\mathbb{Q}$ for every $x\in E\cap f^{-1}(E)$, then the upper Minkowski dimension of $f(E)\cap E$ is strictly less than the Hausdorff dimension of $E$. We also establish a quantitative dimension drop result when $E$ is a missing-digit set and $f$ is an affine map with rational slope satisfying a certain arithmetic condition. Based on these results and a result of Shmerkin [Ann. of Math., 2019], we obtain characterizations of $\gamma$ in various contexts such that $\overline{\dim}_{\rm M}\left((\gamma E+\alpha)\cap E\right)<\dim_{\rm H}E$ for every $\alpha\in\mathbb{R}$.

math.CA

Minkowski geometry of finite Hurwitz continued fractions

We study the Minkowski geometry of finite-level sets of Gaussian rationals defined by the lengths of their Hurwitz continued fraction expansions. For each $m\geq 1$, let $H_m$ be the set of points in the fundamental square whose Hurwitz continued fraction expansions have length exactly $m$. We also introduce the relaxed recursive sets defined by $G_0=\{0\}$ and $$G_m=\Big\{\frac{1}{u+v}: u \in\mathbb{Z}[i],\ v\in G_{m-1},\ |u+v|>1 \Big\}.$$ We prove that for every $m\geq 1$, $$\dim_{\rm M} H_m=\dim_{\rm M} G_m=1.$$ We further determine the critical one-dimensional Minkowski content of these sets. We have ${\mathcal M}^1(H_1)={\mathcal M}^1(G_1)=4\pi\log(1+\sqrt{2})$, whereas ${\mathcal M}^1(H_m)={\mathcal M}^1(G_m)=\infty$ for every $m\geq 2$.

math.CA

FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On

Image-based Virtual Try-On (IVTON) has greatly advanced through diffusion models, yet existing methods require many sampling steps and depend on masks with costly auxiliary networks. In addition, the absence of large-scale mask-free paired datasets further limits the development of mask-free IVTON. We propose FDM-MFVT, a few-step diffusion model for mask-free IVTON, integrating an Outfit-aware Noise Optimization Module (OANO) and an Instruction-driven Try-on Module (IDT) to enhance efficiency and flexibility.The OANO module initializes the alignment space with noise using the input image and only needs 6 steps to generate a higher-fidelity try-on image compared to 30 steps.The IDT module uses virtual try-on prompts and efficient adaptation to generate high-quality results from garment and person images alone. We further introduce MFVT, a 30,000-pair mask-free IVTON dataset. Experiments show that FDM-MFVT achieves superior quantitative and qualitative results with fewer inference steps than mask-based and mask-free baseline methods.

cs.CV

StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots

Climbing hollow stairs remains a challenging problem for quadruped robots due to the high risk of leg trapping, severe depth sparsity, and high-frequency depth-sensing noise. In this paper, we propose StairMaster, a novel three-stage reinforcement learning framework for stable locomotion on such extreme discontinuous terrains. Our architecture integrates a Cross-Attention mechanism to extract structural features from noisy depth data, alongside a Spatial-aware Recurrent Unit (SRU) that maintains robust spatio-temporal memory to mitigate perception blind spots. To bridge the sim-to-real gap in depth perception, we propose a high-fidelity sim-to-real depth sensor modeling pipeline that faithfully replicates real-world sensor artifacts. Additionally, we employ a 3D waypoint-guided active perception reward for proactive sensing, alongside hollow gap kinematic and stair edge penalties to ensure precise foothold placement. We successfully deployed StairMaster on a Unitree Go2 robot, demonstrating its ability to conquer hollow stairs with an unprecedented incline of up to 55$^\circ$ through zero-shot transfer. To the best of our knowledge, this is the first RL-based policy to achieve such steep hollow stair climbing in real-world environments. Project Website: https://sivan666666.github.io/StairMaster/.

cs.RO

Semi-invisible Hyperon Decays in the Effective Lagrangian Approach

We systematically investigate the semi-invisible decays of hyperons (hyperon $\to \pi(/\gamma)\ +$ invisible($\psi$)) in the Mesogenesis mechanism by the effective Lagrangian approach. The one-loop hadronic contribution of triangle diagrams with final-state interactions is fully examined in the present work. Our analysis indicates that the triangle diagram yield sizable corrections to the branching ratio that are as significant as those from tree diagrams. Especially for the $\Sigma^-\to \pi^-\psi$ and $\Xi^0 \to \pi^0 \psi$, their loop contributions cannot be ignored. Consequently, the branching ratios of hyperon hadronic semi-invisible decays are found to be of order $10^{-5}$, particularly for $\Sigma^+\to\pi^+\psi$, $\Xi^0\to\pi^0\psi$, and $\Xi^-\to\pi^-\psi$, whereas those of radiative semi-invisible decays are less than $10^{-7}$.

hep-ph

ACCORD: Action-Conditioned Contextual Grounding for Language Agents

User instructions are often underspecified because humans rely on implicit assumptions about the surrounding environment. For large language model (LLM) agents operating in information-rich digital and physical environments, these assumptions cannot be inferred from the instruction alone; they must be recovered from the current state of tools, data, interfaces, and observations. Effective execution therefore requires agents to identify missing context, ground it in observed evidence, and carry it forward into subsequent actions. We show that current agents often fail to do so. They act from assumed rather than observed specifics, overlook information they could have gathered, and fail to incorporate evidence that has already been returned. Building on this insight, we propose ACCORD (Action-Conditioned Contextual Grounding), a simple and effective agent framework for adaptive grounding. Before each action, ACCORD actively probes the environment for missing information and integrates relevant context from the agent's trajectory that would otherwise be overlooked. Requiring no additional training or task-success signals, ACCORD improves task-goal completion on AppWorld by up to +20.6 points with GPT-5-mini, from 42.0% to 62.6%, compared to strong baselines. These gains persist with a substantially stronger base model (+10.8 with Claude-4.5-sonnet), an open-weight model (+10.1 with Qwen3.5-27B-FP8), and on the embodied AlfWorld benchmark (+7.4 success rate with GPT-5-mini).

cs.CL

SimWeaver: Zero-Shot RGB Sim-to-Real for Deformable Manipulation

RGB sim-to-real for deformable manipulation has remained largely unsolved without real-world fine-tuning. We present SimWeaver, which trains zero-shot RGB VLA policies on 200 simulated demonstrations per task, reaching above 80% per-task and 91% average real-world success across 5 diverse deformable tasks including plastic-bag manipulation, without teleoperation or per-task calibration. SimWeaver combines a reliable measurement-backed simulator (SimWeaver-Sim) with an extensible asset framework supporting single-image generation(SimWeaver-Asset), a deterministic topology-aware trajectory synthesizer (SimWeaver-Syn), and a sim-to-real protocol with ISP-aware photometric augmentation (SimWeaver-Real). On silk grasping, the sim-trained policy reaches 100% under visual distribution shifts where real-data baselines drop to 9-70%, at two orders of magnitude lower per-trajectory cost. We will release SimWeaver and a representative asset subset. Project page: https://simweaver.github.io/

cs.RO

ISAC for AI: A Trade-off Framework Across Data Acquisition and Transfer in Federated Learning

In this paper, we propose a resource allocation framework for federated learning (FL) in integrated sensing and communication (ISAC) systems, where we consider not only the reliability of model transfer through communication, but also the quality of data acquisition through sensing in the first place. Unlike existing works that assume training data is pre-collected or only impose a fixed sensing signal-to-noise ratio (SNR) threshold to reflect data quality, we explicitly characterize the relationship between sensing data quality (measured by sensing SNR), dataset size, and the upload reliability in FL training, and exploit this relationship to allocate resources between sensing and communication under a shared energy budget. This is non-trivial due to the intricate coupling among sensing data quality, transmission reliability, and communication resource allocation; nevertheless, it enables a principled joint optimization framework that directly enhances learning performance. Specifically, we derive a closed-form convergence upper bound that quantifies the joint impact of these factors on the FL optimality gap. Utilizing this upper bound, the original intractable optimization problem can be reformulated into a tractable resource allocation problem that jointly optimizes the sensing transmit power, number of sensing snapshots, and communication transmit power at each device subject to individual energy budget constraints. To solve the reformulated problem, we propose a two-layer optimization algorithm with linear complexity, where the outer layer employs golden section search and the inner layer solves per-device subproblems with closed-form solutions.

eess.SP

Low Light Image Enhancement Challenge at NTIRE 2026

This paper presents a comprehensive review of the NTIRE 2026 Low Light Image Enhancement Challenge, highlighting the proposed solutions and final results. The objective of this challenge is to identify effective networks capable of producing clearer and visually compelling images in diverse and challenging conditions by learning representative visual cues with the purpose of restoring information loss due to low-contrast and noisy images. A total of 195 participants registered for the first track and 153 for the second track of the competition, and 22 teams ultimately submitted valid entries. This paper thoroughly evaluates the state-of-the-art advances in (joint denoising and) low-light image enhancement, showcasing the significant progress in the field, while leveraging samples of our novel dataset.

cs.CV

Awakening the Sleeping Agent: Lean-Specific Agentic Data Reactivates General Tool Use in Goedel Prover

Heavy supervised fine-tuning on a target domain can strongly suppress capabilities that were present in the base model. We study this phenomenon in formal mathematics using Goedel-Prover-V2, an open-source model heavily trained on 1.8 million formal-math examples. After domain specialization, the model almost completely loses its ability to produce valid tool calls, even when explicitly instructed to use tools, dropping from 89.4% function-calling accuracy in the base model to nearly 0%. We ask whether this agentic collapse is permanent or instead reversible. To answer this question, we fine-tune the specialized model on a small amount of Lean-specific tool-use data. Remarkably, as few as 100 agentic traces are sufficient to restore strong tool-calling behavior. Importantly, this recovery is not the result of reward hacking or benchmark-specific optimization: the recovery data is entirely drawn from the Lean setting, where the model uses natural-language queries to search the Mathlib library for relevant theorems and lemmas, yet the regained capability transfers well beyond that domain. In particular, these same 100 Lean-specific traces improve performance on the Berkeley Function Calling Leaderboard from near zero to 83.8%, approaching the base model's 89.4% despite the mismatch in task distribution and protocol. The recovered capability is also practically useful in-domain. On ProofNet, pass@32 improves from 21.51% to 25.81%. Together, these results show that heavy domain supervised fine-tuning can suppress general tool-use ability without permanently erasing it, and that a small amount of domain-specific agentic data can awaken dormant tool-use capabilities.

cs.AI

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On

As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing specialized VTON models. Meanwhile, universal multi-reference image editing models have progressed rapidly and exhibit strong generalization in visual editing, suggesting a promising route toward more flexible VTON systems. However, despite their strong capabilities, the strengths and limitations of universal editors for VTON remain insufficiently explored due to the lack of systematic evaluation benchmarks. To address this gap, we introduce VTEdit-Bench, a comprehensive benchmark designed to evaluate universal multi-reference image editing models across various realistic VTON scenarios. VTEdit-Bench contains 24,220 test image pairs spanning five representative VTON tasks with progressively increasing complexity, enabling systematic analysis of robustness and generalization. We further propose VTEdit-QA, a reference-aware VLM-based evaluator that assesses VTON performance from three key aspects: model consistency, cloth consistency, and overall image quality. Through this framework, we systematically evaluate eight universal editing models and compare them with seven specialized VTON models. Results show that top universal editors are competitive on conventional tasks and generalize more stably to harder scenarios, but remain challenged by complex reference configurations, particularly multi-cloth conditioning.

cs.CV

The Illusion of Forgetting: Attack Unlearned Diffusion via Initial Latent Variable Optimization

Text-to-image diffusion models (DMs) are frequently abused to produce harmful or copyrighted content, violating public interests. Concept erasure (unlearning) is a promising paradigm to alleviate this issue. However, there exists a peculiar forgetting illusion phenomenon with unclear cause. Based on empirical analysis, we formally explain this cause: most unlearning partially disrupt the mapping between linguistic symbols and the underlying internal knowledge, leaving the knowledge intact as dormant memories. We further demonstrate that distributional discrepancy in the denoising process serves as a measurable indicator of how much of the mapping is retained, also reflecting unlearning strength. Inspired by this, we propose IVO (Initial Latent Variable Optimization), a novel attack framework designed to assess the robustness of current unlearning methods. IVO optimizes initial latent variables to realign the noise distribution of unlearned models with that of their vanilla counterparts, which reconstructs the fractured mappings and consequently revives dormant memories. Extensive experiments covering 11 unlearning techniques and 3 concept scenarios show that IVO outperforms state-of-the-art baselines, exposing fundamental flaws in current unlearning mechanisms. Warning: This paper has unsafe images that may offend some readers.

cs.LG

Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks

In recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of downstream tasks. To address this, we propose a new task of Burst Image Quality Assessment (BuIQA), to evaluate the task-driven quality of each frame within a burst sequence, providing reasonable cues for burst image selection. Specifically, we establish the first benchmark dataset for BuIQA, consisting of $7,346$ burst sequences with $45,827$ images and $191,572$ annotated quality scores for multiple downstream scenarios. Inspired by the data analysis, a unified BuIQA framework is proposed to achieve an efficient adaption for BuIQA under diverse downstream scenarios. Specifically, a task-driven prompt generation network is developed with heterogeneous knowledge distillation, to learn the priors of the downstream task. Then, the task-aware quality assessment network is introduced to assess the burst image quality based on the task prompt. Extensive experiments across 10 downstream scenarios demonstrate the impressive BuIQA performance of the proposed approach, outperforming the state-of-the-art. Furthermore, it can achieve $0.33$ dB PSNR improvement in the downstream tasks of denoising and super-resolution, by applying our approach to select the high-quality burst frames.

cs.CV

Representations of rational numbers and Minkowski dimension

In this paper, we investigate the representations of rational numbers via continued fraction, Egyptian fraction, and Engel fraction expansions. Given $m \in \mathbb{N}$, denote by $C_m, E_m, E_m^*$ the sets of rational numbers whose continued fraction, Egyptian fraction, and Engel fraction expansions have length $m$, respectively. We first establish the Minkowski dimensions of these sets, which implies that their global scaling properties are different. We also apply the results to sumsets of decreasing sequences.

math.CA

Box dimension of the graphs of recurrent fractal interpolation functions

Let $f$ be a generalized affine recurrent fractal interpolation function with vertical scaling functions. In this paper, by introducing underlying local iterated function systems of $f$, we define restricted vertical scaling matrices. Then we prove the monotonicity of spectral radii of these matrices without additional conditions. We also prove the irreducibility of these matrices under the assumption that vertical scaling functions are positive. With these results, we estimate the upper and lower box dimensions of the graphs of $f$ by the limits of spectral radii of restricted vertical scaling matrices. In particular, we obtain an explicit formula of the box dimension of the graph of $f$ under certain constraint conditions.

math.CA

Local level sets of the Takagi-van der Waerden function

In this paper, we investigate the Takagi-van der Waerden function, $$ T_r(x) = \sum_{n=0}^{\infty} \frac{\phi(r^n x)}{r^n} ,\quad x\in [0,1], \quad r \in \mathbb{Z}^+, $$ where $\phi(x)={\rm dist}(x,\mathbb{Z})$ represents the distance from $x$ to the nearest integer. %We prove that for every even integer $r \geq 2$, the expected number of local level sets contained in the level set $L_r(y)$ is $1 + 1/r$, if $y$ is a random variable uniformly distributed over the range of $T_r$. Lagarias and Maddock [Level sets of the Takagi function: local level sets, \emph{Monatsh. Math.}, {\bf 166} (2012), No. 2, 201--238] introduced the notion of local level sets for the classical Takagi function $T_2$. They proved that if $y$ is a random variable uniformly distributed over the range of $T_2$, then the expected number of local level sets contained in the level set $L_2(y)$ equals $3/2$. We extend the study by defining an analogous concept of local level sets for all even integers $r$. Then we prove that, for every even integer $r\geq 2$, if $y$ is a random variable uniformly distributed, then the expected number of local level sets contained in the level set $L_r(y)$ equals $1 + 1/r$.

math.CA