SearcharxivSearch

arXiv subjects

Nan Luo

Publications and source records attributed to Nan Luo.

9 recordsLinked to original sources

Y-BotFrame: An Extensible Embodied Agent Framework for Quadruped Robot Assistants

Quadruped robots are capable of traversing a wide range of complex terrains with high flexibility. As highly mobile ground-based intelligent platforms, they can be equipped with modules for navigation control, environmental perception, and intelligent interaction, thereby serving as real-world mobile deployment platforms for various algorithms. In this paper, we introduce Y-BotFrame, an extensible embodied platform that turns a robot into an intelligent ground assistant. Y-BotFrame integrates multimodal perception capabilities, including speech, vision, and LiDAR, and employs a large language model as the cognitive core for environmental understanding, contextual reasoning, and task planning. The system maps user natural-language instructions into executable embodied task units that can be carried out by the robot. Y-BotFrame supports natural interaction through voice commands and visual feedback, removing the need for a remote controller and enabling efficient human-robot collaboration. With a highly extensible framework, Y-BotFrame supports plug-and-play integration of new functional modules as well as modular upgrades and iterative development, offering a reference implementation for the real-world deployment of general-purpose, instruction-driven embodied agents.The supplementary video is available at https://xdei-group.github.io/Y-BotFrame/.

cs.RO

AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents

Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring, and emergency response. However, most UAV applications still rely on pre-defined command sequences or task-specific pipelines, where developers manually connect perception, planning, flight control, simulation, logging, and safety modules. This limits the flexibility, reproducibility, and extensibility of autonomous aerial systems. This paper presents AerialClaw, an open-source software framework that enables UAVs to operate as decision-making aerial agents rather than merely command-following platforms. Given a natural-language mission, AerialClaw allows an LLM-based agent to understand the task, maintain context, invoke executable aerial skills, observe perception and runtime feedback, and iteratively update its decisions in a closed loop. The framework adopts a modular brain-skill-runtime architecture, combining hard skills for atomic UAV operations, Markdown-based soft skills for reusable task strategies, document-driven agent state and capability boundaries, memory-driven reflection, safety-oriented runtime validation, and platform-agnostic execution adapters. AerialClaw supports lightweight mock execution, PX4 SITL with Gazebo, and AirSim-based simulation, together with a web console, pluggable model backends, example missions, simulation assets, and staged deployment scripts. By combining standardized aerial skills, document-driven agent state, memory, and closed-loop LLM decision-making, AerialClaw provides a reproducible and extensible open-source framework for building UAV systems that can interpret missions, make decisions, execute skills, and adapt their behavior from feedback.

cs.RO

Improving Few-Shot Change Detection Visual Question Answering via Decision-Ambiguity-guided Reinforcement Fine-Tuning

Change detection visual question answering (CDVQA) requires answering text queries by reasoning about semantic changes in bi-temporal remote sensing images. A straightforward approach is to boost CDVQA performance with generic vision-language models via supervised fine-tuning (SFT). Despite recent progress, we observe that a significant portion of failures do not stem from clearly incorrect predictions, but from decision ambiguity, where the model assigns similar confidence to the correct answer and strong distractors. To formalize this challenge, we define Decision-Ambiguous Samples (DAS) as instances with a small probability margin between the ground-truth answer and the most competitive alternative. We argue that explicitly optimizing DAS is crucial for improving the discriminability and robustness of CDVQA models. To this end, we propose DARFT, a Decision-Ambiguity-guided Reinforcement Fine-Tuning framework that first mines DAS using an SFT-trained reference policy and then applies group-relative policy optimization on the mined subset. By leveraging multi-sample decoding and intra-group relative advantages, DARFT suppresses strong distractors and sharpens decision boundaries without additional supervision. Extensive experiments demonstrate consistent gains over SFT baselines, particularly under few-shot settings.

cs.CV

Is-NeRF: In-scattering Neural Radiance Field for Blurred Images

Neural Radiance Fields (NeRF) has gained significant attention for its prominent implicit 3D representation and realistic novel view synthesis capabilities. Available works unexceptionally employ straight-line volume rendering, which struggles to handle sophisticated lightpath scenarios and introduces geometric ambiguities during training, particularly evident when processing motion-blurred images. To address these challenges, this work proposes a novel deblur neural radiance field, Is-NeRF, featuring explicit lightpath modeling in real-world environments. By unifying six common light propagation phenomena through an in-scattering representation, we establish a new scattering-aware volume rendering pipeline adaptable to complex lightpaths. Additionally, we introduce an adaptive learning strategy that enables autonomous determining of scattering directions and sampling intervals to capture finer object details. The proposed network jointly optimizes NeRF parameters, scattering parameters, and camera motions to recover fine-grained scene representations from blurry images. Comprehensive evaluations demonstrate that it effectively handles complex real-world scenarios, outperforming state-of-the-art approaches in generating high-fidelity images with accurate geometric details.

cs.GR

Anomalous Transport of Elongated Particles in Oscillatory Vortical Flows

We investigate the transport dynamics of elongated particles in cellular vortical flows that undergo spatial oscillations over time. Experimental flow visualizations reveal mixed flow fields with chaotic and elliptic regions coexisting. Surprisingly, the particle transport rate does not increase monotonically with particle length, even though longer particles are expected to explore neighboring vortices more easily. Numerical simulations in a much larger system produce similar transport anomalies, characterized by subdiffusion due to frequent long-time trapping in vortices at certain lengths, but normal diffusion at others. At moderate oscillation frequencies, these long-time trapping events occur within the chaotic region; at high frequencies, they occur in the elliptic regions, but only for particles whose lengths match these regions. In the latter case, subdiffusion is robust against random noise. Our results reveal new mechanisms for controlling particle diffusion in fluid flows.

physics.flu-dyn

Enhancing Clean Label Backdoor Attack with Two-phase Specific Triggers

Backdoor attacks threaten Deep Neural Networks (DNNs). Towards stealthiness, researchers propose clean-label backdoor attacks, which require the adversaries not to alter the labels of the poisoned training datasets. Clean-label settings make the attack more stealthy due to the correct image-label pairs, but some problems still exist: first, traditional methods for poisoning training data are ineffective; second, traditional triggers are not stealthy which are still perceptible. To solve these problems, we propose a two-phase and image-specific triggers generation method to enhance clean-label backdoor attacks. Our methods are (1) powerful: our triggers can both promote the two phases (i.e., the backdoor implantation and activation phase) in backdoor attacks simultaneously; (2) stealthy: our triggers are generated from each image. They are image-specific instead of fixed triggers. Extensive experiments demonstrate that our approach can achieve a fantastic attack success rate~(98.98%) with low poisoning rate~(5%), high stealthiness under many evaluation metrics and is resistant to backdoor defense methods.

cs.CR

Topological cavity based on slow light topological edge mode for broadband Purcell enhancement

Slow light in topological valley photonic crystal structures offers new possibilities to enhance light-matter interaction. We report a topological cavity based on slow light topological edge mode for broadband Purcell enhancement. The topological edge modes with large group indices over 100 can be realized with a bearded interface between two topologically distinct valley photonic crystals, featuring the greatly enhanced Purcell factor because of the increased local density of states. In the slow light regime, the topological cavity supports much more cavity modes with higher quality factor than that in the fast light regime, which is both demonstrated theoretically and experimentally. We demonstrate the cavity enables the broadband Purcell enhancement together with substantial Purcell factor, benefiting from dense cavity modes with high quality factor in a wide spectral range. It has great benefit to the realization of high-efficiency quantum-dot-based single-photon sources and entangled-photon sources with less restriction on spectral match. Such topological cavity could serve as a significant building block toward the development of photonic integrated circuits with embedded quantum emitters.

physics.optics

Ab initio Study of Ground-State CS Photodissociation Via Highly Excited Electronic States

Photodissociation by ultraviolet radiation is the key destruction pathway for CS in photon-dominated regions, such as diffuse clouds. However, the large uncertainties of photodissociation cross sections and rates of CS, resulting from a lack of both laboratory experiments and theoretical calculations, limit the accuracy of calculated abundances of S-bearing molecules by modern astrochemical models. Here we show a detailed \textit{ab initio} study of CS photodissociation. Accurate potential energy curves of CS electronic states were obtained by choosing an active space CAS(8,10) in MRCI+Q/aug-cc-pV(5+d)Z calculation with additional diffuse functions, with a focus on the \(B\) and \(C\,^1Σ^+\) states. Cross sections for both direct photodissociation and predissociation from the vibronic ground state were calculated by applying the coupled-channel method. We found that the \(C-X\) \((0-0)\) transition has extremely strong absorption due to a large transition dipole moment in the Franck-Condon region and the upper state is resonant with several triplet states via spin-orbit couplings, resulting in predissociation to the main atomic products C \((^3P)\) and S \((^1D)\). Our new calculations show the photodissociation rate under the standard interstellar radiation field is \(2.9\ee{-9}\)\,s\(^{-1}\), with a 57\% contribution from \(C-X\) \((0-0)\) transition. This value is larger than that adopted by the Leiden photodissociation and photoionization database by a factor of 3.0. Our accurate \textit{ab initio} calculations will allow more secure determination of S-bearing molecules in astrochemical models.

astro-ph.GA

Enhanced CNN for image denoising

Owing to flexible architectures of deep convolutional neural networks (CNNs), CNNs are successfully used for image denoising. However, they suffer from the following drawbacks: (i) deep network architecture is very difficult to train. (ii) Deeper networks face the challenge of performance saturation. In this study, the authors propose a novel method called enhanced convolutional neural denoising network (ECNDNet). Specifically, they use residual learning and batch normalisation techniques to address the problem of training difficulties and accelerate the convergence of the network. In addition, dilated convolutions are used in the proposed network to enlarge the context information and reduce the computational cost. Extensive experiments demonstrate that the ECNDNet outperforms the state-of-the-art methods for image denoising.

cs.CV