SearcharxivSearch

arXiv subjects

Kexun Chen

Publications and source records attributed to Kexun Chen.

6 recordsLinked to original sources

Monocular Depth Estimation via Neural Network with Learnable Algebraic Group and Ring Structures

Monocular depth estimation (MDE) has witnessed remarkable progress driven by Convolutional Neural Networks and transformer-based architectures. However, these approaches typically treat the problem as a generic image-to-image regression on Euclidean grids, thereby overlooking the intrinsic algebraic and geometric structures induced by perspective projection. To address this limitation, we propose LAGRNet, a novel framework that fundamentally grounds MDE in algebraic geometry by explicitly embedding learnable group, ring, and sheaf structures into the deep learning pipeline. Modeling feature maps as sections of a sheaf over an approximated image manifold, our method first establishes a Group-defined Feature Manifold (GFM) parameterized by a learned algebraic group action to enforce projective equivariance and robustness against view changes. To facilitate algebraically consistent cross-scale interactions, we subsequently introduce a Ring Convolution Layer (RCL) that formulates feature fusion as a graded ring homomorphism. Furthermore, to ensure global topological consistency, a Sheaf-based Module (SM) aggregates local depth cues via \v{C}ech nerve on the image topology. Extensive zero-shot evaluations across the KITTI, NYU-Depth V2, and ETH3D benchmarks demonstrate that LAGRNet significantly outperforms state-of-the-art methods in both accuracy and generalization capabilities.

cs.CV

Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space

Imitation learning (IL) with human demonstrations is a promising method for robotic manipulation tasks. While minimal demonstrations enable robotic action execution, achieving high success rates and generalization requires high cost, e.g., continuously adding data or incrementally conducting human-in-loop processes with complex hardware/software systems. In this paper, we rethink the state/action space of the data collection pipeline as well as the underlying factors responsible for the prediction of non-robust actions. To this end, we introduce a Hierarchical Data Collection Space (HD-Space) for robotic imitation learning, a simple data collection scheme, endowing the model to train with proactive and high-quality data. Specifically, We segment the fine manipulation task into multiple key atomic tasks from a high-level perspective and design atomic state/action spaces for human demonstrations, aiming to generate robust IL data. We conduct empirical evaluations across two simulated and five real-world long-horizon manipulation tasks and demonstrate that IL policy training with HD-Space-based data can achieve significantly enhanced policy performance. HD-Space allows the use of a small amount of demonstration data to train a more powerful policy, particularly for long-horizon manipulation tasks. We aim for HD-Space to offer insights into optimizing data quality and guiding data scaling. project page: https://hd-space-robotics.github.io.

cs.RO

Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration

Understanding how humans cooperatively utilize semantic knowledge to explore unfamiliar environments and decide on navigation directions is critical for house service multi-robot systems. Previous methods primarily focused on single-robot centralized planning strategies, which severely limited exploration efficiency. Recent research has considered decentralized planning strategies for multiple robots, assigning separate planning models to each robot, but these approaches often overlook communication costs. In this work, we propose Multimodal Chain-of-Thought Co-Navigation (MCoCoNav), a modular approach that utilizes multimodal Chain-of-Thought to plan collaborative semantic navigation for multiple robots. MCoCoNav combines visual perception with Vision Language Models (VLMs) to evaluate exploration value through probabilistic scoring, thus reducing time costs and achieving stable outputs. Additionally, a global semantic map is used as a communication bridge, minimizing communication overhead while integrating observational results. Guided by scores that reflect exploration trends, robots utilize this map to assess whether to explore new frontier points or revisit history nodes. Experiments on HM3D_v0.2 and MP3D demonstrate the effectiveness of our approach. Our code is available at https://github.com/FrankZxShen/MCoCoNav.git.

cs.RO

Channel Pruned YOLOv5-based Deep Learning Approach for Rapid and Accurate Outdoor Obstacles Detection

One-stage algorithm have been widely used in target detection systems that need to be trained with massive data. Most of them perform well both in real-time and accuracy. However, due to their convolutional structure, they need more computing power and greater memory consumption. Hence, we applied pruning strategy to target detection networks to reduce the number of parameters and the size of model. To demonstrate the practicality of the pruning method, we select the YOLOv5 model for experiments and provide a data set of outdoor obstacles to show the effect of model. In this specific data set, in the best circumstances, the volume of the network model is reduced by 49.7% compared with the original model, and the reasoning time is reduced by 52.5%. Meanwhile, it also uses data processing methods to compensate for the drop in accuracy caused by pruning.

cs.CV

Nanostructured germanium with >99 % absorption at 300-1600 nm wavelengths

Near-infrared (NIR) sensors find numerous applications within various industry fields, including optical communications and medical diagnostics. However, the state-of-the-art NIR sensors made of germanium (Ge) suffer from rather poor response, largely due to high reflection from the illuminated device surface. We demonstrate here a method to increase the sensitivity of Ge sensors by implementing nanostructures to the wafer surfaces. The absorbance of nanostructured Ge wafers is measured to be >99 % in the whole UV-VIS-NIR spectrum up to 1600 nm wavelength, which is a significant improvement to bare Ge wafers that reach absorption of only 63 % in maximum. The process is shown to be capable of producing uniform nanostructures covering full 100-mm-diameter substrates as well as wafers with etch mask openings of different sizes and shapes, which demonstrates its applicability to CMOS sensor manufacturing. The results imply that nanostructured Ge has potential to revolutionize the sensitivity of Ge-based sensors.

cond-mat.mtrl-sci

Efficient photon capture on germanium surfaces using industrially feasible nanostructure formation

Nanostructured surfaces are known to provide excellent optical properties for various photonics devices. Fabrication of such nanoscale structures to germanium (Ge) surfaces by metal assisted chemical etching (MACE) is, however, challenging as Ge surface is highly reactive resulting often in micron-level rather than nanoscale structures. Here we show that by properly controlling the process, it is possible to confine the chemical reaction only to the vicinity of the metal nanoparticles and obtain nanostructures also in Ge. Furthermore, it is shown that controlling the density of the nanoparticles, concentration of oxidizing and dissolving agents as well as the etching time plays a crucial role in successful nanostructure formation. We also discuss the impact of high mobility of charge carriers on the chemical reactions taking place on Ge surfaces. As a result we propose a simple one-step MACE process that results in nanoscale structures with less than 10% surface reflectance in the wavelength region between 400 nm and 1600 nm. The method consumes only a small amount of Ge and is thus industrially viable and also applicable to thin Ge layers.

physics.app-ph