SearcharxivSearch

arXiv subjects

Yuan Ren

Publications and source records attributed to Yuan Ren.

At least 19 recordsLinked to original sources

Rydberg-atom microwave angle-of-arrival detection via cylindrical vapor-cell-mediated field redistribution

Microwave angle-of-arrival(AoA) measurement is essential for radar, communication, and spectrum mon-itoring. Existing Rydberg-atom-based AoA schemes employ phase-difference measurements with local oscil-lators, standing-wave fluorescence imaging, or amplitude-ratio readout with internal metal reflectors. Here we demonstrate a new approach: a cylindrical glass vapor cell serving directly as an angle-encoding dielectric structure, eliminating the need for multiple apertures, local oscillators, or imaging optics. The cylindrical geometry produces angle-dependent reflection and field redistribution, mapping the incident AoA onto the effective microwave field sampled by the Rydberg ensemble. This effective field is read out optically via the Autler-Townes(A-T) splitting in the electromagnetically induced transparency(EIT) spectrum. Full-wave simulations and experiments at 11.64 GHz confirm a deterministic, geometry-mediated response over 0{\deg}-90{\deg}.Over the monotonic operating range(55{\deg}-125{\deg}), the angular resolution(minimum distinguishable increment)is 0.05{\deg}, and the angular accuracy(RMSE of repeated measurements) is 0.13{\deg} across the full range, improving to 0.05{\deg} in the optimal region(65{\deg}-115{\deg}). Occupying a sensing volume of~2.5 cm3, this method offers a compact, single-sensor pathway toward Rydberg AoA receivers.

physics.optics

RAEM: Robust Autonomous Exploration for Multi-Floor Environments with a Quadruped Robot

In this paper, we propose RAEM, a robust autonomous exploration framework for quadruped robots operating in multi-floor environments. Most existing ground-robot exploration approaches rely on planar traversability representations, which cannot adequately represent the overlapping structures and cross-floor connectivity of multi-floor buildings. Although tomography-based representations provide effective traversability modeling for multi-floor navigation, maintaining a global tomography map incurs substantial computational overhead for online exploration with frequent replanning. Moreover, sparse and fragmented LiDAR observations in stairwells can degrade local traversability estimation, leading to irregular viewpoint placement and temporary topological disconnections. To address these challenges, RAEM adopts a hybrid local-global traversability representation, in which a local tomography map and an explicitly categorized local 3D grid map are used for online terrain analysis and connectivity evaluation, while an elevation-aware global topological graph is incrementally constructed from these local spatial representations for efficient cross-floor exploration planning. We further introduce a staircase center alignment strategy to reduce abrupt yaw variations during climbing and a dual path searching mechanism to recover guidance paths when the global topology is locally disconnected. Extensive simulation and real-world experiments demonstrate robust and computationally stable autonomous exploration across multi-floor structures, including continuous exploration of a five-floor stairwell.

cs.RO

Scalable native signed optical computing enabled by dual-wavelength incoherent multiplexing

Incoherent photonic neural networks (PNNs) provide a robust platform for analog optical computing, yet efficient implementation of native signed operations remains challenging. Existing incoherent PNNs approaches often require additional spatial channels or temporal encoding steps to represent bipolar input signals, resulting in hardware overhead that scales with system size. Here, we demonstrate a dual-wavelength incoherent photonic architecture that natively supports both signed inputs and signed weights on a thin-film lithium niobate platform. By encoding complementary signal components onto two wavelength channels and performing computation within a shared physical path, the proposed scheme eliminates duplicated weighting units. As a result, the additional hardware overhead associated with signed computation remains constant per multiply accumulate operation, independent of matrix size. The fabricated device exhibits a modulation bandwidth exceeding 40 GHz and achieves four-quadrant optical multiplication with a standard deviation error of 1.27%. System-level functionality is validated through neural-network classification, achieving 95.1% accuracy on the Moons dataset and 91.63% on MNIST. These results establish a practical route toward scalable incoherent photonic computing systems with native bipolar processing capability.

physics.optics

Part-Level 3D Gaussian Vehicle Generation with Joint and Hinge Axis Estimation

Simulation is essential for autonomous driving, yet current frameworks often model vehicles as rigid assets and fail to capture part-level articulation. With perception algorithms increasingly leveraging dynamics such as wheel steering or door opening, realistic simulation requires animatable vehicle representations. Existing CAD-based pipelines are limited by library coverage and fixed templates, preventing faithful reconstruction of in-the-wild instances. We propose a generative framework that, from a single image or sparse multi-view input, synthesizes an animatable 3D Gaussian vehicle. Our method addresses two challenges: (i) large 3D asset generators are optimized for static quality but not articulation, leading to distortions at part boundaries when animated; and (ii) segmentation alone cannot provide the kinematic parameters required for motion. To overcome this, we introduce a part-edge refinement module that enforces exclusive Gaussian ownership and a kinematic reasoning head that predicts joint positions and hinge axes of movable parts. Together, these components enable faithful part-aware simulation, bridging the gap between static generation and animatable vehicle models.

cs.AI

UniScale: Unified Scale-Aware 3D Reconstruction for Multi-View Understanding via Prior Injection for Robotic Perception

We present UniScale, a unified, scale-aware multi-view 3D reconstruction framework for robotic applications that flexibly integrates geometric priors through a modular, semantically informed design. In vision-based robotic navigation, the accurate extraction of environmental structure from raw image sequences is critical for downstream tasks. UniScale addresses this challenge with a single feed-forward network that jointly estimates camera intrinsics and extrinsics, scale-invariant depth and point maps, and the metric scale of a scene from multi-view images, while optionally incorporating auxiliary geometric priors when available. By combining global contextual reasoning with camera-aware feature representations, UniScale is able to recover the metric-scale of the scene. In robotic settings where camera intrinsics are known, they can be easily incorporated to improve performance, with additional gains obtained when camera poses are also available. This co-design enables robust, metric-aware 3D reconstruction within a single unified model. Importantly, UniScale does not require training from scratch, and leverages world priors exhibited in pre-existing models without geometric encoding strategies, making it particularly suitable for resource-constrained robotic teams. We evaluate UniScale on multiple benchmarks, demonstrating strong generalization and consistent performance across diverse environments. We will release our implementation upon acceptance.

cs.CV

First Submillimeter Lights from Dome A: Tracing the Carbon Cycle in the Feedback of Massive Stars

The cycling of carbon between its ionized, atomic, and molecular phases shapes the chemical compositions and physical conditions of the interstellar medium (ISM). However, ground-based studies of the full carbon cycle have been limited by atmospheric absorption. Dome~A, the most promising site for submillimeter astronomy, has long resisted successful submillimeter astronomical observations. Using the 60~cm Antarctic Terahertz Explorer, we present the first successful CO ($4-3$) and [CI] ($^3P_1 - ^3P_0$) mapping observations of two archetypal triggered massive star-formation regions at Dome~A. These data, together with archival [CII], provide the first complete characterization of all three carbon phases in these environments. We find elevated C$^{0}$/CO abundance ratios in high-extinction regions, plausibly driven by deep penetration of intense radiation fields from massive stars into a clumpy ISM. These findings mark a major milestone for submillimeter astronomy at Dome~A and offer valuable insights into the impact of massive star feedback on the surrounding ISM.

astro-ph.GA

Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark

4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what extent can Multimodal Large Language Models (MLLMs) achieve human-level 4D spatial intelligence? In this work, we present Spatial4D-Bench, a versatile 4D spatial intelligence benchmark designed to comprehensively assess the 4D spatial reasoning abilities of MLLMs. Unlike existing spatial intelligence benchmarks that are often small-scale or limited in diversity, Spatial4D-Bench provides a large-scale, multi-task evaluation benchmark consisting of ~40,000 question-answer pairs covering 18 well-defined tasks. We systematically organize these tasks into six cognitive categories: object understanding, scene understanding, spatial relationship understanding, spatiotemporal relationship understanding, spatial reasoning and spatiotemporal reasoning. Spatial4D-Bench thereby offers a structured and comprehensive benchmark for evaluating the spatial cognition abilities of MLLMs, covering a broad spectrum of tasks that parallel the versatility of human spatial intelligence. We benchmark various state-of-the-art open-source and proprietary MLLMs on Spatial4D-Bench and reveal their substantial limitations in a wide variety of 4D spatial reasoning aspects, such as route plan, action recognition, and physical plausibility reasoning. We hope that the findings provided in this work offer valuable insights to the community and that our benchmark can facilitate the development of more capable MLLMs toward human-level 4D spatial intelligence. More resources can be found on our project page.

cs.CV

High-Q Lithium Niobate Microring Resonator with Electro-Optically Reconfigurable Coupling Strength

The development of sophisticated integrated photonic circuits demands microresonators that combine exceptional optical confinement with dynamic operational flexibility. Here, we demonstrate a racetrack resonator on the thin-film lithium niobate platform that achieves an electro-optically tunable coupling strength while maintaining a stable, high intrinsic Q factor on the order of 10^6. By incorporating a Mach-Zehnder interferometer into the coupling region, the device facilitates a continuous and reversible transition across the entire coupling spectrum from under-coupling and critical coupling to deep over-coupling. To ensure high spectral purity, we employ Euler bends to facilitate an adiabatic transition between the straight and curved waveguide sections. This design effectively suppresses the excitation of higher-order modes, resulting in a clean transmission spectrum characterized by exclusive fundamental mode operation. At the critical coupling point, the resonator exhibits a high extinction ratio exceeding 30 dB. The integration of stable ultra-high Q, single-mode purity, and full-range coupling reconfigurability positions this device as a vital component for adaptive microwave photonics, high-efficiency nonlinear optics, and programmable quantum photonic networks.

physics.optics

Transformational astrophysics and exoplanet science with Habitable Worlds Observatory's High Resolution Imager

Habitable Worlds Observatory (HWO) will be NASA's flagship space telescope of the 2040s, designed to search for life on other planets and to transform broad areas of astrophysics. NASA are seeking international partners, and the UK is well-placed to lead the design and construction of its imaging camera - which is likely to produce the mission's most visible public impact. Early participation in the mission would return investment to UK industry, and bring generational leadership for the UK in space science, space technology, and astrophysics.

astro-ph.IM

Near-Zero Crosstalk and Ultra-Low Loss Waveguide Crossings Enabled by three-dimensional Ta2O5-on-LNOI Integrated Photonic Platform

Waveguide crossings represent one of the most critical components in very-large-scale photonic integration (VLSPI). Three-dimensional waveguide crossings, which distribute optical pathways across multiple planes, can achieve near-zero crosstalk and extremely low crossing-induced loss. However, they face an intrinsic trade-off between interlayer crossing performance and coupling efficiency. To address this challenge, we developed a low-cost fabrication method for 3D waveguide crossings by exploiting the edge rounding effect inherent to chemical mechanical polishing (CMP). Using this method, we demonstrate waveguide crossings with average loss below 0.002 dB and crosstalk below -62 dB on Ta2O5-on-LNOI integrated photonic platform. Our method maintains full compatibility with conventional semiconductor manufacturing technology and paves the way for realizing VLSPI on the thin-film lithium niobate platform.

physics.optics

Monocular Visual 8D Pose Estimation for Articulated Bicycles and Cyclists

In Autonomous Driving, cyclists belong to the safety-critical class of Vulnerable Road Users (VRU), and accurate estimation of their pose is critical for cyclist crossing intention classification, behavior prediction, and collision avoidance. Unlike rigid objects, articulated bicycles are composed of movable rigid parts linked by joints and constrained by a kinematic structure. 6D pose methods can estimate the 3D rotation and translation of rigid bicycles, but 6D becomes insufficient when the steering/pedals angles of the bicycle vary. That is because: 1) varying the articulated pose of the bicycle causes its 3D bounding box to vary as well, and 2) the 3D box orientation is not necessarily aligned to the orientation of the steering which determines the actual intended travel direction. In this work, we introduce a method for category-level 8D pose estimation for articulated bicycles and cyclists from a single RGB image. Besides being able to estimate the 3D translation and rotation of a bicycle from a single image, our method also estimates the rotations of its steering handles and pedals with respect to the bicycle body frame. These two new parameters enable the estimation of a more fine-grained bicycle pose state and travel direction. Our proposed model jointly estimates the 8D pose and the 3D Keypoints of articulated bicycles, and trains with a mix of synthetic and real image data to generalize on real images. We include an evaluation section where we evaluate the accuracy of our estimated 8D pose parameters, and our method shows promising results by achieving competitive scores when compared against state-of-the-art category-level 6D pose estimators that use rigid canonical object templates for matching.

cs.CV

Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators

Compute-in-memory (CIM) accelerators have emerged as a promising way for enhancing the energy efficiency of convolutional neural networks (CNNs). Deploying CNNs on CIM platforms generally requires quantization of network weights and activations to meet hardware constraints. However, existing approaches either prioritize hardware efficiency with binary weight and activation quantization at the cost of accuracy, or utilize multi-bit weights and activations for greater accuracy but limited efficiency. In this paper, we introduce a novel binary weight multi-bit activation (BWMA) method for CNNs on CIM-based accelerators. Our contributions include: deriving closed-form solutions for weight quantization in each layer, significantly improving the representational capabilities of binarized weights; and developing a differentiable function for activation quantization, approximating the ideal multi-bit function while bypassing the extensive search for optimal settings. Through comprehensive experiments on CIFAR-10 and ImageNet datasets, we show that BWMA achieves notable accuracy improvements over existing methods, registering gains of 1.44\%-5.46\% and 0.35\%-5.37\% on respective datasets. Moreover, hardware simulation results indicate that 4-bit activation quantization strikes the optimal balance between hardware cost and model performance.

cs.AR

UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation

Estimating the 6D pose of novel objects is a fundamental yet challenging problem in robotics, often relying on access to object CAD models. However, acquiring such models can be costly and impractical. Recent approaches aim to bypass this requirement by leveraging strong priors from foundation models to reconstruct objects from single or multi-view images, but typically require additional training or produce hallucinated geometry. To this end, we propose UnPose, a novel framework for zero-shot, model-free 6D object pose estimation and reconstruction that exploits 3D priors and uncertainty estimates from a pre-trained diffusion model. Specifically, starting from a single-view RGB-D frame, UnPose uses a multi-view diffusion model to estimate an initial 3D model using 3D Gaussian Splatting (3DGS) representation, along with pixel-wise epistemic uncertainty estimates. As additional observations become available, we incrementally refine the 3DGS model by fusing new views guided by the diffusion model's uncertainty, thereby continuously improving the pose estimation accuracy and 3D reconstruction quality. To ensure global consistency, the diffusion prior-generated views and subsequent observations are further integrated in a pose graph and jointly optimized into a coherent 3DGS field. Extensive experiments demonstrate that UnPose significantly outperforms existing approaches in both 6D pose estimation accuracy and 3D reconstruction quality. We further showcase its practical applicability in real-world robotic manipulation tasks.

cs.RO

A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips

Designing lightweight convolutional neural network (CNN) models is an active research area in edge AI. Compute-in-memory (CIM) provides a new computing paradigm to alleviate time and energy consumption caused by data transfer in von Neumann architecture. Among competing alternatives, resistive random-access memory (RRAM) is a promising CIM device owing to its reliability and multi-bit programmability. However, classical lightweight designs such as depthwise convolution incurs under-utilization of RRAM crossbars restricted by their inherently dense weight-to-RRAM cell mapping. To build an RRAM-friendly yet efficient CNN, we evaluate the hardware cost of DenseNet which maintains a high accuracy vs other CNNs at a small parameter count. Observing the linearly increasing channels in DenseNet leads to a low crossbar utilization and causes large latency and energy consumption, we propose a scheme that concatenates feature maps of front layers to form the input of the last layer in each stage. Experiments show that our proposed model consumes less time and energy than conventional ResNet and DenseNet, while producing competitive accuracy on CIFAR and ImageNet datasets.

cs.AR

On the Hulls of Group Codes

Let $\mathbb {F}_q$ be a finite field and $G$ a finte group with $(|G|,q)=1$. By a group code in $\mathbb {F}_q[G]$ we mean a two-sided ideal in $\mathbb {F}_q[G]$. We will prove a general criterion for the existence of group codes with given hull dimension, and then apply it to deduce explicit criterions for existence of group codes with hull dimension $\leq3$. In particular our criterion for the existence of $1$-dimensional hulls generalizes that of privious work which consider only abelian groups $G$.

cs.IT

Monolithically Integrated Optical Convolutional Processors on Thin Film Lithium Niobate

Photonic neural networks (PNNs) of sufficiently large physical dimensions and high operation accuracies are envisaged as an ideal candidate for breaking the major bottlenecks in the current artificial intelligence architectures in terms of latency, energy efficiency and computational power. To achieve this vision, it is of vital importance to scale up the PNNs and in the meantime reduce the high demand on the dimensions required by the PNNs. The underlying cause of this strategy is the enormous gap between the scales of photonic and electronic integrated circuits. Here, we demonstrate monolithically integrated optical convolutional processors on thin film lithium niobate (TFLN) to enable large-scale programmable convolution kernels and in turn greatly reduce the dimensions required by the subsequent fully connected layers. Experimental validation achieves high classification accuracies of 96%/86% on the MNIST/Fashion-MNIST datasets and 84.6% on the AG News dataset, while dramatically reducing the required subsequent fully connected layer dimensions to 196x10 (from 784x10) and 175x4 (from 800x4), respectively. Furthermore, our devices can be driven by commercial field-programmable gate array (FPGA) systems, a unique advantage in addition to their scalable channel number and kernel size, our architecture provides a solution to build practical machine learning photonic devices.

physics.optics

Response to recent comments on Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supp. Info. for Nature 638, 651-655 (2025)

The topological gap protocol (TGP) is a statistical test designed to identify a topological phase with high confidence and without human bias. It is used to determine a promising parameter regime for operating topological qubits. The protocol's key metric is the probability of incorrectly identifying a trivial region as topological, referred to as the false discovery rate (FDR). Two recent manuscripts [arXiv:2502.19560, arXiv:2503.08944] engage with the topological gap protocol and its use in Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supplementary Information for Nature 638, 651-655 (2025), although they do not explicitly dispute the main results of either one. We demonstrate that the objections in arXiv:2502.19560 and arXiv:2503.08944 are unfounded, and we uphold the conclusions of Phys. Rev. B 107, 245423 (2023) and Nature 638, 651-655 (2025). Specifically, we show that no flaws have been identified in our estimate of the false discovery rate (FDR). We provide a point-by-point rebuttal of the comments in arXiv:2502.19560 and arXiv:2503.08944.

cond-mat.mes-hall

SR-LIO++: LiDAR-Inertial Odometry and Quantized Mapping with Caching-Aware Sweep Reconstruction

Addressing the inherent low acquisition frequency limitation of 3D LiDAR to achieve high-frequency output has become a critical research focus in the LiDAR-Inertial Odometry (LIO) domain. To ensure real-time performance, frequency-enhanced LIO systems must process each sweep within significantly reduced timeframe, which presents substantial challenges for deployment on resource-constrained platforms. To address these limitations, we introduce SR-LIO++, an innovative LIO system capable of achieving doubled output frequency relative to input frequency on resource-constrained hardware platforms, including the Raspberry Pi 4B. Our system employs the previously proposed sweep reconstruction methodology to enhance LiDAR sweep frequency, generating high-frequency reconstructed sweeps. Building upon this foundation, we propose a caching mechanism for intermediate results (i.e., surface parameters) of the most recent segments, effectively minimizing redundant processing of common segments in adjacent reconstructed sweeps. This method decouples processing time from the traditionally linear dependence on reconstructed sweep frequency. Furthermore, we present a quantized map point management based on index table mapping, significantly reducing memory usage by converting global 3D point storage from 64-bit double precision to 8-bit char representation. This method also converts the computationally intensive Euclidean distance calculations in nearest neighbor searches from 64-bit double precision to 16-bit short and 32-bit integer formats, reducing computational cost. Extensive experimental evaluations across three distinct computing platforms and four public datasets demonstrate that SR-LIO++ maintains state-of-the-art accuracy while substantially enhancing efficiency. Notably, our system successfully achieves 20 Hz state output on Raspberry Pi 4B hardware.

cs.RO