Searcharxiv⌕ Search

arXiv subjects

Zhongyuan Liu

Publications and source records attributed to Zhongyuan Liu.

At least 19 recordsLinked to original sources

PART: Learning 3D Part Assembly and Retrieval with Transformers

3D assembly is fundamental to modern manufacturing and digital content creation. In this paper, we present PART, a unified transformer-based framework for 3D part retrieval and assembly: given a target shape and a part library, PART automatically selects the appropriate parts and predicts their 6-DoF poses to reconstruct the target. While prior work has achieved impressive progress on assembling a pre-defined set of parts, this more practical retrieval-based setting remains largely unexplored. The task faces three key challenges: (i) a combinatorially explosive search space that grows exponentially with library size; (ii) variable-length outputs, as different targets require different numbers of parts; and (iii) continuous 6-DoF pose estimation for part assembly. To address these, we formulate retrieval and assembly as a set prediction problem and design a novel transformer-based framework that retrieves parts and regresses their poses with variable-length output. Additionally, we exploit the duality between part pose estimation and target segmentation through joint training and a novel segmentation-enhanced optimization module. Finally, We curate a large-scale dataset of 80K+ shapes, and the results show that PART generalizes to scene layouts, image targets, and real-world scans. Project Page: https://iambrc.github.io/PART-project-page/.

cs.CV↗

Simultaneous Determination of Local Magnetic Fields and Sensor Orientation with Nitrogen-Vacancy Centers in Nanodiamond

Nitrogen-vacancy (NV) centers in nanodiamonds have emerged as a promising quantum sensing platform for biomedical imaging applications, yet random orientations of individual particles present significant challenges in large-scale sensor calibration. In this study, we demonstrate a novel approach to simultaneously determine each particle's crystallographic axes and the surrounding local vector magnetic field. Specifically, a minimum of four distinct bias fields is required to unambiguously extract both the orientation and the local field. We validate our method experimentally using NV centers in two scenarios: (1) in a bulk diamond with known crystal orientation as a proof of concept, and (2) on various single nanodiamonds to mimic real-world applications. Our work represents a crucial step towards unlocking the full potential of nanodiamonds for advanced applications such as in-situ biomedical imaging and nanoscale sensing in complex environments.

quant-ph↗

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development trajectories identifies three stages that together span the lifecycle of game development with a coding agent: initial game generation, bug diagnosis and repair, and optimization over multiple turns. Therefore, we introduce GameXpert-Bench, which operationalizes the three lifecycle stages as three complementary benchmark tracks. GameGen evaluates complete game creation from a single request in an empty workspace. GameFix evaluates diagnosis and repair when defects are reported or left for the agent to discover. GameOpt evaluates cumulative optimization through request chains seeded by real development trajectories between users and agents. We evaluate each track using live game interaction, deterministic behavioral tests, or final product criteria with regression checks. The suite contains 97 generation tasks across 11 genres; 100 repair tasks from 50 game levels verified by humans, each with 19-27 injected bugs; and 17 optimization chains with six turns and 102 requests. Across the three tracks, current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.

cs.AI↗

$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce $τ_0$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

cs.RO↗

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation

Recently, Large Language Models (LLMs) have emerged as promising layout agents for 3D scene generation. Existing layout agents still suffer from implausible layout generation because most of them convert 3D assets and 3D layouts into textual descriptions as inputs and outputs, which involves severe information loss due to the modality gap between texts and 3D assets and 3D layouts. We propose NaLA, a native 3D LLM layout Agent for high-quality 3D scene generation by placing 3D assets in the scene. For the inputs, NaLA encodes 3D scene boundaries and 3D assets directly into the LLM, preserving fine-grained geometry and enabling explicit reasoning over relationships like collisions, surface supporting, and containment. To accurately output the positions and orientations of assets, NaLA adopts a coarse-to-fine prediction mechanism that first predicts discrete poses in an autoregressive manner and then refines the discrete poses with a continuous regression. Trained on diverse layout datasets, NaLA attains strong geometric perception and layout coherence. Experiments demonstrate that NaLA outperforms prior layout agents in both generation quality and inference efficiency, with comprehensive ablation studies to verify each component's effectiveness.

cs.CV↗

Qualitative analysis of multi-peak solutions for Nonlinear Schrödinger equations with nearly critical Sobolev exponents

In this paper, we are concerned with qualitative properties of multi-peak solutions of the following nonlinear Schrödinger equations \begin{equation*} -Δu+V(x)u= u^{p-\varepsilon},\,\,\,u>0,\,\,\,\text{in}\,\,\,\mathbb{R}^N, \end{equation*} where $V(x)$ is a nonnegative continuous function, $\varepsilon>0$, $p=\frac{N+2}{N-2}$, $N\geq6$. The existence of multi-peak solutions has been obtained by Cao et al. (Calc. Var. Partial Differential Equations, 64: 139, 2025). The main objective in this paper is to establish the local uniqueness and Morse index of the multi-peak solutions in \cite{CLl1} provided that $V(x)$ possesses $k$ non-degenerate critical points by using the blow-up analysis based on Pohozaev identities.

math.AP↗

Existence, non-degeneracy and local uniqueness of multi-peak solutions to the fractional Schrödinger equation with nearly critical exponent in $\mathbb{R}^N$

In this paper, we consider the following fractional Schrödinger equation \begin{equation*} \left\{ \begin{array}{lcl} (-Δ)^{s}u+V(x)u=u^{{p_s}-ε}\ \ \ &\hbox{in}\ \mathbb{R}^N,\\ u>0\ \ \ &\hbox{in}\ \mathbb{R}^N, \end{array} \right. \end{equation*} where $0 0$, $p_s=(N+2s)/(N-2s)$, $N>4s$ and $V(x)\in C^1(\mathbb{R}^N)\cap L^\infty (\mathbb{R}^N)$ is non-negative. We first use the Lyapunov-Schmidt reduction method to construct multi-peak solutions to the above equation provided that $V(x)$ possesses $k$ stable critical points. Then we prove the non-degeneracy and local uniqueness of the multi-peak solutions, for $\frac{1}{2}<s<1$, $N\geq 6s$, via the blow-up argument based on various local Pohozaev identities. Due to the nonlocal property of the fractional Laplacian, we need to make delicate analysis of the approximate solutions and establish the local Pohozaev identities for the corresponding harmonic extension instead of $u$. This approach not only requires to develop refined estimates for several integrals in the local Pohozaev identities, but also to apply Pohozaev identities through a markedly different way.

math.AP↗

Spin Relaxometry with Solid-State Defects: Theory, Platforms, and Applications

Spin relaxometry using solid-state spin defects, such as the diamond nitrogen-vacancy (NV) center, probes dynamical processes by measuring how environmental fluctuations enhance the spin relaxation rate. In the weak-coupling limit, relaxation rates sample the transverse magnetic-noise power spectral density through a sensor-specific filter function, turning the defect into a local, frequency-selective noise spectrometer. This review bridges theory and experiment, clarifying how measured relaxation rates map onto noise spectra and how near-field geometry shapes the response. We highlight representative applications across condensed-matter physics, chemical and biological sensing, and relaxometry-based magnetic-resonance spectroscopy. We conclude with emerging opportunities and key challenges.

cond-mat.mes-hall↗

FU-MPC: Frontier- and Uncertainty-Aware Model Predictive Control for Efficient and Accurate UAV Exploration with Motorized LiDAR

Efficient UAV exploration in unknown environments requires rapid coverage expansion while maintaining accurate and reliable localization, since safe navigation in complex scenes depends on consistent mapping and pose estimation. However, for conventional LiDAR-equipped UAVs, the observable region is tightly coupled with the UAV pose and motion. Expanding coverage often requires additional translational or rotational maneuvers, which can reduce exploration efficiency and increase the risk of localization degradation in geometrically challenging environments. Motorized rotating LiDARs provide a promising solution by actively adjusting the sensor viewing direction without changing the UAV motion, thereby introducing an additional sensing degree of freedom. Nevertheless, existing exploration systems rarely exploit this scanning freedom as an explicit decision variable linked to both exploration progress and localization quality. To address this gap, we develop a UAV platform equipped with an independently actuated rotating LiDAR and propose a hierarchical exploration framework. The global planner organizes frontiers into representative viewpoints and sequences them using topology-aware transition costs. Built upon this planner, FU-MPC serves as a local receding-horizon scan controller that optimizes LiDAR rotation along the predicted flight trajectory. The controller jointly considers frontier-aware exploration utility and direction-dependent localization uncertainty, while lightweight surrogate evaluation enables real-time onboard execution. Experiments in complex environments demonstrate that the proposed system improves exploration efficiency while maintaining robust localization performance compared with fixed-pattern scanning and uncertainty-only baselines. The project page can be found at https://kafeiyin00.github.io/FU-MPC/.

cs.RO↗

S3KF: Spherical State-Space Kalman Filtering for Panoramic 3D Multi-Object Tracking

Panoramic multi-object tracking is important for industrial safety monitoring, wide-area robotic perception, and infrastructure-light deployment in large workspaces. In these settings, the sensing system must provide full-surround coverage, metric geometric cues, and stable target association under wide field-of-view distortion and occlusion. Existing image-plane trackers are tightly coupled to the camera projection and become unreliable in panoramic imagery, while conventional Euclidean 3D formulations introduce redundant directional parameters and do not naturally unify angular, scale, and depth estimation. In this paper, we present $\mathbf{S^3KF}$, a panoramic 3D multi-object tracking framework built on a motorized rotating LiDAR and a quad-fisheye camera rig. The key idea is a geometry-consistent state representation on the unit sphere $\mathbb{S}^2$, where object bearing is modeled by a two-degree-of-freedom tangent-plane parameterization and jointly estimated with box scale and depth dynamics. Based on this state, we derive an extended spherical Kalman filtering pipeline that fuses panoramic camera detections with LiDAR depth observations for multimodal tracking. We further establish a map-based ground-truth generation pipeline using wearable localization devices registered to a shared global LiDAR map, enabling quantitative evaluation without motion-capture infrastructure. Experiments on self-collected real-world sequences show decimeter-level planar tracking accuracy, improved identity continuity over a 2D panoramic baseline in dynamic scenes, and real-time onboard operation on a Jetson AGX Orin platform. These results indicate that the proposed framework is a practical solution for panoramic perception and industrial-scale multi-object tracking.The project page can be found at https://kafeiyin00.github.io/S3KF/.

cs.RO↗

A Deep-Learning-Boosted Framework for Quantum Sensing with Nitrogen-Vacancy Centers in Diamond

Nitrogen-vacancy (NV) centers in diamond are a versatile quantum sensing platform for high sensitivity measurements of magnetic fields, temperature and strain with nanoscale spatial resolution. A common bottleneck is the analysis of optically detected magnetic resonance (ODMR) spectra, where target quantities are encoded in resonance features. Conventional nonlinear fitting is often computationally expensive, sensitive to initialization, and prone to failure at low signal-to-noise ratio (SNR). Here we introduce a robust, efficient machine learning (ML) framework for real-time ODMR analysis based on a one-dimensional convolutional neural network (1D-CNN). The model performs direct parameter inference without initial guesses or iterative optimization, and is naturally parallelizable on graphics processing units (GPU) for high-throughput processing. We validate the approach on both synthetic and experimental datasets, showing improved throughput, accuracy and robustness than standard nonlinear fitting, with the largest gains in the low-SNR regime. We further validate our methods in two representative sensing applications: diagnosing intracellular temperature changes using nanodiamond probes and widefield magnetic imaging of superconducting vortices in a high-temperature superconductor. This deep-learning inference framework enables fast and reliable extraction of physical parameters from complex ODMR data and provides a scalable route to real-time quantum sensing and imaging.

quant-ph↗

Co-Layout: LLM-driven Co-optimization for Interior Layout

We present a novel framework for automated interior design that combines large language models (LLMs) with grid-based integer programming to jointly optimize room layout and furniture placement. Given a textual prompt, the LLM-driven agent workflow extracts structured design constraints related to room configurations and furniture arrangements. These constraints are encoded into a unified grid-based representation inspired by ``Modulor". Our formulation accounts for key design requirements, including corridor connectivity, room accessibility, spatial exclusivity, and user-specified preferences. To improve computational efficiency, we adopt a coarse-to-fine optimization strategy that begins with a low-resolution grid to solve a simplified problem and guides the solution at the full resolution. Experimental results across diverse scenarios demonstrate that our joint optimization approach significantly outperforms existing two-stage design pipelines in solution quality, and achieves notable computational efficiency through the coarse-to-fine strategy.

cs.CV↗

Improving the Energy and Angular Resolutions of X-ray Telescopes with Nitrogen-Vacancy Centers in Diamond

We introduce a focal-plane detector for advancing the energy and angular resolutions of current X-ray telescopes. The architecture integrates a metallic magnetic microcalorimeter (MMC) array of paramagnetic absorber pads with a thin layer of nitrogen-vacancy (NV) centers in diamond for simultaneous optical readout. An impinging X-ray photon induces a temperature transient in an absorber pad, kept at ~35 mK. This time- and temperature-dependent magnetic field transient is then optically imaged by diamond NV centers, kept at 4 K and positioned directly below the pad. For a 10 $μ$m absorber length used with a 12 m focal length telescope, our design yields an optimal angular resolution of ~0.17 arcseconds and energy resolution of ~0.70 eV. Our NV-MMC design improves upon current transition-edge sensors (TES) or MMCs read-out by superconducting quantum interference devices (SQUID) by enabling simultaneous optical readout of the entire MMC array. Because no additional cryogenic multiplexing electronics are required, our approach scales naturally to larger and finer arrays, supporting finer angular resolutions and wider fields of view.

astro-ph.IM↗

Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation

Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, while deep generative models face challenges in producing content with richness and diversity. Furthermore, approaches that utilize large language models frequently lack robustness and fail to accurately capture complex spatial relationships. To address these challenges, this paper presents a novel vision-guided 3D layout generation system. We first construct a high-quality asset library containing 2,037 scene assets and 147 3D scene layouts. Subsequently, we employ an image generation model to expand prompt representations into images, fine-tuning it to align with our asset library. We then develop a robust image parsing module to recover the 3D layout of scenes based on visual semantics and geometric information. Finally, we optimize the scene layout using scene graphs and overall visual semantics to ensure logical coherence and alignment with the images. Extensive user testing demonstrates that our algorithm significantly outperforms existing methods in terms of layout richness and quality. The code and dataset will be available at https://github.com/HiHiAllen/Imaginarium.

cs.CV↗

Universal Reconstruction of Complex Magnetic Profiles with Minimum Prior Assumptions

Understanding intricate magnetic structures in materials is essential for advancing materials science, spintronics, and geology. Recent developments of quantum-enabled magnetometers, such as nitrogen-vacancy (NV) centers in diamond, have enabled direct imaging of magnetic field distributions across a wide range of magnetic profiles. However, reconstructing the magnetization from an experimentally measured magnetic field map is a complex inverse problem, further complicated by measurement noise, finite spatial resolution, and variations in sample-to-sensor distance. In this work, we present a novel and efficient GPU-accelerated method for reconstructing spatially varying magnetization density from measured magnetic fields with minimal prior assumptions. We validate our method by simulating diverse magnetic structures under realistic experimental conditions, including multi-domain ferromagnetism and magnetic spin textures such as skyrmion, anti-skyrmion, and meron. Experimentally, we reconstruct the magnetization of a micrometer-scale Apollo lunar mare basalt (sample 10003,184) and a nanometer-scale twisted double-trilayer CrI3. The basalt exhibits soft ferromagnetic domains consistent with previous paleomagnetic studies, whereas the CrI3 system reveals a well-defined hexagonal magnetic Moire superlattice. Our approach provides a versatile and universal tool for investigating complex magnetization profiles, paving the way for future quantum sensing experiments.

cond-mat.mes-hall↗

STARC: See-Through-Wall Augmented Reality Framework for Human-Robot Collaboration in Emergency Response

In emergency response missions, first responders must navigate cluttered indoor environments where occlusions block direct line-of-sight, concealing both life-threatening hazards and victims in need of rescue. We present STARC, a see-through AR framework for human-robot collaboration that fuses mobile-robot mapping with responder-mounted LiDAR sensing. A ground robot running LiDAR-inertial odometry performs large-area exploration and 3D human detection, while helmet- or handheld-mounted LiDAR on the responder is registered to the robot's global map via relative pose estimation. This cross-LiDAR alignment enables consistent first-person projection of detected humans and their point clouds - rendered in AR with low latency - into the responder's view. By providing real-time visualization of hidden occupants and hazards, STARC enhances situational awareness and reduces operator risk. Experiments in simulation, lab setups, and tactical field trials confirm robust pose alignment, reliable detections, and stable overlays, underscoring the potential of our system for fire-fighting, disaster relief, and other safety-critical operations. Code and design will be open-sourced upon acceptance.

cs.RO↗

PERAL: Perception-Aware Motion Control for Passive LiDAR Excitation in Spherical Robots

Autonomous mobile robots increasingly rely on LiDAR-IMU odometry for navigation and mapping, yet horizontally mounted LiDARs such as the MID360 capture few near-ground returns, limiting terrain awareness and degrading performance in feature-scarce environments. Prior solutions - static tilt, active rotation, or high-density sensors - either sacrifice horizontal perception or incur added actuators, cost, and power. We introduce PERAL, a perception-aware motion control framework for spherical robots that achieves passive LiDAR excitation without dedicated hardware. By modeling the coupling between internal differential-drive actuation and sensor attitude, PERAL superimposes bounded, non-periodic oscillations onto nominal goal- or trajectory-tracking commands, enriching vertical scan diversity while preserving navigation accuracy. Implemented on a compact spherical robot, PERAL is validated across laboratory, corridor, and tactical environments. Experiments demonstrate up to 96 percent map completeness, a 27 percent reduction in trajectory tracking error, and robust near-ground human detection, all at lower weight, power, and cost compared with static tilt, active rotation, and fixed horizontal baselines. The design and code will be open-sourced upon acceptance.

cs.RO↗

Adaptive Motorized LiDAR Scanning Control for Robust Localization with OpenStreetMap

LiDAR-to-OpenStreetMap (OSM) localization has gained increasing attention, as OSM provides lightweight global priors such as building footprints. These priors enhance global consistency for robot navigation, but OSM is often incomplete or outdated, limiting its reliability in real-world deployment. Meanwhile, LiDAR itself suffers from a limited field of view (FoV), where motorized rotation is commonly used to achieve panoramic coverage. Existing motorized LiDAR systems, however, typically employ constant-speed scanning that disregards both scene structure and map priors, leading to wasted effort in feature-sparse regions and degraded localization accuracy. To address these challenges, we propose Adaptive LiDAR Scanning with OSM guidance, a framework that integrates global priors with local observability prediction to improve localization robustness. Specifically, we augment uncertainty-aware model predictive control with an OSM-aware term that adaptively allocates scanning effort according to both scene-dependent observability and the spatial distribution of OSM features. The method is implemented in ROS with a motorized LiDAR odometry backend and evaluated in both simulation and real-world experiments. Results on campus roads, indoor corridors, and urban environments demonstrate significant reductions in trajectory error compared to constant-speed baselines, while maintaining scan completeness. These findings highlight the potential of coupling open-source maps with adaptive LiDAR scanning to achieve robust and efficient localization in complex environments.

cs.RO↗