Searcharxiv⌕ Search

arXiv subjects

Zhongxiang Zhou

Publications and source records attributed to Zhongxiang Zhou.

At least 19 recordsLinked to original sources

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a significant challenge. In this paper, we present a generalizable deformation learning framework that reconstructs 3D objects by explicitly deforming a category-level shape template to match the target observation. To address complex shape variations between the template and the target, we introduce a geometry-guided feature modeling mechanism. This process first enriches foundation features with template topology to yield a geometry-aware representation, which is then explicitly correlated with the target observation to guide precise deformation. Furthermore, to bridge the disparity between the fixed template and arbitrary target views, we propose a view-adaptive feature aggregation module. This module leverages multi-view template features and their corresponding camera poses to enrich the canonical template representation, ensuring robust feature alignment regardless of the target's perspective. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods in handling large shape variations and diverse viewpoints, exhibiting strong generalization to novel categories and effectively supporting downstream real-world dexterous robotic manipulation tasks. Project homepage: https://GODeform.github.io/

cs.CV↗

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).

cs.RO↗

Spatiotemporal Co-reflection with Spacetime Discontinuities at Moving Interfaces

The control of reflection and refraction at interfaces using engineered media is central to numerous optical technologies, with negative refraction and the suppression of backscattering representing two prominent research frontiers. In this work, we demonstrate that an effective negative refraction accompanied by an absence of backscattering can be realized at a moving spatiotemporal interface when temporal and spatial reflections occur concurrently. While such spatiotemporal co-reflection is prohibited in one-dimensional linear dispersive media, we show that it becomes permissible under oblique incidence within a specific range of traveling-wave modulation velocities. Leveraging this mechanism, we propose a spatiotemporal flat lens capable of nonreciprocal electromagnetic wave focusing. These findings provide a framework for developing advanced spatiotemporal metamaterials and time-varying metasurfaces.

physics.optics↗

Breaking the Limitations of Temporal Modulation via Mixed Continuity Conditions

The conventional description of time-varying media assumes that electromagnetic fields evolve according to fixed continuity conditions during parameter jumps. Here we reveal that these conditions are not physical constraints but tunable design degrees of freedom. By developing a unified framework that treats continuity rules as engineerable parameters, we expand the scope of time-varying metamaterials and enable wave phenomena previously considered impossible. For instance, non-resonant, reflectionless wave amplification without momentum bandgaps, and reversible conversion between propagating waves and static fields for optical memory, etc. This work opens a new dimension for controlling light-matter interactions.

physics.optics↗

Topological Braiding and Dynamic Probing of Phase Transitions at Temporal Interfaces in Non-Hermitian Synthetic Dimensions

Non-Hermitian systems give rise to distinct topological phenomena, yet their manifestations at temporal interfaces characterized by abrupt changes in system parameters remain largely unex plored. Upon an abrupt alteration of the Hamiltonian in a one-dimensional non-Hermitian sys tem,the ensuring temporal interface excites both reflected and refracted wave modes. By intro ducing a chiral-symmetric Hamiltonian, this study reveals the topological effects at such temporal interfaces. We find that the reflection and refraction coefficients exhibit a topological braiding struc ture. This structure is directly determined by the difference in the topological invariants across the interface, establishing a bulk-boundary correspondence for temporal interfaces in non-Hermitian systems. Furthermore, we propose a dynamical probe that leverages the geometric similarity of eigenstates at the temporal interface to detect topological phase transitions. These findings estab lish a fundamental connection between topological braiding and nonreciprocal dynamics at temporal interfaces, providing a platform to explore phase transition detection and nonreciprocal phenomena in time-varying non-Hermitian systems.

physics.optics↗

Guest metal-driven quantum anharmonic effects on stability and two-gap superconductivity in carbon-boron clathrates

Traditionally, strong quantum anharmonic effects have been considered a characteristic of hydrogen-rich compounds. Here we propose that these effects also play a decisive role in boron-carbon clathrates. The stability and superconducting transition temperature (Tc) of carbon-boron clathrates XYB6C6, whose metal atoms have an average oxidation state of +1.5, have long remained under debate. At this oxidation state, some combinations (e.g., RbSrB6C6) are dynamically stable, whereas others (e.g., RbPbB6C6) are not. Using the stochastic self-consistent harmonic approximation combined with machine learning, we find that the anharmonicity originates primarily from guest metal atoms. For comparison, we find that quantum fluctuations have negligible influence on SrB3C3, but remove the lattice instability of RbPbB6C6. The predicted Tc of RbPbB6C6 (88 K) is nearly twice that of SrB3C3. Moreover, RbPbB6C6 exhibits two-gap superconductivity due to the higher C/B ratio in the density of states at the Fermi level compared to SrB3C3, weakening the sp3 hybridization. These findings demonstrate that quantum anharmonicity crucially governs the stability and superconductivity of XYB6C6 clathrates.

cond-mat.supr-con↗

Toward Embodiment Equivariant Vision-Language-Action Policy

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying less attention to the design of action spaces. This leads to the configuration generalization problem, which requires costly adaptation. We address this challenge by formulating cross-embodiment pre-training as designing policies equivariant to embodiment configuration transformations. Building on this principle, we propose a framework that (i) establishes a embodiment equivariance theory for action space and policy design, (ii) introduces an action decoder that enforces configuration equivariance, and (iii) incorporates a geometry-aware network architecture to enhance embodiment-agnostic spatial reasoning. Extensive experiments in both simulation and real-world settings demonstrate that our approach improves pre-training effectiveness and enables efficient fine-tuning on novel robot embodiments. Our code is available at https://github.com/hhcaz/e2vla

cs.RO↗

ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models

The advancement of embodied intelligence is accelerating the integration of robots into daily life as human assistants. This evolution requires robots to not only interpret high-level instructions and plan tasks but also perceive and adapt within dynamic environments. Vision-Language Models (VLMs) present a promising solution by combining visual understanding and language reasoning. However, existing VLM-based methods struggle with interactive exploration, accurate perception, and real-time plan adaptation. To address these challenges, we propose ExploreVLM, a novel closed-loop task planning framework powered by Vision-Language Models (VLMs). The framework is built around a step-wise feedback mechanism that enables real-time plan adjustment and supports interactive exploration. At its core is a dual-stage task planner with self-reflection, enhanced by an object-centric spatial relation graph that provides structured, language-grounded scene representations to guide perception and planning. An execution validator supports the closed loop by verifying each action and triggering re-planning. Extensive real-world experiments demonstrate that ExploreVLM significantly outperforms state-of-the-art baselines, particularly in exploration-centric tasks. Ablation studies further validate the critical role of the reflective planner and structured perception in achieving robust and efficient task execution.

cs.RO↗

TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots

Humanoid robots have the potential capability to perform a diverse range of manipulation tasks, but this is based on a robust and precise standing controller. Existing methods are either ill-suited to precisely control high-dimensional upper-body joints, or difficult to ensure both robustness and accuracy, especially when upper-body motions are fast. This paper proposes a novel time optimization policy (TOP), to train a standing manipulation control model that ensures balance, precision, and time efficiency simultaneously, with the idea of adjusting the time trajectory of upper-body motions but not only strengthening the disturbance resistance of the lower-body. Our approach consists of three parts. Firstly, we utilize motion prior to represent upper-body motions to enhance the coordination ability between the upper and lower-body by training a variational autoencoder (VAE). Then we decouple the whole-body control into an upper-body PD controller for precision and a lower-body RL controller to enhance robust stability. Finally, we train TOP method in conjunction with the decoupled controller and VAE to reduce the balance burden resulting from fast upper-body motions that would destabilize the robot and exceed the capabilities of the lower-body RL policy. The effectiveness of the proposed approach is evaluated via both simulation and real world experiments, which demonstrate the superiority on standing manipulation tasks stably and accurately. The project page can be found at https://anonymous.4open.science/w/top-258F/.

cs.RO↗

A new approach for solving the problem of creation of inverse electron distribution function and practical recommendations for experimental searches for such media in glow discharges with hollow and flat cathodes

This paper proposes a novel approach for creating an inverse electron distribution function (EDF). Based on the obtained criteria for the formation of an inverse EDF in a non-uniform plasma, studies are conducted in low- and medium-pressure glow discharges with flat and hollow cathodes. The results of the numerical modeling and theoretical analysis are used to present reliable criteria and scaling for the evaluation of the possible inversion of the EDF under specific conditions. By solving the nonlocal Boltzmann kinetic equation in energy and coordinate variables, it is shown that the simplest way to implement the inversion of the EDF is in a glow discharge with a hollow cathode. For such discharges, practical recommendations are developed and specific conditions for the experimental detection of an inverse EDF are identified.

physics.plasm-ph↗

Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning

Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data, despite its low costs, is limited by the sim-to-real gap. We identify the root cause of gripper state ambiguity as the lack of tactile feedback. To address this, we propose a novel approach employing pseudo-tactile as feedback, inspired by the idea of using a force-controlled gripper as a tactile sensor. This method enhances policy robustness without additional data collection and hardware involvement, while providing a noise-free binary gripper state observation for the policy and thus facilitating pure simulation learning to unleash the power of simulation. Experimental results across three real-world grasp-based tasks demonstrate the necessity, effectiveness, and efficiency of our approach.

cs.RO↗

CNSv2: Probabilistic Correspondence Encoded Neural Image Servo

Visual servo based on traditional image matching methods often requires accurate keypoint correspondence for high precision control. However, keypoint detection or matching tends to fail in challenging scenarios with inconsistent illuminations or textureless objects, resulting significant performance degradation. Previous approaches, including our proposed Correspondence encoded Neural image Servo policy (CNS), attempted to alleviate these issues by integrating neural control strategies. While CNS shows certain improvement against error correspondence over conventional image-based controllers, it could not fully resolve the limitations arising from poor keypoint detection and matching. In this paper, we continue to address this problem and propose a new solution: Probabilistic Correspondence Encoded Neural Image Servo (CNSv2). CNSv2 leverages probabilistic feature matching to improve robustness in challenging scenarios. By redesigning the architecture to condition on multimodal feature matching, CNSv2 achieves high precision, improved robustness across diverse scenes and runs in real-time. We validate CNSv2 with simulations and real-world experiments, demonstrating its effectiveness in overcoming the limitations of detector-based methods in visual servo tasks.

cs.CV↗

Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior

We focus on the task of unknown object rearrangement, where a robot is supposed to re-configure the objects into a desired goal configuration specified by an RGB-D image. Recent works explore unknown object rearrangement systems by incorporating learning-based perception modules. However, they are sensitive to perception error, and pay less attention to task-level performance. In this paper, we aim to develop an effective system for unknown object rearrangement amidst perception noise. We theoretically reveal that the noisy perception impacts grasp and place in a decoupled way, and show such a decoupled structure is valuable to improve task optimality. We propose GSP, a dual-loop system with the decoupled structure as prior. For the inner loop, we learn a see policy for self-confident in-hand object matching. For the outer loop, we learn a grasp policy aware of object matching and grasp capability guided by task-level rewards. We leverage the foundation model CLIP for object matching, policy learning and self-termination. A series of experiments indicate that GSP can conduct unknown object rearrangement with higher completion rates and fewer steps.

cs.RO↗

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works require object labels or visual attributes for grounding, which calls for handcrafted rules in planner and restricts the range of language instructions. In this paper, we propose to jointly model vision, language and action with object-centric representation. Our method is applicable under more flexible language instructions, and not limited by visual grounding error. Besides, by utilizing the powerful priors from the pre-trained multi-modal model and grasp model, sample efficiency is effectively improved and the sim2real problem is relived without additional data for transfer. A series of experiments carried out in simulation and real world indicate that our method can achieve better task success rate by less times of motion under more flexible language instructions. Moreover, our method is capable of generalizing better to scenarios with unseen objects and language instructions. Our code is available at https://github.com/xukechun/Vision-Language-Grasping

cs.RO↗

Topological States Decorated by Twig Boundary in Plasma Photonic Crystals

The twig edge states in graphene-like structures are viewed as the fourth states complementary to their zigzag, bearded, and armchair counterparts. In this work, we study a rod-in-plasma system in honeycomb lattice with twig edge truncation under external magnetic fields and lattice scaling and show that twig edge states can exist in different phases of the system, such as quantum Hall phase, quantum spin Hall phase and insulating phase. The twig edge states in the negative permittivity background exhibit robust one-way transmission property immune to backscattering and thus provide a novel avenue for solving the plasma communication blackout problem. Moreover, we demonstrate that corner and edge states can exist within the shrunken structure by modulating the on-site potential of the twig edges. Especially, helical edge states with the unique feature of pseudospin-momentum locking that could be excited by chiral sources are demonstrated at the twig edges. Our results show that the twig edges and interface engineering can bring new opportunities for more flexible manipulation of electromagnetic waves.

cond-mat.mes-hall↗

Open-Set Object Detection Using Classification-free Object Proposal and Instance-level Contrastive Learning

Detecting both known and unknown objects is a fundamental skill for robot manipulation in unstructured environments. Open-set object detection (OSOD) is a promising direction to handle the problem consisting of two subtasks: objects and background separation, and open-set object classification. In this paper, we present Openset RCNN to address the challenging OSOD. To disambiguate unknown objects and background in the first subtask, we propose to use classification-free region proposal network (CF-RPN) which estimates the objectness score of each region purely using cues from object's location and shape preventing overfitting to the training categories. To identify unknown objects in the second subtask, we propose to represent them using the complementary region of known categories in a latent space which is accomplished by a prototype learning network (PLN). PLN performs instance-level contrastive learning to encode proposals to a latent space and builds a compact region centering with a prototype for each known category. Further, we note that the detection performance of unknown objects can not be unbiasedly evaluated on the situation that commonly used object detection datasets are not fully annotated. Thus, a new benchmark is introduced by reorganizing GraspNet-1billion, a robotic grasp pose detection dataset with complete annotation. Extensive experiments demonstrate the merits of our method. We finally show that our Openset RCNN can endow the robot with an open-set perception ability to support robotic rearrangement tasks in cluttered environments. More details can be found in https://sites.google.com/view/openset-rcnn/

cs.CV↗

Learning adaptive manipulation of objects with revolute joint: A case study on varied cabinet doors opening

This paper introduces a learning-based framework for robot adaptive manipulating the object with a revolute joint in unstructured environments. We concentrate our discussion on various cabinet door opening tasks. To improve the performance of Deep Reinforcement Learning in this scene, we analytically provide an efficient sampling manner utilizing the constraints of the objects. To open various kinds of doors, we add encoded environment parameters that define the various environments to the input of out policy. To transfer the policy into the real world, we train an adaptation module in simulation and fine-tune the adaptation module to cut down the impact of the policy-unaware environment parameters. We design a series of experiments to validate the efficacy of our framework. Additionally, we testify to the model's performance in the real world compared to the traditional door opening method.

cs.RO↗