SearcharxivSearch

arXiv subjects

Xiaotao Zhang

Publications and source records attributed to Xiaotao Zhang.

9 recordsLinked to original sources

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations should evolve under its predicted motion. Generative world-action models (WAMs) jointly predict future observations and actions, yet existing WAMs for continuous VLN do not condition joint future-view and action generation on geometry-aware representations inferred from the observed history. We present WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN. To consolidate past observations into persistent scene context, a frozen feed-forward geometry encoder extracts geometry-aware representations from the monocular egocentric RGB history, and a trainable 3D Scene-to-Token Adapter converts them into a fixed-length prefix in the token space of the world-action Diffusion Transformer. Through block-causal attention, this prefix conditions every future video-action block, providing a shared geometric context for both future-view and action generation. We train WNM-3D through supervised world-action fine-tuning on A*-generated demonstrations, DAgger-style adaptation on policy-visited states, and Counterfactual DanceGRPO refinement for closed-loop execution. Experiments on GN-Bench show that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation. Stage-wise ablations further show that DAgger-SFT provides the larger success-rate gain, while Counterfactual DanceGRPO subsequently improves both navigation success and path efficiency.

cs.AI

GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation

Embodied navigation connects intelligent agents with the physical world and is fundamental for general robotic intelligence. Limited availability and quality of navigation data have constrained Vision-and-Language Navigation (VLN) systems' generalization and long-horizon capabilities. To address this, we curate diverse 3D scenes and develop an automated pipeline for large-scale navigation data, resulting in the GN-Matrix dataset. Building on a 3D Gaussian Splatting (3DGS) engine, we introduce a high-fidelity simulation platform supporting interactive roaming and collision-aware navigation. We further propose GN-Bench, the first BEV-based benchmark incorporating dynamic 3DGS avatars for human-robot interaction evaluation. To leverage the simulator, we develop an RL-driven navigation foundation model, Break and Establish (BAE). After supervised learning, DAgger exposes the model to rollout-induced states, breaking narrow expert-centric distributions and enabling downstream RL exploration. This unified VLN paradigm integrates map-based and map-free tasks, including instruction following, human following, and goal navigation. GN-BAE formalizes high-fidelity 3DGS-rendered Bird's Eye View representations as compact memory, unlocking latent spatial reasoning in VLMs. Extensive evaluations on GN-Bench and VLN-CE show that GN0 outperforms state-of-the-art VLN methods. Overall, GN-Matrix offers a unified framework spanning data, simulation, and learning, advancing embodied navigation in research and industrial applications.

cs.RO

AcademiClaw: When Students Set Challenges for AI Agents

Benchmarks within the OpenClaw ecosystem have thus far evaluated exclusively assistant-level tasks, leaving the academic-level capabilities of OpenClaw largely unexamined. We introduce AcademiClaw, a bilingual benchmark of 80 complex, long-horizon tasks sourced directly from university students' real academic workflows -- homework, research projects, competitions, and personal projects -- that they found current AI agents unable to solve effectively. Curated from 230 student-submitted candidates through rigorous expert review, the final task set spans 25+ professional domains, ranging from olympiad-level mathematics and linguistics problems to GPU-intensive reinforcement learning and full-stack system debugging, with 16 tasks requiring CUDA GPU execution. Each task executes in an isolated Docker sandbox and is scored on task completion by multi-dimensional rubrics combining six complementary techniques, with an independent five-category safety audit providing additional behavioral analysis. Experiments on six frontier models show that even the best achieves only a 55\% pass rate. Further analysis uncovers sharp capability boundaries across task domains, divergent behavioral strategies among models, and a disconnect between token consumption and output quality, providing fine-grained diagnostic signals beyond what aggregate metrics reveal. We hope that AcademiClaw and its open-sourced data and code can serve as a useful resource for the OpenClaw community, driving progress toward agents that are more capable and versatile across the full breadth of real-world academic demands. All data and code are available at https://github.com/GAIR-NLP/AcademiClaw.

cs.AI

From the Landau-de Gennes theory to the Ericksen-Leslie theory in dimension two

In this paper, we study the connection between the Ericksen-Leslie equations and the Beris-Edwards equations in dimension two. It is shown that the weak solutions to the Beris-Edwards equations converge to the one to the Ericksen-Leslie equations as the elastic coefficient tends to zero. Moreover, the limiting weak solutions to the Ericksen-Leslie equations may have singular points.

math.AP

Research progress of rubrene as an excellent multifunctional organic semiconductor

Rubrene, a superstar in organic semiconductors, has achieved unprecedented achievements in the application of electronic devices, and research based on its various photoelectric properties is still in progress. In this review, we introduced the preparation of rubrene crystal, summarized the applications in organic optoelectronic devices with the latest research achievements based on rubrene semiconductors. An outlook of future research directions and challenges of rubrene semiconductor for applications is also provided.

physics.app-ph

Local well-posedness of isentropic compressible Navier-Stokes equations with vacuum

In this paper, the local well-posedness of strong solutions to the Cauchy problem of the isentropic compressible Navier-Stokes equations is proved with the initial date being allowed to have vacuum. The main contribution of this paper is that the well-posedness is established without assuming any compatibility condition on the initial data, which was widely used before in many literatures concerning the well-posedness of compressible Navier-Stokes equations in the presence of vacuum.

math.AP

On optimal boundary control of Ericksen-Leslie system in dimension two

In this paper, we consider the boundary value problem of a simplified Ericksen-Leslie system in dimension two with non-slip boundary condition for the velocity field $u$ and time-dependent boundary condition for the director field $d$ of unit length. For such a system, we first establish the existence of a global weak solution that is smooth away from finitely many singular times, then establish the existence of a unique global strong solution that is smooth for $t>0$ under the assumption that the image of boundary data is contained in the hemisphere $\mathbb S^2_+$. Finally, we apply these theorems to the problem of optimal boundary control of the simplified Ericksen-Leslie system and show both the existence and a necessary condition of an optimal boundary control.

math.AP