Searcharxiv⌕ Search

arXiv subjects

Zhiyuan Qin

Publications and source records attributed to Zhiyuan Qin.

14 recordsLinked to original sources

MSEditor: Toward Consistent Multi-Shot Video Editing

In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos consist of discontinuous temporal segments that vary significantly in viewpoint, camera scale, and subject pose, leading to severe identity drift and cumulative error propagation. Achieving coherent edits requires establishing reliable cross-shot semantic awareness to maintain stable subject appearance and visual continuity across these disjointed boundaries. To address this, we propose MSEditor, the first framework designed specifically for consistent multi-shot video editing. To overcome the scarcity of high-quality multi-shot training data, we repurpose existing multi-view video datasets to provide robust cross-shot supervision. Architecturally, we introduce a Supervisory Adapter that injects this cross-shot information into the diffusion backbone, enabling the model to learn identity-consistent representations. Furthermore, to effectively mitigate cumulative errors and ensure long-range temporal coherence, we design a Cross-Shot Packing strategy that dynamically aggregates information from semantically related shots within the self-attention window. Extensive experiments demonstrate that MSEditor significantly outperforms existing methods on our curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.

cs.CV↗

High-quality Nano-patterning of Oxide Interfaces Using Transferred Gold Mask

Complex oxide interfaces, such as $\mathrm{SrTiO_3}$ and $\mathrm{KTaO_3}$ based heterostructures, host rich correlated phenomena with strong potential for advanced device applications. However, these interfaces are extremely susceptible to contamination and defect formation during nanofabrication, which often compromises device performance. Here, we present a solvent-free method for patterning oxide interfaces by employing high-resolution transferable thin metal masks in conjunction with oxygen-enriched $\mathrm{Ar^+}$ ion milling, which enables a clean and well-controlled nanofabrication process. Transport measurements demonstrate that the fabricated devices preserve their intrinsic properties, including high carrier mobilities, with negligible degradation compared to the pristine interfaces. This technique offers a convenient and robust route for engineering high-performance oxide electronic devices with precisely tailored transport characteristics.

cond-mat.mtrl-sci↗

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs

Multimodal Large Language Models (MLLMs) have significantly advanced the landscape of embodied AI, yet transitioning to synchronized bimanual coordination introduces formidable challenges in multi-stream multimodal integration. We introduce ST-BiBench, a comprehensive multi-tier framework for evaluating spatio-temporal multimodal coordination. Our approach centers on Strategic Coordination Planning, assessing high-level cross-modal reasoning over multiple action and perception streams. To investigate the "proximity paradox"-where semantically coherent plans fail to align with spatially grounded visual inputs-we incorporate Foundational Spatial Grounding to verify workspace awareness and arm-selection logic. Furthermore, we probe model frontiers through Fine-Grained Action Control, investigating whether MLLMs can directly synthesize high-dimensional continuous action modalities (16-Dim) from complex multimodal metadata. Evaluating 30+ state-of-the-art MLLMs, we uncover a persistent and pervasive "coordination paradox"-a significant gap between high-level strategic reasoning and fine-grained physical execution. Results reveal that while frontier MLLMs excel at logic-driven strategy, they frequently suffer from perception-logic disconnection and multi-stream interference during multimodal fusion. ST-BiBench provides a platform for identifying critical bottlenecks in multi-stream multimodal fusion and cross-modal alignment for complex embodied tasks.

cs.RO↗

TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion

The vision-language-action (VLA) paradigm has enabled powerful robotic control by leveraging vision-language models, but its reliance on large-scale, high-quality robot data limits its generalization. Generative world models offer a promising alternative for general-purpose embodied AI, yet a critical gap remains between their pixel-level plans and physically executable actions. To this end, we propose the Tool-Centric Inverse Dynamics Model (TC-IDM). By focusing on the tool's imagined trajectory as synthesized by the world model, TC-IDM establishes a robust intermediate representation that bridges the gap between visual planning and physical control. TC-IDM extracts the tool's point cloud trajectories via segmentation and 3D motion estimation from generated videos. Considering diverse tool attributes, our architecture employs decoupled action heads to project these planned trajectories into 6-DoF end-effector motions and corresponding control signals. This plan-and-translate paradigm not only supports a wide range of end-effectors but also significantly improves viewpoint invariance. Furthermore, it exhibits strong generalization capabilities across long-horizon and out-of-distribution tasks, including interacting with deformable objects. In real-world evaluations, the world model with TC-IDM achieves an average success rate of 61.11 percent, with 77.7 percent on simple tasks and 38.46 percent on zero-shot deformable object tasks. It substantially outperforms end-to-end VLA-style baselines and other inverse dynamics models.

cs.RO↗

Reconfigurable Oxide Nanoelectronics by Tip-induced Electron Delocalization

Reconfigurable oxide nanoelectronics, enabled by conductive atomic force microscope (cAFM) lithography, have established complex oxide interfaces as a promising platform for quantum engineering that harnesses emergent phenomena for advanced functionalities. However, this cAFM nanofabrication process can only occur in the air, with simultaneous device decay described under the "water-cycle" writing mechanism. These restrictions pose ongoing challenges for device optimization in the quantum regime at mK temperatures. Here, we demonstrate a "waterless" cAFM lithography approach that is compatible with vacuum and cryogenic environments. Through oxygen vacancy engineering at the LaAlO$_3$/SrTiO$_3$ interface, we have achieved nonvolatile and reconfigurable cAFM control of nanoscale interfacial polaron-electron liquid transition at mK temperatures with an ultrafine line resolution of 0.85 nm. Supported by first-principles calculations and drift-diffusion modeling, we show that tip-controlled oxygen vacancy electromigration plays a key role. This advancement bridges reconfigurable device fabrication and concurrent characterization in situ at mK temperatures, and establishes a versatile Hubbard toolbox for engineering programmable quantum phases in correlated oxides.

cond-mat.mes-hall↗

Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test

As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D prediction or interactive generation. However, before exploring these downstream tasks, video foundation models still have two critical questions unanswered: (1) whether their generative generalization is sufficient to maintain perceptual fidelity in the eyes of human observers, and (2) whether they are robust enough to serve as a universal prior for real-world embodied agents. To provide a standardized framework for answering these questions, we introduce the Embodied Turing Test benchmark: WoW-World-Eval (Wow,wo,val). Building upon 609 robot manipulation data, Wow-wo-val examines five core abilities, including perception, planning, prediction, generalization, and execution. We propose a comprehensive evaluation protocol with 22 metrics to assess the models' generation ability, which achieves a high Pearson Correlation between the overall score and human preference (>0.93) and establishes a reliable foundation for the Human Turing Test. On Wow-wo-val, models achieve only 17.27 on long-horizon planning and at best 68.02 on physical consistency, indicating limited spatiotemporal consistency and physical reasoning. For the Inverse Dynamic Model Turing Test, we first use an IDM to evaluate the video foundation models' execution accuracy in the real world. However, most models collapse to $\approx$ 0% success, while WoW maintains a 40.74% success rate. These findings point to a noticeable gap between the generated videos and the real world, highlighting the urgency and necessity of benchmarking World Model in Embodied AI.

cs.RO↗

SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents

Self-evolution, the ability of agents to autonomously improve their reasoning and behavior, is essential for the embodied domain with long-horizon, real-world tasks. Despite current advancements in reinforcement fine-tuning (RFT) showing strong performance in enhancing reasoning in LLMs, its potential to enable self-evolving embodied intelligence with multi-modal interactions remains largely unexplored. Specifically, reinforcement fine-tuning faces two fundamental obstacles in embodied settings: (i) the lack of accessible intermediate rewards in multi-step reasoning tasks limits effective learning signals, and (ii) reliance on hand-crafted reward functions restricts generalization to novel tasks and environments. To address these challenges, we present Self-Evolving Embodied Agents-R1, SEEA-R1, the first RFT framework designed for enabling the self-evolving capabilities of embodied agents. Specifically, to convert sparse delayed rewards into denser intermediate signals that improve multi-step reasoning, we propose Tree-based group relative policy optimization (Tree-GRPO) integrates Monte Carlo Tree Search into GRPO. To generalize reward estimation across tasks and scenes, supporting autonomous adaptation and reward-driven self-evolution, we further introduce Multi-modal Generative Reward Model (MGRM). To holistically evaluate the effectiveness of SEEA-R1, we evaluate on the ALFWorld benchmark, surpassing state-of-the-art methods with scores of 85.07% (textual) and 46.27% (multi-modal), outperforming prior models including GPT-4o. SEEA-R1 also achieves scores of 80.3% (textual) and 44.03% (multi-modal) without ground truth reward, surpassing all open-source baselines and highlighting its scalability as a self-evolving embodied agent. Additional experiments and qualitative analysis further support the potential of SEEA-R1 for future research in scalable embodied intelligence.

cs.AI↗

WoW: Towards a World omniscient World model Through Embodied Interaction

Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping physical causality. This observation leads to our central hypothesis: authentic physical intuition of the world model must be grounded in extensive, causally rich interactions with the real world. To test this hypothesis, we present WoW, a 14-billion-parameter generative world model trained on 2 million robot interaction trajectories. Our findings reveal that the model's understanding of physics is a probabilistic distribution of plausible outcomes, leading to stochastic instabilities and physical hallucinations. Furthermore, we demonstrate that this emergent capability can be actively constrained toward physical realism by SOPHIA, where vision-language model agents evaluate the DiT-generated output and guide its refinement by iteratively evolving the language instructions. In addition, a co-trained Inverse Dynamics Model translates these refined plans into executable robotic actions, thus closing the imagination-to-action loop. We establish WoWBench, a new benchmark focused on physical consistency and causal reasoning in video, where WoW achieves state-of-the-art performance in both human and autonomous evaluation, demonstrating strong ability in physical causality, collision dynamics, and object permanence. Our work provides systematic evidence that large-scale, real-world interaction is a cornerstone for developing physical intuition in AI. Models, data, and benchmarks will be open-sourced.

cs.RO↗

WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation

Wrist-view observations are crucial for VLA models as they capture fine-grained hand-object interactions that directly enhance manipulation performance. Yet large-scale datasets rarely include such recordings, resulting in a substantial gap between abundant anchor views and scarce wrist views. Existing world models cannot bridge this gap, as they require a wrist-view first frame and thus fail to generate wrist-view videos from anchor views alone. Amid this gap, recent visual geometry models such as VGGT emerge with geometric and cross-view priors that make it possible to address extreme viewpoint shifts. Inspired by these insights, we propose WristWorld, the first 4D world model that generates wrist-view videos solely from anchor views. WristWorld operates in two stages: (i) Reconstruction, which extends VGGT and incorporates our Spatial Projection Consistency (SPC) Loss to estimate geometrically consistent wrist-view poses and 4D point clouds; (ii) Generation, which employs our video generation model to synthesize temporally coherent wrist-view videos from the reconstructed perspective. Experiments on Droid, Calvin, and Franka Panda demonstrate state-of-the-art video generation with superior spatial consistency, while also improving VLA performance, raising the average task completion length on Calvin by 3.81% and closing 42.4% of the anchor-wrist view gap.

cs.CV↗

Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis

With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real-world data remains costly and time-consuming, hindering the development of downstream applications. While some works attempt to collect task-specific data via a rendering process, most approaches still rely on manual scene construction, limiting their scalability and accuracy. To address these challenges, we propose Follow-Your-Instruction, a Multimodal Large Language Model (MLLM)-driven framework for automatically synthesizing high-quality 2D, 3D, and 4D data. Our \textbf{Follow-Your-Instruction} first collects assets and their associated descriptions through multimodal inputs using the MLLM-Collector. Then it constructs 3D layouts, and leverages Vision-Language Models (VLMs) for semantic refinement through multi-view scenes with the MLLM-Generator and MLLM-Optimizer, respectively. Finally, it uses MLLM-Planner to generate temporally coherent future frames. We evaluate the quality of the generated data through comprehensive experiments on the 2D, 3D, and 4D generative tasks. The results show that our synthetic data significantly boosts the performance of existing baseline models, demonstrating Follow-Your-Instruction's potential as a scalable and effective data engine for generative intelligence.

cs.CV↗

Time-Reversal Symmetry Protected Transport at Correlated Oxide Interfaces

Time-reversal symmetry (TRS) protection is core to topological physics, yet its role in correlated oxides-typically non-topological-remains underexplored. This limit hampers the potential in engineering exotic quantum states by fusing TRS protection and the rich emergent phenomena in the oxide platform. Here, we report evidence of a TRS-protected subband at oxygen vacancy-free LaAlO3/SrTiO3 interfaces. This subband causes a low-field quantum oscillation with anomalous characters: exceptionally light electron mass, aperiodicity, and susceptibility to magnetic fields. All findings align with a Rashba model in which TRS-protected transport occurs along quasi-1D ferroelastic domain walls, which possess a Dirac band topology and a giant Rashba spin-orbit coupling, two orders stronger than the 2D interface. Our results deepen the understanding of SrTiO3-based electron systems, unveiling an appealing new platform for quantum engineering.

cond-mat.mes-hall↗

EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, we propose EmbodiedVSR (Embodied Visual Spatial Reasoning), a novel framework that integrates dynamic scene graph-guided Chain-of-Thought (CoT) reasoning to enhance spatial understanding for embodied agents. By explicitly constructing structured knowledge representations through dynamic scene graphs, our method enables zero-shot spatial reasoning without task-specific fine-tuning. This approach not only disentangles intricate spatial relationships but also aligns reasoning steps with actionable environmental dynamics. To rigorously evaluate performance, we introduce the eSpatial-Benchmark, a comprehensive dataset including real-world embodied scenarios with fine-grained spatial annotations and adaptive task difficulty levels. Experiments demonstrate that our framework significantly outperforms existing MLLM-based methods in accuracy and reasoning coherence, particularly in long-horizon tasks requiring iterative environment interaction. The results reveal the untapped potential of MLLMs for embodied intelligence when equipped with structured, explainable reasoning mechanisms, paving the way for more reliable deployment in real-world spatial applications. The codes and datasets will be released soon.

cs.RO↗

A Milli-Kelvin Atomic Force Microscope Made of Glass

Milli-Kelvin atomic force microscopy (mK-AFM) presents an ongoing experimental challenge due to the intense vibrations in a cryogen-free dilution refrigerator and the low cooling power available at mK temperatures. A viable approach is to make the system exceptionally rigid and thermally insulating to decouple external vibrations and isolate heat dissipation from the piezo elements. Here, we present a low-cost and large scan-range mK-AFM that operates below 100 mK. All the essential parts of our mK-AFM, including the scanners, tip assembly, and microscope body, are custom-made of fused silica glass by taking advantage of its high specific modulus, extremely low thermal expansion coefficient, and excellent thermal insulation properties. We carefully balance the scan range (25 $μ$m $\times$ 25 $μ$m), heat dissipation, and stiffness of the system to reach optimal performance at mK temperatures.

physics.ins-det↗

Domain wall nature of sketched LaAlO3/SrTiO3 nanowires

The rich electron correlations and highly coherent transport in reconfigurable devices sketched by a conductive atomic force microscope tip at the LaAlO3/SrTiO3 interface have enabled the oxide platform an ideal playground for studying correlated electrons and quantum technological applications. Why these one-dimensional devices possess enhanced properties over the two-dimensional interface, however, has remained elusive. Here we provide evidence that one-dimensional LaAlO3/SrTiO3 nanowires are intrinsically ferroelastic domain walls by nature through thermodynamic study. We have observed spreading resistance anomalies under thermo-stimulus and temperature cycles, with characteristic temperatures matching domain wall polarity. This information is crucial in understanding the novel phenomena including superconductivity and high mobility quantum transport.

cond-mat.mes-hall↗