SearcharxivSearch

arXiv subjects

Tianrui Zhang

Publications and source records attributed to Tianrui Zhang.

9 recordsLinked to original sources

Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation

Generalization remains a fundamental challenge in robotic manipulation. To tackle this challenge, recent Vision-Language-Action (VLA) models build policies on top of Vision-Language Models (VLMs), seeking to transfer their open-world semantic knowledge. However, their zero-shot capability lags significantly behind the base VLMs, as the instruction-vision-action data is too limited to cover diverse scenarios, tasks, and robot embodiments. In this work, we present Goal-VLA, a zero-shot framework that leverages Image-Generative VLMs as world models to generate desired goal states, from which the target object pose is derived to enable generalizable manipulation. The key insight is that object state representation is the golden interface, naturally separating a manipulation system into high-level and low-level policies. This representation abstracts away explicit action annotations, allowing the use of highly generalizable VLMs while simultaneously providing spatial cues for training-free low-level control. To further improve robustness, we introduce a Reflection-through-Synthesis process that iteratively validates and refines the generated goal image before execution. Both simulated and real-world experiments demonstrate that our \name achieves strong performance and inspiring generalizability in manipulation tasks. Supplementary materials are available at https://nus-lins-lab.github.io/goalvlaweb/.

cs.RO

CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving

Generative models have been widely applied to world modeling for environment simulation and future state prediction. With advancements in autonomous driving, there is a growing demand not only for high-fidelity video generation under various controls, but also for producing diverse and meaningful information such as depth estimation. To address this, we propose CVD-STORM, a cross-view video diffusion model utilizing a spatial-temporal reconstruction Variational Autoencoder (VAE) that generates long-term, multi-view videos with 4D reconstruction capabilities under various control inputs. Our approach first fine-tunes the VAE with an auxiliary 4D reconstruction task, enhancing its ability to encode 3D structures and temporal dynamics. Subsequently, we integrate this VAE into the video diffusion process to significantly improve generation quality. Experimental results demonstrate that our model achieves substantial improvements in both FID and FVD metrics. Additionally, the jointly-trained Gaussian Splatting Decoder effectively reconstructs dynamic scenes, providing valuable geometric information for comprehensive scene understanding. Our project page is https://sensetime-fvg.github.io/CVD-STORM.

cs.CV

T(R,O) Grasp: Efficient Graph Diffusion of Robot-Object Spatial Transformation for Cross-Embodiment Dexterous Grasping

Dexterous grasping remains a central challenge in robotics due to the complexity of its high-dimensional state and action space. We introduce T(R,O) Grasp, a diffusion-based framework that efficiently generates accurate and diverse grasps across multiple robotic hands. At its core is the T(R,O) Graph, a unified representation that models spatial transformations between robotic hands and objects while encoding their geometric properties. A graph diffusion model, coupled with an efficient inverse kinematics solver, supports both unconditioned and conditioned grasp synthesis. Extensive experiments on a diverse set of dexterous hands show that T(R,O) Grasp achieves average success rate of 94.83%, inference speed of 0.21s, and throughput of 41 grasps per second on an NVIDIA A100 40GB GPU, substantially outperforming existing baselines. In addition, our approach is robust and generalizable across embodiments while significantly reducing memory consumption. More importantly, the high inference speed enables closed-loop dexterous manipulation, underscoring the potential of T(R,O) Grasp to scale into a foundation model for dexterous grasping.

cs.RO

RAG4ITOps: A Supervised Fine-Tunable and Comprehensive RAG Framework for IT Operations and Maintenance

With the ever-increasing demands on Question Answering (QA) systems for IT operations and maintenance, an efficient and supervised fine-tunable framework is necessary to ensure the data security, private deployment and continuous upgrading. Although Large Language Models (LLMs) have notably improved the open-domain QA's performance, how to efficiently handle enterprise-exclusive corpora and build domain-specific QA systems are still less-studied for industrial applications. In this paper, we propose a general and comprehensive framework based on Retrieval Augmented Generation (RAG) and facilitate the whole business process of establishing QA systems for IT operations and maintenance. In accordance with the prevailing RAG method, our proposed framework, named with RAG4ITOps, composes of two major stages: (1) Models Fine-tuning \& Data Vectorization, and (2) Online QA System Process. At the Stage 1, we leverage a contrastive learning method with two negative sampling strategies to fine-tune the embedding model, and design the instruction templates to fine-tune the LLM with a Retrieval Augmented Fine-Tuning method. At the Stage 2, an efficient process of QA system is built for serving. We collect enterprise-exclusive corpora from the domain of cloud computing, and the extensive experiments show that our method achieves superior results than counterparts on two kinds of QA tasks. Our experiment also provide a case for applying the RAG4ITOps to real-world enterprise-level applications.

cs.AI

Spin-torque memristors based on perpendicular magnetic tunnel junctions with a hybrid chiral texture

Spin-torque memristors were proposed in 2009, which could provide fast, low-power and infinite memristive behavior for large-density non-volatile memory and neuromorphic computing. However, the strict requirements of combining high magnetoresistance, stable intermediate states and spin-polarized current switching in a single device pose difficulties in physical implementation. Here, we experimentally demonstrate a nanoscale spin-torque memristor based on a perpendicular-anisotropy magnetic tunnel junction with a CoFeB/W/CoFeB composite free layer structure. Its tunneling magnetoresistance is higher than 200%, and memristive behavior can be realized by spin-transfer torque switching. Memristive states are maintained by robust domain wall pinning around clusters of W atoms, where nanoscale vertical chiral spin textures could be formed through the competition between opposing Dzyaloshinskii-Moriya interactions and the fluctuating interlayer coupling caused by the Ruderman-Kittel-Kasuya-Yosida interaction between the two CoFeB free layers. Spike-timing-dependent plasticity is also demonstrated in this device.

physics.app-ph

Field-free switching of perpendicular magnetic tunnel junction by the interplay of spin orbit and spin transfer torques

Spin-orbit torque and spin-transfer torque are leading the pathway to the future of spintronic memories. However, both of the mechanisms are suffering from intrinsic limitations. In particular, an external magnetic field is required for spin-orbit torque to execute deterministic switching in perpendicular magnetic tunnel junctions; the demand for reduced spin-transfer torque switching current to realize ultralow power is still remaining. Thus, a more advanced switching mechanism is urgently needed to move forward spintronics for wide applications. Here, we experimentally demonstrate the field-free switching of three-terminal perpendicular-anisotropy nanopillar devices through the interaction between spin-orbit and spin-transfer torques. The threshold current density of spin-transfer torque switching is reduced to 0.94 MA.cm-2, which is a significant decrease compared to that in other current-induced magnetization switching mechanisms. In addition, thanks to the interplay of spin-orbit and spin-transfer torques, lower current-induced switching in the conventional two-terminal perpendicular magnetic tunnel junctions is also achieved.

physics.app-ph

Assessing the risk of advanced persistent threats

As a new type of cyber attacks, advanced persistent threats (APTs) pose a severe threat to modern society. This paper focuses on the assessment of the risk of APTs. Based on a dynamic model characterizing the time evolution of the state of an organization, the organization's risk is defined as its maximum possible expected loss, and the risk assessment problem is modeled as a constrained optimization problem. The influence of different factors on an organization's risk is uncovered through theoretical analysis. Based on extensive experiments, we speculate that the attack strategy obtained by applying the hill-climbing method to the proposed optimization problem, which we call the HC strategy, always leads to the maximum possible expected loss. We then present a set of five heuristic attack strategies and, through comparative experiments, show that the HC strategy causes a higher risk than all these heuristic strategies do, which supports our conjecture. Finally, the impact of two factors on the attacker's HC cost profit is determined through computer simulations. These findings help understand the risk of APTs in a quantitative manner.

cs.CR

On the effectiveness of the truth-spreading/rumor-blocking strategy for restraining rumors

Spreading truths and blocking rumors are two typical strategies for inhibiting rumors. In practice, a tradeoff between the two strategies, which is known as the TSRB strategy, may achieve a better cost-effectiveness. This paper is devoted to assessing the effectiveness of the TSRB strategy. For that purpose, an individual-level spreading model (the generic URQT model) capturing the interaction between a rumor and the truth is established. Under the model, a set of criteria for the dying out of a rumor is presented. These criteria capture the combined influence of the basic parameters and the network structures on the effectiveness of the TSRB strategy. Experimental results show that, when the rumor dies out, the dynamics of a simplified URQT model (the linear URQT model) fits well with the actual rumor-truth interacting process. Therefore, the generic URQT model and sometimes the linear URQT model provide a proper basis for assessing the effectiveness of the TSRB strategy.

cs.SI

A discount strategy in word-of-mouth marketing and its assessment

This paper addresses the discount pricing in word-of-mouth (WOM) marketing. A new discount strategy known as the Infection-Based Discount (IBD) strategy is proposed. The basic idea of the IBD strategy lies in that each customer enjoys a discount that is linearly proportional to his/her influence in the WOM network. To evaluate the performance of the IBD strategy, the WOM spreading process is modeled as a dynamic model known as the DPA model, and the performance of the IBD strategy is modeled as a function of the basic discount. Next, the influence of different factors, including the basic discount and the WOM network, on the dynamics of the DPA model is revealed experimentally. Finally, the influence of different factors on the performance of the IBD strategy is uncovered experimentally. On this basis, some promotional measures are recommended.

cs.SI