SearcharxivSearch

arXiv subjects

Qingming Liu

Publications and source records attributed to Qingming Liu.

6 recordsLinked to original sources

EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards

Video generative models are increasingly used as world models for robotics, where a model generates a future visual rollout conditioned on the current observation and task instruction, and an inverse dynamics model (IDM) converts the generated frames into executable robot actions. However, current video world models lack explicit executability constraints. As a result, visually coherent rollouts may still violate rigid-body and kinematic consistency, producing unstable or infeasible control commands when decoded by an IDM. We refer to this mismatch between visual generation and physically executable control as the executability gap. While this gap can be mitigated at inference time using techniques such as rejection sampling, such approaches are inefficient due to the high cost of video generation. In this paper, we leverage the executability gap as a training signal and introduce Executable Video Alignment (EVA), a reinforcement-learning post-training framework for aligning video world models. EVA trains an inverse dynamics model on real robot trajectories and repurposes it as a reward model that evaluates generated videos through the action sequences they induce, encouraging smooth motions measured by velocity, acceleration, and jerk while penalizing actions that violate embodiment constraints. Importantly, the reward remains informative even when generated videos contain severe visual artifacts, since such artifacts typically translate into unstable or out-of-bound actions. Experiments on the RoboTwin benchmark and a real bimanual robot show that EVA reduces embodiment-specific artifacts in generated rollouts and improves downstream task execution success.

cs.RO

PAct: Part-Decomposed Single-View Articulated Object Generation

Articulated objects are central to interactive 3D applications, including embodied AI, robotics, and VR/AR, where functional part decomposition and kinematic motion are essential. Yet producing high-fidelity articulated assets remains difficult to scale because it requires reliable part decomposition and kinematic rigging. Existing approaches largely fall into two paradigms: optimization-based reconstruction or distillation, which can be accurate but often takes tens of minutes to hours per instance, and inference-time methods that rely on template or part retrieval, producing plausible results that may not match the specific structure and appearance in the input observation. We introduce a part-centric generative framework for articulated object creation that synthesizes part geometry, composition, and articulation under explicit part-aware conditioning. Our representation models an object as a set of movable parts, each encoded by latent tokens augmented with part identity and articulation cues. Conditioned on a single image, the model generates articulated 3D assets that preserve instance-level correspondence while maintaining valid part structure and motion. The resulting approach avoids per-instance optimization, enables fast feed-forward inference, and supports controllable assembly and articulation, which are important for embodied interaction. Experiments on common articulated categories (e.g., drawers and doors) show improved input consistency, part accuracy, and articulation plausibility over optimization-based and retrieval-driven baselines, while substantially reducing inference time.

cs.CV

Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards

Generating high-quality and photorealistic 3D assets remains a longstanding challenge in 3D vision and computer graphics. Although state-of-the-art generative models, such as diffusion models, have made significant progress in 3D generation, they often fall short of human-designed content due to limited ability to follow instructions, align with human preferences, or produce realistic textures, geometries, and physical attributes. In this paper, we introduce Nabla-R2D3, a highly effective and sample-efficient reinforcement learning alignment framework for 3D-native diffusion models using 2D rewards. Built upon the recently proposed Nabla-GFlowNet method, which matches the score function to reward gradients in a principled manner for reward finetuning, our Nabla-R2D3 enables effective adaptation of 3D diffusion models using only 2D reward signals. Extensive experiments show that, unlike vanilla finetuning baselines which either struggle to converge or suffer from reward hacking, Nabla-R2D3 consistently achieves higher rewards and reduced prior forgetting within a few finetuning steps.

cs.GR

MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors

In this paper, we propose MoDGS, a new pipeline to render novel views of dy namic scenes from a casually captured monocular video. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid move ment of input cameras to construct multiview consistency but struggle to recon struct dynamic scenes on casually captured input videos whose cameras are either static or move slowly. To address this challenging task, MoDGS adopts recent single-view depth estimation methods to guide the learning of the dynamic scene. Then, a novel 3D-aware initialization method is proposed to learn a reasonable deformation field and a new robust depth loss is proposed to guide the learning of dynamic scene geometry. Comprehensive experiments demonstrate that MoDGS is able to render high-quality novel view images of dynamic scenes from just a casually captured monocular video, which outperforms state-of-the-art meth ods by a significant margin. The code is publicly available now.

cs.CV

The energy distribution structure and dynamic characteristics of energy release in electrostatic discharge process

The detail structure of energy output and the dynamic characteristics of electric spark discharge process have been studied to calculate the energy of electric spark induced plasma under different discharge condition accurately. A series of electric spark discharge experiments were conducted with the capacitor stored energy in the range of 10J 100J and 1000J respectively. And the resistance of wire, switch and plasma between electrodes were evaluated by different methods. An optimized method for electric resistance evaluation of the full discharge circuit, three poles switch and electric spark induced plasma during the discharge process was put forward. The electric energy consumed by wire, electric switch and electric spark induced plasma between electrodes were obtained by Joules law. The structure of energy distribution and the dynamic process of energy release during the capacitor discharge process have been studied. Experiments results showed that, with the increase of capacitor released energy, the duration of discharge process becomes longer, and the energy of plasma accounts for more in the capacitor released energy. The dynamic resistance of plasma and three poles switch obtained by energy conversation law is more precise than that obtained by the parameters of electric current oscillation during the discharge process.

physics.plasm-ph

Unsteady self-sustained detonation waves in flake aluminum dust/air mixtures

Self-sustained detonation waves in flake aluminum dust/air mixtures have been studied in a tube of diameter 199 mm and length 32.4 m. A pressure sensor array of 32 sensors mounted around certain circumferences of the tube was used to measure the shape of the detonation front in the circumferential direction and pressure histories of the detonation wave. A two-head spin detonation wave front was observed for the aluminum dust/air mixtures, and the cellular structure resulting from the spinning movement of the triple point was analyzed. The variations in velocity and overpressure of the detonation wave with propagation distance in a cell were studied. The interactions of waves in triple-point configurations were analyzed and the flow-field parameters were calculated. Three types of triple-point configuration exist in the wave front of the detonation wave of an aluminum dust/air mixture. Both strong and weak transverse waves exist in the unstable self-sustained detonation wave.

physics.flu-dyn