SearcharxivSearch

arXiv subjects

Jiaxi Song

Publications and source records attributed to Jiaxi Song.

6 recordsLinked to original sources

A conservative coupling method of sharp-interface and multi-species model for compressible reacting gas-liquid flows with phase change

In this paper, a conservative sharp-interface and diffuse-interface coupling method is developed for compressible two-phase multi-species flows with phase change and chemical reactions. The liquid--gas interface is represented by a sharp-interface model, whereas a diffuse-interface model treats the transport and chemical reactions of gas-phase species. Conservation is enforced by coupling the two phases through interfacial fluxes obtained from a multi-species phase-change Riemann problem. The original single-species four-wave Riemann solver is extended to multi-species gas mixtures by modifying the interfacial energy jump condition. Interfacial mass transfer is restricted to the condensable vapor species. Accordingly, both the interfacial energy jump condition and the gas-mixture energy-exchange flux are constructed using the internal energy of the phase-changing vapor species rather than the mixture internal energy. With this species-selective energy coupling, an approximate multi-species Riemann solver is constructed that retains the four-wave structure while avoiding the multidimensional nonlinear root-finding required by the exact solution. A series of numerical tests, including impulsive evaporation and condensation, reacting aluminum vaporization, shock-droplet interaction, and detonation-droplet interaction, are performed to assess the accuracy and robustness of the method. The numerical results agree well with reference solutions and benchmark data, demonstrating that the present method resolves the effects of phase change and chemical reactions in compressible multi-species multiphase flows while preserving conservative interfacial coupling.

physics.flu-dyn

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.

cs.RO

Scaling World Model for Hierarchical Manipulation Policies

Vision-Language-Action (VLA) models are promising for generalist robot manipulation but remain brittle in out-of-distribution (OOD) settings, especially with limited real-robot data. To resolve the generalization bottleneck, we introduce a hierarchical Vision-Language-Action framework \our{} that leverages the generalization of large-scale pre-trained world model for robust and generalizable VIsual Subgoal TAsk decomposition VISTA. Our hierarchical framework \our{} consists of a world model as the high-level planner and a VLA as the low-level executor. The high-level world model first divides manipulation tasks into subtask sequences with goal images, and the low-level policy follows the textual and visual guidance to generate action sequences. Compared to raw textual goal specification, these synthesized goal images provide visually and physically grounded details for low-level policies, making it feasible to generalize across unseen objects and novel scenarios. We validate both visual goal synthesis and our hierarchical VLA policies in massive out-of-distribution scenarios, and the performance of the same-structured VLA in novel scenarios could boost from 14% to 69% with the guidance generated by the world model. Results demonstrate that our method outperforms previous baselines with a clear margin, particularly in out-of-distribution scenarios. Project page: \href{https://vista-wm.github.io/}{https://vista-wm.github.io}

cs.RO

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning capabilities required for complex, real-world problemsolving, particularly in multi-disciplinary and multimodal contexts. In this paper, we introduce a graduate-level, multi-disciplinary, EnglishChinese benchmark, dubbed as Reasoning Bench (R-Bench), for assessing the reasoning capability of both language and multimodal models. RBench spans 1,094 questions across 108 subjects for language model evaluation and 665 questions across 83 subjects for multimodal model testing in both English and Chinese. These questions are meticulously curated to ensure rigorous difficulty calibration, subject balance, and crosslinguistic alignment, enabling the assessment to be an Olympiad-level multi-disciplinary benchmark. We evaluate widely used models, including OpenAI o1, GPT-4o, DeepSeek-R1, etc. Experimental results indicate that advanced models perform poorly on complex reasoning, especially multimodal reasoning. Even the top-performing model OpenAI o1 achieves only 53.2% accuracy on our multimodal evaluation. Data and code are made publicly available at here.

cs.CV

AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset

Preference learning is critical for aligning large language models (LLMs) with human values, yet its success hinges on high-quality datasets comprising three core components: Preference \textbf{A}nnotations, \textbf{I}nstructions, and \textbf{R}esponse Pairs. Current approaches conflate these components, obscuring their individual impacts and hindering systematic optimization. In this work, we propose \textbf{AIR}, a component-wise analysis framework that systematically isolates and optimizes each component while evaluating their synergistic effects. Through rigorous experimentation, AIR reveals actionable principles: annotation simplicity (point-wise generative scoring), instruction inference stability (variance-based filtering across LLMs), and response pair quality (moderate margins + high absolute scores). When combined, these principles yield +5.3 average gains over baseline method, even with only 14k high-quality pairs. Our work shifts preference dataset design from ad hoc scaling to component-aware optimization, offering a blueprint for efficient, reproducible alignment.

cs.CL

Effect of phase change on shock wave and n-dodecane droplet interaction with numerical investigation

In a real propulsion system, shock-droplet interaction is often accompanied by phase change, which has a significant effect on the deformation and fragmentation of the droplet. In this paper, we study the effect of phase change on the n-dodecane droplet propulsion, deformation and fragmentation impacted by shock waves with high-resolution numerical simulations. First, we conduct a comparative study on shock waves and n-dodecane droplets interaction with and without phase change model. The impact of the shock wave changes the pressure and temperature distribution around the droplet, causing the droplet liquefaction on the windward side. With the influence of phase change, the Kelvin-Helmholtz instability (KHI) waves on the windward surface are enhanced, the development of KHI waves on the leeward surface of droplet is inhibited by vaporization. Furthermore, it is found that phase change suppresses both the flattening of the cylinder and shearing of the sheet at droplet equator. Next, we investigate the effect of Mach number on shock-droplet interaction with consideration of phase change. As the shock Mach number increases, the flattening and vaporization of droplets are suppressed, the KHI waves on the windward surface and the shear stripping of the sheet at the droplet equator are enhanced. The shear stripping of the liquid sheet plays a more dominant role in the deformation and breakup process than the flattening of the droplet under the SIE breakup mechanism in a higher Mach number.

physics.flu-dyn