SearcharxivSearch

arXiv subjects

Lixin Xu

Publications and source records attributed to Lixin Xu.

At least 19 recordsLinked to original sources

How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

Multi-fingered dexterous manipulation remains a frontier for real-world reinforcement learning (RL) due to the high-dimensional action space and the prohibitive cost of hardware failures. While human-in-the-loop (HIL) RL allows operators to intervene before failures occur, current pipelines often treat these interventions as reactive corrections, discarding the rich safety signal inherent in the operator's decision to take control. In this paper, we ask: How can we learn from what a human would avoid? We present WHIRL, a safety-aware RL framework that transforms binary human interventions into forward-predictive signals for proactive risk avoidance. Our approach centers on an intervention-aware latent world model with four prediction heads: dynamics, reward, termination, and a novel per-state intervention-probability head that learns to predict the likelihood of a human takeover at future states. This head provides an actor-side risk-shaping term that discourages the policy from entering "intervention-prone" regions, modeling the operator's internal safety threshold. We evaluate our framework on a 16-DoF LEAP Hand across tasks spanning convex and irregular object grasping, prismatic manipulation, and long-horizon multi-stage tasks. Our results show that predictive risk-shaping enables the system to achieve a 96.7 percent success rate on complex grasping tasks while reducing the operator intervention burden by up to 84 percent in step-weighted terms. By closing the loop between human intuition and predictive world modeling, this work provides a practical safety-aware recipe for training complex dexterous agents in the real world while reducing operator fatigue and hardware-risk exposure.

cs.RO

Proto-Area Response Beyond Bulk Entropy

We study how much of the proto-area response can be inferred from the bulk von Neumann entropy. In the perturbative holographic code setting considered by CCKLP and Witten, we keep the recovery frame fixed and analyze the response averaged over the GUE at second order in the encoding perturbation. Near the maximally mixed bulk spectrum, the response is fixed by the entropy through quadratic order. For $d_1\geq3$, spectral information first enters at cubic order through the third centered moment, so states with the same entropy can have different responses. We determine the resulting response range on small equal-entropy level sets and obtain its sharp leading $D^{3/2}$ dependence, including the coefficient. The corresponding minimax error for prediction from entropy alone is one half of this range. For $d_1=2$, the entropy instead determines the response exactly.

hep-th

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical deployment faces a critical yet frequently overlooked hurdle: the pose misalignment dilemma. In real-world scenarios, the initial pose of a user-provided static mesh rarely aligns with the starting frame of a reference video. Naively forcing a mesh to follow a mismatched trajectory inevitably leads to severe geometric distortion or animation failure. To address this, we present Rectified Dynamic Mesh (R-DMesh), a unified framework designed to generate high-fidelity 4D meshes that are ``rectified'' to align with video context. Unlike standard motion transfer approaches, our method introduces a novel VAE that explicitly disentangles the input into a conditional base mesh, relative motion trajectories, and a crucial rectification jump offset. This offset is learned to automatically transform the arbitrary pose of the input mesh to match the video's initial state before animation begins. We process these components via a Triflow Attention mechanism, which leverages vertex-wise geometric features to modulate the three orthogonal flows, ensuring physical consistency and local rigidity during the rectification and animation process. For generation, we employ a Rectified Flow-based Diffusion Transformer conditioned on pre-trained video latents, effectively transferring rich spatio-temporal priors to the 3D domain. To support this task, we construct Video-RDMesh, a large-scale dataset of over 500k dynamic mesh sequences specifically curated to simulate pose misalignment. Extensive experiments demonstrate that R-DMesh not only solves the alignment problem but also enables robust downstream applications, including pose retargeting and holistic 4D generation.

cs.CV

Redshift-Dependent Intrinsic Dispersion in the Quasar UV/X-ray Luminosity Relation

Accurate modeling of the intrinsic dispersion in the quasar UV/X-ray luminosity relation is essential for reliable cosmological inference. We investigate its redshift dependence using luminosity distances reconstructed from cosmic chronometer and baryon acoustic oscillation measurements through Gaussian-process (GP) regression. Bayesian model comparison and posterior constraints show that the intrinsic dispersion is not well described by a single redshift-independent constant over $0.7<z<2.6$. It remains approximately constant at $0.7<z<1.6$, but shows an overall decreasing trend in the higher-redshift interval $1.6<z<2.6$, where the redshift-dependent intrinsic-dispersion model is decisively favored. This conclusion remains qualitatively robust against changes in the scaling-relation parameterization, GP kernel, and redshift binning scheme. We further examine its impact on cosmological inference in the flat $Λ$CDM model and find that, under the adopted calibration setup, the redshift-dependent intrinsic-dispersion model shifts the posterior median of $Ω_{\rm m0}$ by $ΔΩ_{\rm m0}\simeq 0.025$. This indicates that intrinsic-dispersion modeling is a non-negligible component of the systematic-error budget for quasar cosmology and should be accounted for in future precision analyses.

astro-ph.CO

Higher-Order Modes Are More Robust: The Origin of Bending Insensitivity in Multimode Fibers and Its Exploitation for Imaging

Multimode fibers (MMFs) hold great promise for minimally invasive imaging, yet their extreme sensitivity to bending perturbations severely hinders practical applications. Starting from mode coupling theory, we theoretically prove and experimentally verify that higher-order modes exhibit significantly greater bending robustness than lower-order modes. We reveal that bending-induced changes in the output speckle arise predominantly from alterations in the mode effective propagation constants, whereas the modal power redistribution caused by inter-modal coupling is negligible. Since the bending-induced changes are dominated by alterations in the mode effective propagation constants, which are inherently independent of the input field, and the modal power redistribution caused by inter-modal coupling is negligible, it follows that, under moderate bending, the perturbation imposed by bending on light propagation inside the fiber is independent of the input field distribution. Based on this input-independent bending perturbation property, we propose a physics-inspired dual-encoder neural network. The network separately extracts features from a reference speckle pattern (corresponding to a circular field input) and an information-carrying speckle pattern acquired under the same fiber state, then employs a differential fusion module to decouple the image information from environmental perturbations, and subsequently recovers the image after excluding the perturbation. Our method significantly enhances the anti-bending capability of MMF imaging systems, offering a new paradigm that integrates physical insights with deep learning for robust fiber-optic imaging.

physics.optics

Deep Learning Calibration of the Quasar X-ray/UV Luminosity Relation for Cosmological Applications

Quasars can serve as standard candles through an empirical scaling relation between their ultraviolet (UV) and X-ray luminosities. As high-redshift probes, it is critical to test whether this relation evolves with redshift. In this work, we reconstruct the Hubble diagram of the Pantheon+ sample using the deep learning--based LADDER algorithm and use it as a reference to investigate the quasar scaling relation. Our results, which are consistent with those from Gaussian process regression and narrow-bin analyses, show that the potentially contaminated sample at $z<0.7$ differs significantly from the $z>0.7$ sample; thus, it should be further screened or excluded when quasars are used as cosmological probes. We find that the scaling relation exhibits a non-linear redshift dependence that cannot be accounted for by a simple linear correction, and that this behavior is a feature of the current data sample rather than a consequence of cosmological model misspecification. To use quasars as standardizable candles, further modeling of the scaling relation and intrinsic dispersion, or more advanced data processing techniques, is required.

astro-ph.CO

Redshift Evolution of the HII Galaxy $L$-$σ$ Relation: Gaussian Process Analysis and Cosmological Implications

The empirical correlation between the H$β$ luminosity ($L$) and the ionized gas velocity dispersion ($σ$) in HII starburst galaxies (HIIGs) provides a foundation for using them as cosmological standard candles. A key unresolved issue is whether this $L$-$σ$ relation changes with redshift, which would impact its application at high redshifts. We test for possible evolution using cosmology-independent distance estimates up to $z \sim 1.8$, obtained from Gaussian Process regression of the Pantheon+ Type Ia supernovae Hubble diagram. These distances allow us to compare the standard $L$-$σ$ relation with three redshift-dependent extensions through Bayesian model comparison. We find that a logarithmic redshift correction is statistically preferred when the intrinsic dispersion of the relation is explicitly modeled, significantly improving the fit to high-$z$ data. However, the evidence for evolution strongly depends on how the likelihood function accounts for this intrinsic dispersion and is weaker if it is ignored. We also show that Malmquist bias significantly affects comparisons between low- and high-$z$ samples, reducing -- though not eliminating -- the statistical preference for redshift evolution after matching luminosity ranges. These results indicate that current HIIG data favor a redshift-dependent modification of the standard $L$-$σ$ relation, while highlighting the critical role of selection effects and intrinsic dispersion modeling in establishing HIIGs as precise cosmological probes.

astro-ph.CO

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

We present Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets. This system includes two foundation components: a large-scale shape generation model -- Hunyuan3D-DiT, and a large-scale texture synthesis model -- Hunyuan3D-Paint. The shape generative model, built on a scalable flow-based diffusion transformer, aims to create geometry that properly aligns with a given condition image, laying a solid foundation for downstream applications. The texture synthesis model, benefiting from strong geometric and diffusion priors, produces high-resolution and vibrant texture maps for either generated or hand-crafted meshes. Furthermore, we build Hunyuan3D-Studio -- a versatile, user-friendly production platform that simplifies the re-creation process of 3D assets. It allows both professional and amateur users to manipulate or even animate their meshes efficiently. We systematically evaluate our models, showing that Hunyuan3D 2.0 outperforms previous state-of-the-art models, including the open-source models and closed-source models in geometry details, condition alignment, texture quality, and etc. Hunyuan3D 2.0 is publicly released in order to fill the gaps in the open-source 3D community for large-scale foundation generative models. The code and pre-trained weights of our models are available at: https://github.com/Tencent/Hunyuan3D-2

cs.CV

Gravitational formulation of stress-tensor deformed field theories

Stress-tensor deformations suggest a geometric origin of emergent gravity but are typically non-local for $d>2$. We couple a seed QFT to Einstein gravity with deformation parameter $λ$ and evaluate the gravitational path integral at the metric saddle. Around a fixed reference background, the leading deformation is universal: a bilocal term quadratic in the stress tensor with kernel set by the graviton Green's function, plus a systematic higher-order expansion. Expressed on the saddle-point (deformed) metric, the flow becomes local. We then provide two constructive completions on deformed backgrounds--Palatini $f(R)$ gravity and an eigenvalue method for general Ricci-based theories--and apply them to scalar generalized Nambu-Goto and $\det T$ deformations (arbitrary $d$), two-dimensional multi-scalar ModMax and Born-Infeld models, and four-dimensional root-$T\bar T$ and $T\bar T$ flows of Maxwell theory yielding ModMax and Born-Infeld electrodynamics. In free field theory, an off-shell analysis further shows that the leading quantum correction generates an Einstein-Hilbert term with controlled higher-derivative terms.

hep-th

DexFormer: Cross-Embodied Dexterous Manipulation via History-Conditioned Transformer

Dexterous manipulation remains one of the most challenging problems in robotics, requiring coherent control of high-DoF hands and arms under complex, contact-rich dynamics. A major barrier is embodiment variability: different dexterous hands exhibit distinct kinematics and dynamics, forcing prior methods to train separate policies or rely on shared action spaces with per-embodiment decoder heads. We present DexFormer, an end-to-end, dynamics-aware cross-embodiment policy built on a modified transformer backbone that conditions on historical observations. By using temporal context to infer morphology and dynamics on the fly, DexFormer adapts to diverse hand configurations and produces embodiment-appropriate control actions. Trained over a variety of procedurally generated dexterous-hand assets, DexFormer acquires a generalizable manipulation prior and exhibits strong zero-shot transfer to Leap Hand, Allegro Hand, and Rapid Hand. Our results show that a single policy can generalize across heterogeneous hand embodiments, establishing a scalable foundation for cross-embodiment dexterous manipulation. Project website: https://davidlxu.github.io/DexFormer-web/.

cs.RO

PDF-HR: Pose Distance Fields for Humanoid Robots

Pose and motion priors play a crucial role in humanoid robotics. Although such priors have been widely studied in human motion recovery (HMR) domain with a range of models, their adoption for humanoid robots remains limited, largely due to the scarcity of high-quality humanoid motion data. In this work, we introduce Pose Distance Fields for Humanoid Robots (PDF-HR), a lightweight prior that represents the robot pose distribution as a continuous and differentiable manifold. Given an arbitrary pose, PDF-HR predicts its distance to a large corpus of retargeted robot poses, yielding a smooth measure of pose plausibility that is well suited for optimization and control. PDF-HR can be integrated as a reward shaping term, a regularizer, or a standalone plausibility scorer across diverse pipelines. We evaluate PDF-HR on various humanoid tasks, including single-trajectory motion tracking, general motion tracking, style-based motion mimicry, and general motion retargeting. Experiments show that this plug-and-play prior consistently and substantially strengthens strong baselines. Code and models will be released.

cs.RO

Equivariant Cohomology, BRST Quantization, and Analytic Localization: A Unified Framework

This paper provides a detailed exposition of the two main models for equivariant cohomology -- the Cartan and Weil models -- and their explicit isomorphism via the Kalkman (Mathai--Quillen) transformation. We then connect this framework to the BRST quantization of gauge theories, showing how the BRST complex can be identified with the Cartan model. Viewing both the Kalkman transformation and Witten's Morse-theoretic deformation as gauge-fixing procedures leads naturally to the \emph{equivariant Witten deformation}. This combined perspective yields a transparent analytic proof of the Atiyah--Bott--Berline--Vergne (ABBV) localization formula for integrals of equivariantly closed forms.The theory is richly illustrated with computations on $\mathbb{CP}^1$ and $\mathbb{CP}^n$, supplemented by explicit coordinate calculations.

hep-th

Collective dynamics in holographic fractonic solids

Fractonic phases of matter, a class of states in which collective excitations with constrained mobility exist, were originally discovered in the study of quantum error-correcting codes in solvable lattice spin models such as Haah's code and the X-cube model. Recently, they have also drawn the attention of the high-energy physics community due to the UV/IR mixing that arises when coarse-graining these lattice models. In this work, we consider a (3+1)-dimensional holographic model of fractonic solids and investigate the low-energy collective dynamics systematically. By computing the quasinormal modes of black holes, we obtain all the hydrodynamic excitations on the boundary, including two acoustic phonons, a longitudinal diffusive mode, and a subdiffusive collective mode with the dispersion $ω\sim-ik^4$. In addition, it is found that the latter remains gapless when translational symmetry is explicitly broken. These results suggest that the subdiffusive mode is inherently protected by the crystal-dipole symmetry in solids and is qualitatively unaffected by broken spacetime symmetries.

hep-th

DexSinGrasp: Learning a Unified Policy for Dexterous Object Singulation and Grasping in Densely Cluttered Environments

Grasping objects in cluttered environments remains a fundamental yet challenging problem in robotic manipulation. While prior works have explored learning-based synergies between pushing and grasping for two-fingered grippers, few have leveraged the high degrees of freedom (DoF) in dexterous hands to perform efficient singulation for grasping in cluttered settings. In this work, we introduce DexSinGrasp, a unified policy for dexterous object singulation and grasping. DexSinGrasp enables high-dexterity object singulation to facilitate grasping, significantly improving efficiency and effectiveness in cluttered environments. We incorporate clutter arrangement curriculum learning to enhance success rates and generalization across diverse clutter conditions, while policy distillation enables a deployable vision-based grasping strategy. To evaluate our approach, we introduce a set of cluttered grasping tasks with varying object arrangements and occlusion levels. Experimental results show that our method outperforms baselines in both efficiency and grasping success rate, particularly in dense clutter. Codes, appendix, and videos are available on our website https://nus-lins-lab.github.io/dexsingweb/.

cs.RO

Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization

We introduce Auto-Connect, a novel approach for automatic rigging that explicitly preserves skeletal connectivity through a connectivity-preserving tokenization scheme. Unlike previous methods that predict bone positions represented as two joints or first predict points before determining connectivity, our method employs special tokens to define endpoints for each joint's children and for each hierarchical layer, effectively automating connectivity relationships. This approach significantly enhances topological accuracy by integrating connectivity information directly into the prediction framework. To further guarantee high-quality topology, we implement a topology-aware reward function that quantifies topological correctness, which is then utilized in a post-training phase through reward-guided Direct Preference Optimization. Additionally, we incorporate implicit geodesic features for latent top-k bone selection, which substantially improves skinning quality. By leveraging geodesic distance information within the model's latent space, our approach intelligently determines the most influential bones for each vertex, effectively mitigating common skinning artifacts. This combination of connectivity-preserving tokenization, reward-guided fine-tuning, and geodesic-aware bone selection enables our model to consistently generate more anatomically plausible skeletal structures with superior deformation properties.

cs.CV

Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation

The creation of high-quality 3D assets, a cornerstone of modern game development, has long been characterized by labor-intensive and specialized workflows. This paper presents Hunyuan3D Studio, an end-to-end AI-powered content creation platform designed to revolutionize the game production pipeline by automating and streamlining the generation of game-ready 3D assets. At its core, Hunyuan3D Studio integrates a suite of advanced neural modules (such as Part-level 3D Generation, Polygon Generation, Semantic UV, etc.) into a cohesive and user-friendly system. This unified framework allows for the rapid transformation of a single concept image or textual description into a fully-realized, production-quality 3D model complete with optimized geometry and high-fidelity PBR textures. We demonstrate that assets generated by Hunyuan3D Studio are not only visually compelling but also adhere to the stringent technical requirements of contemporary game engines, significantly reducing iteration time and lowering the barrier to entry for 3D content creation. By providing a seamless bridge from creative intent to technical asset, Hunyuan3D Studio represents a significant leap forward for AI-assisted workflows in game development and interactive media.

cs.CV

Taming VR Teleoperation and Learning from Demonstration for Multi-Task Bimanual Table Service Manipulation

This technical report presents the champion solution of the Table Service Track in the ICRA 2025 What Bimanuals Can Do (WBCD) competition. We tackled a series of demanding tasks under strict requirements for speed, precision, and reliability: unfolding a tablecloth (deformable-object manipulation), placing a pizza into the container (pick-and-place), and opening and closing a food container with the lid. Our solution combines VR-based teleoperation and Learning from Demonstrations (LfD) to balance robustness and autonomy. Most subtasks were executed through high-fidelity remote teleoperation, while the pizza placement was handled by an ACT-based policy trained from 100 in-person teleoperated demonstrations with randomized initial configurations. By carefully integrating scoring rules, task characteristics, and current technical capabilities, our approach achieved both high efficiency and reliability, ultimately securing the first place in the competition.

cs.RO

DexFlow: A Unified Approach for Dexterous Hand Pose Retargeting and Interaction

Despite advances in hand-object interaction modeling, generating realistic dexterous manipulation data for robotic hands remains a challenge. Retargeting methods often suffer from low accuracy and fail to account for hand-object interactions, leading to artifacts like interpenetration. Generative methods, lacking human hand priors, produce limited and unnatural poses. We propose a data transformation pipeline that combines human hand and object data from multiple sources for high-precision retargeting. Our approach uses a differential loss constraint to ensure temporal consistency and generates contact maps to refine hand-object interactions. Experiments show our method significantly improves pose accuracy, naturalness, and diversity, providing a robust solution for hand-object interaction modeling.

cs.RO