SearcharxivSearch

arXiv subjects

Runqing Wang

Publications and source records attributed to Runqing Wang.

15 recordsLinked to original sources

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, without coupling them into a closed decision-and-learning loop for policy improvement. We present Motus2, a self-evolving general world model for dexterous manipulation. Motus2 advances world modeling through model scaling and data scaling. For model scaling, a single model with shared weights exposes three control interfaces: a policy (world-action model), a simulator (action-conditioned world model), and an evaluator (value model). The policy proposes candidate action chunks, the simulator predicts their visual consequences, and the evaluator assesses the predicted outcomes. Their coupling forms a closed decision-and-learning loop for policy improvement. This formulation uses curated expert demonstrations for action learning, while failed and suboptimal interactions provide valuable evidence for dynamics modeling and value learning. For data scaling, Motus2 progresses from large-scale monocular egocentric data to synchronized stereo egocentric data, followed by robot-domain adaptation with robot trajectories and supplementary human-robot alignment data. Motus2 further studies global-autoregressive and hybrid-memory extensions of its sliding-window context, adds tactile feedback for contact-aware control, and is instantiated on a fully biomimetic platform with stereo vision, dual arms, dual dexterous hands, and tactile sensing. Together, egocentric data scaling and closed-loop general world model scaling provide a general path toward self-evolving dexterous manipulation.

cs.RO

FlowPilot: Real-Time World-Action Modeling for Agile UAV Navigation

We present FlowPilot, a compact world-action model for real-time onboard UAV navigation from depth. Unlike map-then-optimize pipelines that require local reconstruction or end-to-end policies that lack explicit scene prediction, FlowPilot jointly denoises future depth observations and executable trajectories with flow matching. A dual-stream mixture-of-transformers couples video and action experts through shared attention, allowing future-scene prediction and trajectory generation to inform each other. At deployment, the model runs action-centrically and outputs only a trajectory. To ensure trackability, actions are parameterized as degree-7 Bernstein polynomials: the current state constrains the initial control points, and the network predicts five free control points, yielding C^2-continuous references with closed-form velocity, acceleration and jerk. FlowPilot is trained on a three-level depth pyramid spanning high-throughput simulation, photorealistic simulation, and real onboard data. In closed-loop simulation, it outperforms learning- and optimization-based baselines under increasing clutter and commanded speeds up to 8m/s. On a physical quadrotor, the full perception-to-action pipeline runs in under 18ms on a Jetson Orin NX and reaches 5.5m/s in cluttered indoor and forest environments using only onboard sensing and computation.

cs.RO

AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight

Language-conditioned quadrotor flight requires a policy to ground semantic goals, anticipate the visual consequences of ego-motion, and output control references that remain smooth and dynamically executable under rapidly changing first-person views. Existing aerial vision-language navigation and vision-language-action methods commonly use discrete actions, high-level waypoints, or instantaneous velocity commands, which provide limited supervision about how flight actions change future observations. We present AeroAct, an action-centered world-action model (WAM) for quadrotor navigation. To the best of our knowledge, AeroAct is the first WAM instantiated and demonstrated for real-world aerial flight. The model adapts a pretrained video diffusion Transformer to predict local trajectory-action chunks from egocentric visual history, proprioception, and language. Future first-person frames are used during training as dense consequence supervision, while deployment directly decodes actions without generating future video. To obtain aligned visual, state, language, and dynamically feasible action data, we build a DiffAero-based pipeline with complementary Isaac Lab and 3D Gaussian splatting renderers. We further introduce a low-cost handheld collection device that couples camera observations with motion estimates to recreate flight-like egocentric trajectories, and a self-guidance procedure that improves temporal consistency across overlapping trajectory chunks. Closed-loop simulation and real-world experiments show that temporal visual context improves target tracking and object-search performance, and that WAM-based policies can be executed on a physical quadrotor.

cs.RO

MAD: Mapping-Aware World Models for Agile Quadrotor Flight

Agile quadrotor flight in cluttered scenes requires more than a reactive mapping from a depth image to a control command: the vehicle must remember which regions have been observed, infer nearby occupied space, and act under partial visibility and tight latency. In this paper, we present Mapping-Aware Dreamer (MAD), a geometry-aware world model for vision-based quadrotor flight. Instead of using raw-image reconstruction as the main self-supervised objective, MAD learns recurrent latent dynamics that reconstruct robocentric occupancy and visibility grid maps together with proprioceptive states. This design forces the latent state to encode local geometry, visibility history, and ego-motion in a form that is directly relevant to collision avoidance. MAD is trained in DiffAero using a GPU-parallel map-construction module that provides high-throughput supervision for occupancy and visibility. The learned representation is used in three policy-learning modes: imagination-based MAD-Dreamer and feature-extractor variants based on PPO and SHAC. Across visual navigation and racing tasks, MAD-based agents achieve higher success rates, faster flight, and better cross-task transfer than corresponding vision-only baselines. The model also produces interpretable map predictions and accurate ego-motion estimates from depth observations. We further deploy the learned policy on a physical quadrotor with an Intel RealSense D435i and demonstrate safe indoor and outdoor flight under limited sensing, reaching 9.66 m/s in simulation and 5.05 m/s in real-world forest experiments. These results show that mapping-aware world models provide a practical middle ground between modular aerial navigation and end-to-end learning.

cs.RO

Motubrain: An Advanced World Action Model for Robot Control

Vision-Language-Action (VLA) models generalize semantically well but often lack fine-grained modeling of world dynamics. We present Motubrain, a unified World Action Model that jointly models video and action under a UniDiffuser formulation with a three-stream Mixture-of-Transformers architecture. A single model supports policy learning, world modeling, video generation, inverse dynamics, and joint video-action prediction, while scaling to heterogeneous multimodal data such as video-only, task-agnostic, and cross-embodiment robot data. Building on Motus, Motubrain further introduces unified multiview modeling, an independent text stream for stronger language-action coupling, a shared cross-embodiment action representation, and an efficient post-training and deployment recipe for long-horizon real-world control. Our inference stack combines step reduction, compilation, FP8 quantization, DiT caching, V2A-style action-only inference, and real-time chunked closed-loop execution, achieving over 50x speedup over a naive baseline and up to 11 Hz inference. Experimentally, Motubrain achieves 95.8% and 96.1% average success on RoboTwin 2.0 under clean and randomized settings, respectively, attains the strongest reported EWMScore in our WorldArena comparison, and adapts to new humanoid embodiments with only 50--100 trajectories. These results show that unified world action models can scale in generality, predictive accuracy, and real-world deployability.

cs.RO

Geometric Amplitudes: A Covariant Functional Approach for Massless Scalar Theories

Functional geometry is a framework using concepts from geometry to understand the invariance of amplitudes in quantum field theory under a large class of field redefinitions, including those involving derivatives. It is inspired by recursion relations among correlation functions, where higher-point functions depend iteratively upon smaller correlators. Previous work has shown that, with suitable modifications, these correlation functions become covariant under field redefinitions, provided they are evaluated at the physical ``on-shell" point. In this paper, we show how to further modify correlation functions in massless scalar field theories to achieve ``off-shell" covariance. We investigate the conditions required for the framework to work and discuss the geometric interpretation of this construction -- which prioritizes the covariant transformation of observables under field redefinitions over the role of a metric tensor and its derivatives. While analogous modifications may exist for massive theories, we show that framework developed here does not extend straightforwardly to that case.

hep-th

DyMoDreamer: World Modeling with Dynamic Modulation

A critical bottleneck in deep reinforcement learning (DRL) is sample inefficiency, as training high-performance agents often demands extensive environmental interactions. Model-based reinforcement learning (MBRL) mitigates this by building world models that simulate environmental dynamics and generate synthetic experience, improving sample efficiency. However, conventional world models process observations holistically, failing to decouple dynamic objects and temporal features from static backgrounds. This approach is computationally inefficient, especially for visual tasks where dynamic objects significantly influence rewards and decision-making performance. To address this, we introduce DyMoDreamer, a novel MBRL algorithm that incorporates a dynamic modulation mechanism to improve the extraction of dynamic features and enrich the temporal information. DyMoDreamer employs differential observations derived from a novel inter-frame differencing mask, explicitly encoding object-level motion cues and temporal dynamics. Dynamic modulation is modeled as stochastic categorical distributions and integrated into a recurrent state-space model (RSSM), enhancing the model's focus on reward-relevant dynamics. Experiments demonstrate that DyMoDreamer sets a new state-of-the-art on the Atari $100$k benchmark with a $156.6$\% mean human-normalized score, establishes a new record of $832$ on the DeepMind Visual Control Suite, and gains a $9.5$\% performance improvement after $1$M steps on the Crafter benchmark. Our code is released at https://github.com/Ultraman-Tiga1/DyMoDreamer.

cs.LG

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

This letter introduces DiffAero, a lightweight, GPU-accelerated, and fully differentiable simulation framework designed for efficient quadrotor control policy learning. DiffAero supports both environment-level and agent-level parallelism and integrates multiple dynamics models, customizable sensor stacks (IMU, depth camera, and LiDAR), and diverse flight tasks within a unified, GPU-native training interface. By fully parallelizing both physics and rendering on the GPU, DiffAero eliminates CPU-GPU data transfer bottlenecks and delivers orders-of-magnitude improvements in simulation throughput. In contrast to existing simulators, DiffAero not only provides high-performance simulation but also serves as a research platform for exploring differentiable and hybrid learning algorithms. Extensive benchmarks and real-world flight experiments demonstrate that DiffAero and hybrid learning algorithms combined can learn robust flight policies in hours on consumer-grade hardware. The code is available at https://github.com/flyingbitac/diffaero.

cs.RO

Boiling After the Dust Settles: Constraining First-Order Phase Transitions During Dark Energy Domination

A first-order phase transition could occur in the late universe when vacuum energy begins dominating the energy density ($z \lesssim 0.3$) and convert some latent heat into other forms such as invisible radiation. This generic possibility also has concrete motivation in particle physics models which invoke a multitude of vacua to address theoretical puzzles. The na\"{i}ve constraint on such an event comes from measurements of the Hubble expansion rate, but this can only probe transitions involving $\mathcal{O}(10)\%$ of the dark energy. In this work, we show that significantly tighter constraints appear when accounting for phase transition fluctuations affecting CMB photon propagation anisotropically, akin to the integrated Sachs-Wolfe effect. For instance, if a completed phase transition has $\beta/H_\star\lesssim 25$, current CMB data limits the associated vacuum energy released to less than $1\%$ of the dark energy. A transition to negative vacuum energy (quasi-anti-de Sitter) is allowed only for $\beta/H_\star \gtrsim 300$. For $\beta/H_\star \lesssim 500$, the universe will not crunch for at least $14$ Gyr.

hep-ph

Non-factorizable Superamplitudes for Massive N = 1 Superstates

In this paper we study non-factorizable N = 1 superamplitudes for massive chiral superstates. We demonstrate how little group scaling and the supersymmetric Ward identities determine the form of non-factorizable massless superamplitudes, then extrapolate to massive superamplitudes by requiring they reduce to the massless form when we send all masses to zero. This technique does not depend on whether or not the superstates are self-conjugate (so that the fermionic components are either Dirac or Majorana) or whether the superamplitude is dressed with a form-factor.

hep-th

Basis for Non-Factorizable Superamplitudes in N = 1 Supersymmetry

In this paper we develop a semi-standard Young tableau (SSYT) approach to construct a basis of non-factorizable superamplitudes in N = 1 massless supersymmetry. This amplitude basis can be directly translated to a basis for higher dimensional supersymmetric operators, yielding both the number of independent operators and their form. We deal with distinguishable (massless) chiral/vector superfields at first, then generalize the result to the indistinguishable case. Finally, we discuss the advantages and disadvantages of this method compared to the previously studied Hilbert series approach.

hep-th

Hidden U(N) Symmetry behind N=1 Superamplitudes

In this paper we develop a Young diagram approach to constructing higher dimensional operators formed from massless superfields and their superderivatives in $\mathcal{N}=1$ supersymmetry. These operators are in one-to-one correspondence with non-factorizable terms in on-shell superamplitudes, which can be studied with massless spinor helicity techniques. By relating all spin-helicity variables to certain representations under a hidden $U(N)$ symmetry behind the theory, we show each non-factorizable superamplitude can be identified with a specific Young tableau. The desired tableau is picked out of a more general set of $U(N)$ tensor products by enforcing the supersymmetric Ward identities. We then relate these Young tableaux to higher dimensional superfield operators and list the rules to read operators directly from Young tableau. Using this method, we present several illustrative examples.

hep-th

Flexible Job Shop Scheduling via Dual Attention Network Based Reinforcement Learning

Flexible manufacturing has given rise to complex scheduling problems such as the flexible job shop scheduling problem (FJSP). In FJSP, operations can be processed on multiple machines, leading to intricate relationships between operations and machines. Recent works have employed deep reinforcement learning (DRL) to learn priority dispatching rules (PDRs) for solving FJSP. However, the quality of solutions still has room for improvement relative to that by the exact methods such as OR-Tools. To address this issue, this paper presents a novel end-to-end learning framework that weds the merits of self-attention models for deep feature extraction and DRL for scalable decision-making. The complex relationships between operations and machines are represented precisely and concisely, for which a dual-attention network (DAN) comprising several interconnected operation message attention blocks and machine message attention blocks is proposed. The DAN exploits the complicated relationships to construct production-adaptive operation and machine features to support high-quality decisionmaking. Experimental results using synthetic data as well as public benchmarks corroborate that the proposed approach outperforms both traditional PDRs and the state-of-the-art DRL method. Moreover, it achieves results comparable to exact methods in certain cases and demonstrates favorable generalization ability to large-scale and real-world unseen FJSP tasks.

cs.LG

Counting Operators in $N=1$ Supersymmetric Gauge Theories

Following a recent publication, in this paper we count the number of independent operators at arbitrary mass dimension in $N=1$ supersymmetric gauge theories and derive their field and derivative content. This work uses Hilbert series machinery and extends a technique from our previous work on handling integration by parts redundancies to vector superfields. The method proposed here can be applied to both abelian and non-abelian gauge theories and for any set of (chiral/antichiral) matter fields. We work through detailed steps for the abelian case with single flavor chiral superfield at mass dimension eight, and provide other examples in the appendices.

hep-th

Constructing Operator Basis in Supersymmetry: A Hilbert Series Approach

In this paper we introduce a Hilbert series approach to build the operator basis for a N = 1 supersymmetry theory with chiral superfields. We give explicitly the form of the corrections that remove redundancies due to the equations of motion and integration by parts. In addition, we derive the maps between the correction spaces. This technique allows us to calculate the number of independent operators involving chiral and antichiral superfields to arbitrarily high mass dimension. Using this method, we give several illustrative examples.

hep-th