SearcharxivSearch

arXiv subjects

John Zhang

Publications and source records attributed to John Zhang.

12 recordsLinked to original sources

Hybrid Feedback Sampling for Sample-Efficient Model Predictive Control

Thanks to its parallelizability and flexibility, sampling-based Model Predictive Control (MPC) has become widely popular for controlling real-world robotic systems. However, for high-dimensional and open-loop unstable dynamical systems, the required number of samples to improve the control sequence will grow exponentially with the horizon, leading to poor sample efficiency and numerical instability. This paper investigates the instability of shooting methods in sampling-based MPC and shows that the optimal sampling proposal distribution can be realized by sampling with an optimized feedback policy. We refer to this algorithm as Feedback Sampling MPC (FS-MPC). FS-MPC involves a hybrid sampling design which balances local and global search based on the system stability and the available computation budget. Our theoretical analysis shows that our hybrid sampling approach achieves faster convergence than standard MPPI and better optimality than standard feedback sampling. Empirically, in diverse contact-rich control tasks like humanoid loco-manipulation and dexterous manipulation, we show that FS-MPC successfully tackles dynamically unstable tasks where standard sample-based approaches struggle, and strictly outperforms feedback policies alone. Finally, we validate our method on humanoid robot locomotion and manipulation tasks in the real world.

cs.RO

Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination

High-performing human-human teams learn intelligent and efficient communication and coordination strategies to maximize their joint utility. These teams implicitly understand the different roles of heterogeneous team members and adapt their communication protocols accordingly. Multi-Agent Reinforcement Learning (MARL) has attempted to develop computational methods for synthesizing such joint coordination-communication strategies, but emulating heterogeneous communication patterns across agents with different state, action, and observation spaces has remained a challenge. Without properly modeling agent heterogeneity, as in prior MARL work that leverages homogeneous graph networks, communication becomes less helpful and can even deteriorate the team's performance. In the past, we proposed Heterogeneous Policy Networks (HetNet) to learn efficient and diverse communication models for coordinating cooperative heterogeneous teams. In this extended work, we extend Heterogeneous Policy Networks (HetNet) to support scaling heterogeneous robot teams. Building on heterogeneous graph-attention networks, we show that HetNet not only facilitates learning heterogeneous collaborative policies but also enables end-to-end training for learning highly efficient binarized messaging. Our empirical evaluation shows that HetNet sets a new state of the art in learning coordination and communication strategies for heterogeneous multi-agent teams by achieving an 5.84% to 707.65% performance improvement over the next-best baseline across multiple domains while simultaneously achieving a 200x reduction in the required communication bandwidth.

cs.RO

Temporally Decoupled Diffusion Planning for Autonomous Driving

Motion planning in dynamic urban environments requires balancing immediate safety with long-term goals. While diffusion models effectively capture multi-modal decision-making, existing approaches treat trajectories as monolithic entities, overlooking heterogeneous temporal dependencies where near-term plans are constrained by instantaneous dynamics and far-term plans by navigational goals. To address this, we propose Temporally Decoupled Diffusion Model (TDDM), which reformulates trajectory generation via a noise-as-mask paradigm. By partitioning trajectories into segments with independent noise levels, we implicitly treat high noise as information voids and weak noise as contextual cues. This compels the model to reconstruct corrupted near-term states by leveraging internal correlations with better-preserved temporal contexts. Architecturally, we introduce a Temporally Decoupled Adaptive Layer Normalization (TD-AdaLN) to inject segment-specific timesteps. During inference, our Asymmetric Temporal Classifier-Free Guidance utilizes weakly noised far-term priors to guide immediate path generation. Evaluations on the nuPlan benchmark show TDDM approaches or exceeds state-of-the-art baselines, particularly excelling in the challenging Test14-hard subset.

cs.RO

Computer-Using World Model

Agents operating in complex software environments benefit from reasoning about the consequences of their actions, as even a single incorrect user interface (UI) operation can derail long, artifact-preserving workflows. This challenge is particularly acute for computer-using scenarios, where real execution does not support counterfactual exploration, making large-scale trial-and-error learning and planning impractical despite the environment being fully digital and deterministic. We introduce the Computer-Using World Model (CUWM), a world model for desktop software that predicts the next UI state given the current state and a candidate action. CUWM adopts a two-stage factorization of UI dynamics: it first predicts a textual description of agent-relevant state changes, and then realizes these changes visually to synthesize the next screenshot. CUWM is trained on offline UI transitions collected from agents interacting with real Microsoft Office applications, and further refined with a lightweight reinforcement learning stage that aligns textual transition predictions with the structural requirements of computer-using environments. We evaluate CUWM via test-time action search, where a frozen agent uses the world model to simulate and compare candidate actions before execution. Across a range of Office tasks, world-model-guided test-time scaling improves decision quality and execution robustness.

cs.SE

Constraints on Reversing the Thermodynamic Arrow of Time from Black Hole Thermodynamics, Wormholes, and Time-Symmetric Quantum Mechanics

Can the thermodynamic arrow of time in a single universe be reversed, even temporarily, within semiclassical gravity without invoking additional universes or branches? We address this question in a single, connected spacetime where quantum field theory is coupled to classical general relativity, and where black holes, traversable wormholes, and time-symmetric or retrocausal formulations of quantum mechanics might naively appear to open channels for entropy export or cancellation. After distinguishing fine-grained, coarse-grained, and generalized gravitational entropy, and formulating a cosmological coarse-grained entropy, we treat black hole evaporation, wormholes constrained by quantum energy inequalities, and two-time boundary-value frameworks (including absorber-type and two-state-vector formalisms) within a common information-theoretic language. We then introduce a "Global Entropy Transport" (GET) framework and derive a sectoral inequality that bounds the net decrease of matter-plus-radiation entropy in terms of changes in horizon area and correlation (mutual-information) terms, assuming the generalized second law and modern focusing and energy conditions. Within this framework, black holes, wormholes, and retrocausal protocols can at most redistribute entropy among matter, radiation, and gravitational sectors and reshape the local pattern of entropy production. They do not, under current semiclassical, holographic, and statistical-mechanical constraints, permit a genuine reversal of the universal thermodynamic arrow in a single connected universe.

gr-qc

A robust generalizable device-agnostic deep learning model for sleep-wake determination from triaxial wrist accelerometry

Study Objectives: Wrist accelerometry is widely used for inferring sleep-wake state. Previous works demonstrated poor wake detection, without cross-device generalizability and validation in different age range and sleep disorders. We developed a robust deep learning model for to detect sleep-wakefulness from triaxial accelerometry and evaluated its validity across three devices and in a large adult population spanning a wide range of ages with and without sleep disorders. Methods: We collected wrist accelerometry simultaneous to polysomnography (PSG) in 453 adults undergoing clinical sleep testing at a tertiary care sleep laboratory, using three devices. We extracted features in 30-second epochs and trained a 3-class model to detect wake, sleep, and sleep with arousals, which was then collapsed into wake vs. sleep using a decision tree. To enhance wake detection, the model was specifically trained on randomly selected subjects with low sleep efficiency and/or high arousal index from one device recording and then tested on the remaining recordings. Results: The model showed high performance with F1 Score of 0.86, sensitivity (sleep) of 0.87, and specificity (wakefulness) of 0.78, and significant and moderate correlation to PSG in predicting total sleep time (R=0.69) and sleep efficiency (R=0.63). Model performance was robust to the presence of sleep disorders, including sleep apnea and periodic limb movements in sleep, and was consistent across all three models of accelerometer. Conclusions: We present a deep model to detect sleep-wakefulness from actigraphy in adults with relative robustness to the presence of sleep disorders and generalizability across diverse commonly used wrist accelerometers.

q-bio.QM

Mirror Dark Sector Solution of the Hubble Tension with Time-varying Fine-structure Constant

We explore a model introduced by Cyr-Racine, Ge, and Knox (arXiv:2107.13000(2)) that resolves the Hubble tension by invoking a ``mirror world" dark sector with energy density a fixed fraction of the ``ordinary" sector of Lambda-CDM. Although it reconciles cosmic microwave background and large-scale structure observations with local measurements of the Hubble constant, the model requires a value of the primordial Helium mass fraction that is discrepant with observations and with the predictions of Big Bang Nucleosynthesis (BBN). We consider a variant of the model with standard Helium mass fraction but with the value of the electromagnetic fine-structure constant slightly different during photon decoupling from its present value. If $α$ at that epoch is lower than its current value by $Δα\simeq -2\times 10^{-5}$, then we can achieve the same Hubble tension resolution as in Cyr-Racine, et al. but with consistent Helium abundance. As an example of such time-evolution, we consider a toy model of an ultra-light scalar field, with mass $m <4\times 10^{-29}$ eV, coupled to electromagnetism, which evolves after photon decoupling and that appears to be consistent with late-time constraints on $α$ variation and the weak equivalence principle.

astro-ph.CO

Fast Contact-Implicit Model-Predictive Control

We present a general approach for controlling robotic systems that make and break contact with their environments. Contact-implicit model predictive control (CI-MPC) generalizes linear MPC to contact-rich settings by utilizing a bi-level planning formulation with lower-level contact dynamics formulated as time-varying linear complementarity problems (LCPs) computed using strategic Taylor approximations about a reference trajectory. These dynamics enable the upper-level planning problem to reason about contact timing and forces, and generate entirely new contact-mode sequences online. To achieve reliable and fast numerical convergence, we devise a structure-exploiting interior-point solver for these LCP contact dynamics and a custom trajectory optimizer for the tracking problem. We demonstrate real-time solution rates for CI-MPC and the ability to generate and track non-periodic behaviours in hardware experiments on a quadrupedal robot. We also show that the controller is robust to model mismatch and can respond to disturbances by discovering and exploiting new contact modes across a variety of robotic systems in simulation, including a pushbot, planar hopper, planar quadruped, and planar biped.

cs.RO

A Knowledge-Based Decision Support System for In Vitro Fertilization Treatment

In Vitro Fertilization (IVF) is the most widely used Assisted Reproductive Technology (ART). IVF usually involves controlled ovarian stimulation, oocyte retrieval, fertilization in the laboratory with subsequent embryo transfer. The first two steps correspond with follicular phase of females and ovulation in their menstrual cycle. Therefore, we refer to it as the treatment cycle in our paper. The treatment cycle is crucial because the stimulation medications in IVF treatment are applied directly on patients. In order to optimize the stimulation effects and lower the side effects of the stimulation medications, prompt treatment adjustments are in need. In addition, the quality and quantity of the retrieved oocytes have a significant effect on the outcome of the following procedures. To improve the IVF success rate, we propose a knowledge-based decision support system that can provide medical advice on the treatment protocol and medication adjustment for each patient visit during IVF treatment cycle. Our system is efficient in data processing and light-weighted which can be easily embedded into electronic medical record systems. Moreover, an oocyte retrieval oriented evaluation demonstrates that our system performs well in terms of accuracy of advice for the protocols and medications.

cs.AI

Affine lines over derivators: Properties

In this paper we detail a number of properties of the affine line of a derivator, including a number of morphisms between $\mathbb{D}$ and $\mathbb{A}^1_{\mathbb{D}}$, a monoidal structure on $\mathbb{A}^1_{\mathbb{D}}$ if $\mathbb{D}$ is monoidal, and a universal property of $\mathbb{A}^1_{\mathbb{D}}$ over $\mathbb{D}$ akin to the universal property that $R[t]$ has over $R$. These results also extend naturally to affine space $\mathbb{A}^n$ over a derivator $\mathbb{D}$.

math.CT