Searcharxiv⌕ Search

arXiv subjects

Xiaojun Zhu

Publications and source records attributed to Xiaojun Zhu.

11 recordsLinked to original sources

Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model

Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for teleport-like motion. We trace the failure to the training signal, not the representation: anchor-sparse, latent-only supervision never says where along the horizon change belongs. Decode-augmented rollout training (DART) repairs this while keeping the representation frozen, retraining only the flow with decode-path supervision. DART outperforms its latent only parent on the full protocol, restores the temporal structure of motion, and re-couples predicted motion to the scene; at larger scale it further improves prediction quality, closing nearly half the remaining gap to an oracle-informed interpolation reference. Finally, we report an unexpected finding about evaluation: pixel error alone rewards frozen predictions.

cs.CV↗

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9\% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20$\times$; typed heads are also 6.68--6.90$\times$ faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a $π_{0.5}$ action policy, Pointing-VLA raises autonomous real-robot success from 52.7\% to 80.7\% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution.

cs.RO↗

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47\% to 3.00\%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies \emph{prompt-form collapse}: changing the instruction form, rather than adding useful semantics, can dominate execution. We introduce TOWN-VLA (Think Only When Needed), a prompt-authority interface that separates candidate generation from permission to alter the policy input. A fixed compatibility rule authorizes a canonical compact instruction; otherwise, the interface restores the original Base prompt exactly. Across 900 audited routes, every route follows this contract: 525 routes recover Base with matching hashes, and all 375 authorized prompts preserve the task signature. On a matched $4\times7$ LIBERO-Plus evaluation with 10{,}030 episodes per method, success rises from 69.5\% to 73.1\% ($+362$ episodes; 95\% CI 1.89--5.45 points), improving on six perturbation axes and all four suites. On a physical PiPER arm with a frozen \pizerofive{} checkpoint, success rises from 52.7\% to 78.7\% over 150 trials per method ($p=3.16\times10^{-6}$). Prompt authority is enforceable for a frozen controller; oracle-free admission calibration is the next deployment target.

cs.RO↗

Rethinking Demonstration Unlearning in Imitation Learning for Robotics

Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. Metrics inherited from machine unlearning, such as forgetting loss or a single membership attack, do not establish what an edit removed from a policy acting in closed loop. We therefore introduce a retrain-calibrated audit that reads demonstration unlearning along two axes: behavior, whether the edited policy acts like one retrained without the removed demonstrations, and evidence, whether an auditor can still detect it was trained on them. The behavior axis measures action divergence to that retrain at matched states, calibrated by a floor built from independent retrains, so a policy at the floor is as close to a retrain as retrains are to each other. The evidence axis applies a per-demonstration membership attack against a retrain null, reporting both its rank and its absolute member-loss level, since rank alone accepts operators that inflate member losses past the null. A conformal test then combines both axes into one hypothesis of joint retrain consistency, against a fleet of independent retrains large enough to reject at conventional significance. Across five preregistered conditions on three real-robot policy classes and two simulation suites, the axes dissociate in both directions on one checkpoint, as an edit may repair task behavior while leaving evidence unchanged, or reduce evidence while moving behavior away from retraining. On the ACT arm, a redirect edit restores blind-scored robot success to 18 of 20 trials.

cs.RO↗

$D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction

Predicting single-cell transcriptomic responses to genetic perturbations is central to functional genomics and virtual-cell modeling. Existing approaches, however, typically predict an entire expression profile as a whole, leaving the order in which individual gene responses are generated unmodeled. To address this problem, we introduce \textbf{$D^{2}R^{2}$} (\textbf{D}iscrete \textbf{D}iffusion with \textbf{R}egulation \textbf{R}einforcement), which reformulates perturbation prediction as regulation-guided gene-wise progressive generation. A Masked Discrete Diffusion Model represents expression as ordinal tokens and reconstructs a fully masked profile step by step, allowing generated gene responses to condition those that remain masked. A Regulatory Policy Module initializes the generation policy from a gene regulatory network inferred from control cells and adapts it to the perturbation and current partially generated state. Then, group-relative policy optimization refines only the ordering policy using final perturbation-effect agreement as reward. Across Norman19 and VCC-H1, $D^{2}R^{2}$ achieves the best performance on all five metrics on Norman19 and remains competitive on H1. Controlled ablations holding the generator and generation budget fixed show that biological-prior ordering improves over random ordering and is more reliable than uncertainty-based heuristics, whereas reversing the biological-prior ordering degrades every metric. Biological analyses further show that the refined policy prioritizes regulatory genes early while promoting perturbation-specific transcription factors and responsive genes. These results establish gene generation order as an effective, controllable, and biologically interpretable dimension of single-cell perturbation prediction.

cs.AI↗

A novel Phase I clinical trial design with unequal cohort sizes

This paper introduces a new Phase I design aimed at enhancing the performance of existing methods, including algorithm-based, model-based, and model-assisted designs. The design, developed by integrating the concept of Fisher information, is easily operationalized. The new design addresses the issue of the classical designs'slow dosage escalation. Simulation demonstrate that the proposed design markedly enhances performance in terms of efficiency, accuracy, and reliability. Moreover, the trial duration has been notably reduced with a large sample size.

stat.AP↗

Photonic Integrated Neuro-Synaptic Core for Convolutional Spiking Neural Network

Neuromorphic photonic computing has emerged as a competitive computing paradigm to overcome the bottlenecks of the von-Neumann architecture. Linear weighting and nonlinear spiking activation are two fundamental functions of a photonic spiking neural network (PSNN). However, they are separately implemented with different photonic materials and devices, hindering the large-scale integration of PSNN. Here, we propose, fabricate and experimentally demonstrate a photonic neuro-synaptic chip enabling the simultaneous implementation of linear weighting and nonlinear spiking activation based on a distributed feedback (DFB) laser with a saturable absorber (DFB-SA). A prototypical system is experimentally constructed to demonstrate the parallel weighted function and nonlinear spike activation. Furthermore, a four-channel DFB-SA array is fabricated for realizing matrix convolution of a spiking convolutional neural network, achieving a recognition accuracy of 87% for the MNIST dataset. The fabricated neuro-synaptic chip offers a fundamental building block to construct the large-scale integrated PSNN chip.

physics.optics↗

Hardware-algorithm collaborative computing with photonic spiking neuron chip based on integrated Fabry-Pérot laser with saturable absorber

Photonic neuromorphic computing has emerged as a promising avenue toward building a low-latency and energy-efficient non-von-Neuman computing system. Photonic spiking neural network (PSNN) exploits brain-like spatiotemporal processing to realize high-performance neuromorphic computing. However, the nonlinear computation of PSNN remains a significant challenging. Here, we proposed and fabricated a photonic spiking neuron chip based on an integrated Fabry-Pérot laser with a saturable absorber (FP-SA) for the first time. The nonlinear neuron-like dynamics including temporal integration, threshold and spike generation, refractory period, and cascadability were experimentally demonstrated, which offers an indispensable fundamental building block to construct the PSNN hardware. Furthermore, we proposed time-multiplexed spike encoding to realize functional PSNN far beyond the hardware integration scale limit. PSNNs with single/cascaded photonic spiking neurons were experimentally demonstrated to realize hardware-algorithm collaborative computing, showing capability in performing classification tasks with supervised learning algorithm, which paves the way for multi-layer PSNN for solving complex tasks.

cs.ET↗

Fluorescence temperature sensing based on thermally activated singlet-triplet intersystem crossing in crystalline anthracene

The temperature dependence of the steady-state fluorescence spectrum of anthracene crystals range from 300K to 500K had been investigated, which was in the temperature range of most tabletop laser-driven shock wave experiments. The interesting finding is that the fluorescence intensity of the 2-0 transition increases more rapidly than other transitions with the rising temperature. In particular, the transition intensity ratios γn all shows a perfect exponential increasing curve, which can be used for fluorescence temperature sensing. The analysis of sensitivity η and random uncertainty ΔT has demonstrated that the intensity ratio γ2 is the best comprehensive performance physical quantity for temperature sensing. The theoretical analysis and experimental results demonstrated that unusual intensity increasing of 2-0 transition was originated from the second excited triplet state T2, which was thermally coupled with the first excited singlet sate S1. In a word, we established a new fluorescence temperature sensing method based on the intensity ratio and clarified the mechanism of this method was the thermally activated singlet-triplet intersystem crossing.

physics.app-ph↗

Online Vector Scheduling and Generalized Load Balancing

We give a polynomial time reduction from vector scheduling problem (VS) to generalized load balancing problem (GLB). This reduction gives the first non-trivial online algorithm for VS where vectors come in an online fashion. The online algorithm is very simple in that each vector only needs to minimize the $L_{\ln(md)}$ norm of the resulting load when it comes, where $m$ is the number of partitions and $d$ is the dimension of vectors. It has an approximation bound of $e\log(md)$, which is in $O(\ln(md))$, so it also improves the $O(\ln^2d)$ bound of the existing polynomial time algorithm for VS. Additionally, the reduction shows that GLB does not have constant approximation algorithms that run in polynomial time unless $P=NP$.

cs.CC↗

Influence of baryons on spatial distribution of matter: higher order correlation functions

Baryonic physical processes could leave non-negligible imprint on cosmic matter distribution pattern. Series of high precision simulation data sets with identical initial condition are employed for count-in-cell (CIC) analysis, including one N-body dark matter run, one with adiabatic gas only and one with dissipative processes. Variances and higher order correlation functions of dark matter and gas are estimated. It is found that baryon physical processes mainly affected dark matter distribution at scales less than $1h^{-1}$Mpc. In comparison with the pure dark matter run, adiabatic process alone strengthens variance of dark matter by \sim 10% at scale $0.1h^{-1}$Mpc, while $S_n$s of dark matter deviate from pure dark matter case only mildly at a few percentages. Dissipative gas run does not differ much to the adiabatic run in dark matter variance, but renders significantly different $S_n$ parameters of dark matter, bringing about more than 10% enhancement to $S_3$ at $0.1h^{-1}$Mpc and $z=0$. Distribution patterns of gas in two hydrodynamical simulations are prominently different. Variance of gas at $z=0$ decreases by $\sim 30%$ in adiabatic simulation while by $\sim 60%$ in non-adiabatic simulation at $0.1h^{-1}$Mpc, the attenuation is weaker at larger scales but still obvious at $\sim 10h^{-1}$Mpc. $S_n$ parameters of gas are biased upward at scales $< \sim 4h^{-1}$Mpc, dissipative processes give $\sim 84%$ promotion at $z=0$ to $S_3$ at $0.1h^{-1}$Mpc against the moderate $\sim 7%$ in adiabatic simulation. The clustering segregation we observed between gas and dark matter could have intricate implication on modeling galaxy distribution and relevant cosmological application demanding fine details of matter distribution in strongly nonlinear regime.

astro-ph.CO↗