SearcharxivSearch

arXiv subjects

Gu

Publications and source records attributed to Gu.

2 recordsLinked to original sources

X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling

Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive world modeling enables VLA to internalize physical dynamics and long-term causality by predicting future video from past observations. However, naive next-frame prediction faces two challenges: 1) unlike semantically distinct text tokens, video tokens are low-entropy and redundant, causing prediction to degenerate into trivial extrapolation. 2) world modeling poses a temporal dilemma: dense prediction captures instantaneous dynamics, but cannot efficiently model long-horizon causality. To learn world knowledge effectively, we introduce X-Foresight, a predictive world model integrated directly into the VLA architecture to jointly learn world modeling and real-time action control. At its core lies a long-horizon chunk-wise auto-regressive strategy that addresses both challenges: by predicting semantically distant chunks rather than adjacent frames, it escapes trivial extrapolation, while preserving dense intra-chunk frames for instantaneous dynamics and sparse inter-chunk transitions for long-term causality. A curriculum learning schedule progressively extends prediction horizons and stabilizes long-horizon training. To capture long-term causality effectively, we present temporal importance sampling, which concentrates supervision on safety-critical chunks identified by ego-motion and behavioral signals. We further delegate photorealistic synthesis to a diffusion-based multi-view renderer, improving photorealistic appearance. Comprehensive experiments demonstrate that X-Foresight significantly outperforms VLA baselines in planning performance while maintaining strong generative fidelity, establishing a robust paradigm for world-knowledge-driven autonomous systems.

cs.CV

Photoacoustic Imaging Based on AlN MF-PMUT with Broadened Bandwidth

This paper reports an aluminum nitride (AlN) multi-frequency piezoelectric micromachined ultrasound transducers (MF-PMUT) array for photoacoustic (PA) imaging, where the broadened bandwidth is beneficial to improve imaging resolution. Specifically, PMUT based on micro-electromechanical systems (MEMS) technology is suitable for PA endoscopic imaging of blood vessels and bronchi due to its miniature size. More importantly, AlN is a non-toxic material, which makes it harmless for biomedical applications. In this work, a MF-PMUT array are designed and fabricated for PAI. The device's vibration mode impedance and bandwidth are analyzed. The MF-PMUT sensor provides a wider bandwidth (65%) signal detection, which increases the resolution of PAI compared with traditional PMUT. We conduct an experiment on agar sample to present sensor's performance in images' axial resolution.

physics.med-ph