SearcharxivSearch

arXiv subjects

Yingjie Cai

Publications and source records attributed to Yingjie Cai.

At least 19 recordsLinked to original sources

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and controllable synthesis of dynamic urban scenes from a single frame. In practice, our approach first reconstructs a 3D occupancy representation from the input multi-view frame. This representation serves as a foundation for autoregressive scene extension along arbitrary trajectories. Subsequently, a video diffusion model translates the coarse occupancy grid into realistic, spatiotemporally consistent video sequences. Moreover, we propose a hierarchical sketch-and-refine paradigm, in which the generated videos are re-projected as image-conditioned feedback to enhance the 3D occupancy representation, establishing cross-modal alignment and mutual enhancement between the visual and spatial domains. Extensive evaluations on the Waymo Open Dataset and nuScenes demonstrate that InfiniVerse achieves state-of-the-art performance, with a FID of 6.4 and FVD of 67.97, significantly outperforming existing benchmarks in both duration and stability.

cs.CV

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs are typically trained in a variational autoencoder (VAE) latent space. However, the VAE latent space is optimized for pixel reconstruction, which rewards fine appearance detail and leaves the action prediction fragile under visual shifts. Recent works build WAMs in semantic latent space, which are more robust to appearance shifts. However, these models cannot leverage the large-scale VGM pretraining that exists only in VAE space. To overcome this dilemma, we propose Robust-WAM, a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream. This retains the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics that stay reliable under illumination shifts and other visual out-of-distribution conditions. Specifically, we employ learnable query tokens to bring future-scene semantics into the action stream by aligning their output hidden states with the semantic foresight of future ground-truth frames. To establish the temporal correspondence between each query and the future step it describes, we give it the positional encoding of the matching action tokens. Experiments on out-of-distribution generalization simulation benchmarks and a real-robot setup show that our Robust-WAM consistently improves the success rates of multiple WAM baselines without sacrificing in-distribution performance.

cs.CV

Revisiting Approaches to Stellar White-Light Flare Energy Based on Spatiotemporally Resolved Solar Observations

Accurately estimating the bolometric energy of solar and stellar white-light flares (WLFs) is crucial for understanding their physical nature and impact on surrounding planets. However, the lack of spatial resolution in stellar observations forced pioneering stellar WLF studies to adopt simplified energy estimation methods, typically assuming either a constant flare temperature or a fixed radiating area. To assess the physical plausibility of these assumptions, we utilize high-spatiotemporal-resolution solar observations to analyze the true evolution of source region's radiating area and temperature of 70 solar WLFs. It is revealed that both area and temperature of most solar WLFs undergo significant temporal evolution, and the flare area strongly correlates with the flare's peak optical continuum flux. Therefore, we propose a new energy estimation method that permits both flare area and temperature to evolve. Compared with existing methods, our dynamic approach yields systematically lower flare energies, which then prompts us to revisit classical macroscopic scaling laws related to the flare energy. It is further revealed that different energy estimation approaches can systematically alter these scaling relations, calling for a re-examination of these established statistical results and their targeted testing or revision in future work.

astro-ph.SR

Can We Distinguish the Source Region Location of Filament/Prominence Eruptions from the Sun-as-a-star H$α$ Spectrum?

Solar filament/prominence eruptions can significantly perturb geospace when originating from favorable source locations and directions. While stellar analogs have been recently reported, the disk locations and magnetic environments of their source regions remain spatially unresolved on other stars. To bridge this gap, we investigate the typical Sun-as-a-star H$α$ temporal spectral characteristics of solar filament/prominence eruptions with different source region locations (on-disk vs. limb, active region vs. quiet-Sun region). It is revealed that limb eruptions are characterized by blueshifted/redshifted emission caused by the bright off-limb erupting structures, whereas on-disk eruptions may show blueshifted absorptions due to the dark erupting filaments. Among the limb eruptions, front-side limb eruptions usually display line center emission before the blueshifted/redshifted emission, while far-side limb eruptions show the opposite sequence. Moreover, the magnetic environment at source also shapes the spectral characteristics. On-disk filament eruptions from active region exhibit much more intense flare-ribbon-dominated line center emission features compared with those from quiet-Sun region. Limb active region eruptions often show single-wing emissions, whereas large-scale quiet-Sun region (quiescent) prominence eruptions frequently display expansion-induced emission in both wings followed by line center absorption due to the disappearance of bright prominence. These distinct Sun-as-a-star H$α$ spectral characteristics, dependent on eruption location, provide a diagnostic basis for inferring source regions of stellar filament/prominence eruptions from spatially unresolved H$α$ spectra.

astro-ph.SR

Magnetic Evolution of Highly-Sheared Region in Active Region 13842 Producing Large X9.0 Flare

Shearing motion and magnetic flux cancellation around the polarity inversion line (PIL) play significant roles in the build-up of free magnetic energy and magnetic flux rope (MFR) in source region of major solar flares. Here we investigate the magnetic evolution of a highly-sheared PIL in active region (AR) 13842, hosting the largest X9.0 flare of Solar Cycle 25. Since 2024 September 29, a positive-polarity pore persistently drifted northward along the western side of the AR's main negative-polarity sunspot. The main sunspot remained stationary until negative-polarity patches successively emerged to its east and approached. Rear-ended by these same-polarity patches, the sunspot then began moving westward toward the opposite-polarity pore around October 1, forming a collisional PIL. Meanwhile, on the PIL's other side, the pore was also rear-ended by same-polarity patches sequentially emerging behind it, accelerating the shearing motion around the PIL, where frequent flux cancellations were also observed. Synchronous rapid accumulation of free magnetic energy and formation of MFR were then observed in the PIL, where multiple major flares successively occurred within two days. Before these large flares, the area and total free energy of the high-free-energy-density PIL region gradually decreased in the photosphere, which could be caused by the initial ascent of MFR before eruption and serve as a precursor of solar eruptions. These results suggest that persistent flux emergences with cross separation directions facilitates rapid formation of collisional shearing PIL and frequent flux cancellations, leading to repeated MFR formations and multiple large flares in a relatively short time.

astro-ph.SR

DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving

Vision-based autonomous driving has gained much attention due to its low costs and excellent performance. Compared with dense BEV (Bird's Eye View) or sparse query models, Gaussian-centric method is a comprehensive yet sparse representation by describing scene with 3D semantic Gaussians. In this paper, we introduce DLWM, a novel paradigm with Dual Latent World Models specifically designed to enable holistic gaussian-centric pre-training in autonomous driving using two stages. In the first stage, DLWM predicts 3D Gaussians from queries by self-supervised reconstructing multi-view semantic and depth images. Equipped with fine-grained contextual features, in the second stage, two latent world models are trained separately for temporal feature learning, including Gaussian-flow-guided latent prediction for downstream occupancy perception and forecasting tasks, and ego-planning-guided latent prediction for motion planning. Extensive experiments in SurroundOcc and nuScenes benchmarks demonstrate that DLWM shows significant performance gains across Gaussian-centric 3D occupancy perception, 4D occupancy forecasting and motion planning tasks.

cs.CV

DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models

Vision-Language-Action (VLA) models map visual observations and language instructions directly to robotic actions. While effective for simple tasks, standard VLA models often struggle with complex, multi-step tasks requiring logical planning, as well as precise manipulations demanding fine-grained spatial perception. Recent efforts have incorporated Chain-of-Thought (CoT) reasoning to endow VLA models with a ``thinking before acting'' capability. However, current CoT-based VLA models face two critical limitations: 1) an inability to simultaneously capture low-level visual details and high-level logical planning due to their reliance on isolated, single-modal CoT; 2) high inference latency with compounding errors caused by step-by-step autoregressive decoding. To address these limitations, we propose DualCoT-VLA, a visual-linguistic CoT method for VLA models with a parallel reasoning mechanism. To achieve comprehensive multi-modal reasoning, our method integrates a visual CoT for low-level spatial understanding and a linguistic CoT for high-level task planning. Furthermore, to overcome the latency bottleneck, we introduce a parallel CoT mechanism that incorporates two sets of learnable query tokens, shifting autoregressive reasoning to single-step forward reasoning. Extensive experiments demonstrate that our DualCoT-VLA achieves state-of-the-art performance on the LIBERO and RoboCasa GR1 benchmarks, as well as in real-world platforms.

cs.CV

Lean Learning Beyond Clouds: Efficient Discrepancy-Conditioned Optical-SAR Fusion for Semantic Segmentation

Cloud occlusion severely degrades the semantic integrity of optical remote sensing imagery. While incorporating Synthetic Aperture Radar (SAR) provides complementary observations, achieving efficient global modeling and reliable cross-modal fusion under cloud interference remains challenging. Existing methods rely on dense global attention to capture long-range dependencies, yet such aggregation indiscriminately propagates cloud-induced noise. Improving robustness typically entails enlarging model capacity, which further increases computational overhead. Given the large-scale and high-resolution nature of remote sensing applications, such computational demands hinder practical deployment, leading to an efficiency-reliability trade-off. To address this dilemma, we propose EDC, an efficiency-oriented and discrepancy-conditioned optical-SAR semantic segmentation framework. A tri-stream encoder with Carrier Tokens enables compact global context modeling with reduced complexity. To prevent noise contamination, we introduce a Discrepancy-Conditioned Hybrid Fusion (DCHF) mechanism that selectively suppresses unreliable regions during global aggregation. In addition, an auxiliary cloud removal branch with teacher-guided distillation enhances semantic consistency under occlusion. Extensive experiments demonstrate that EDC achieves superior accuracy and efficiency, improving mIoU by 0.56\% and 0.88\% on M3M-CR and WHU-OPT-SAR, respectively, while reducing the number of parameters by 46.7\% and accelerating inference by 1.98$\times$. Our implementation is available at https://github.com/mengcx0209/EDC.

cs.CV

S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi-step video generation or noisy one-step feature extraction, cannot simultaneously guarantee real-time inference and high-fidelity foresight. To address this limitation, we propose S-VAM, a shortcut video-action model that foresees coherent geometric and semantic representations via a single forward pass. Serving as a stable blueprint, these foreseen representations significantly simplify the action prediction. To enable this efficient shortcut, we introduce a novel self-distillation strategy that condenses structured generative priors of multi-step denoising into one-step inference. Specifically, vision foundation model (VFM) representations extracted from the diffusion model's own multi-step generated videos provide teacher targets. Lightweight decouplers, as students, learn to directly map noisy one-step features to these targets. Extensive experiments in simulation and the real world demonstrate that our S-VAM outperforms state-of-the-art methods, enabling efficient and precise manipulation in complex environments. Our project page is https://haodong-yan.github.io/S-VAM/

cs.CV

Optical Continuum Light Curves and Bolometric Energy Estimates of Solar White-light Flares

Solar white-light flares (WLFs) are solar flares exhibiting enhanced emission in the optical continuum. They are critical for understanding energy release and transport mechanisms in solar flares and for conducting comparative studies with stellar WLFs. However, the scarcity of accurately and reliably measured optical continuum light curves for solar WLFs significantly hampers related studies. Based on the optimized solar WLF identification method, we construct a dataset of optical continuum light curves for 70 solar WLFs using 6173 Å continuum intensity images from the Solar Dynamics Observatory. Moreover, for each solar WLF event, we also provide the location of the white-light emission enhancement signals and key parameters including bolometric energies and durations derived from both the traditional fixed-temperature blackbody model and the refined variable-temperature blackbody model. This dataset will serve as a valuable resource for future statistical investigations of solar WLFs and for comparative studies between solar and stellar flares.

astro-ph.SR

SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving

Sparse Perception Models (SPMs) adopt a query-driven paradigm that forgoes explicit dense BEV or volumetric construction, enabling highly efficient computation and accelerated inference. In this paper, we introduce SQS, a novel query-based splatting pre-training specifically designed to advance SPMs in autonomous driving. SQS introduces a plug-in module that predicts 3D Gaussian representations from sparse queries during pre-training, leveraging self-supervised splatting to learn fine-grained contextual features through the reconstruction of multi-view images and depth maps. During fine-tuning, the pre-trained Gaussian queries are seamlessly integrated into downstream networks via query interaction mechanisms that explicitly connect pre-trained queries with task-specific queries, effectively accommodating the diverse requirements of occupancy prediction and 3D object detection. Extensive experiments on autonomous driving benchmarks demonstrate that SQS delivers considerable performance gains across multiple query-based 3D perception tasks, notably in occupancy prediction and 3D object detection, outperforming prior state-of-the-art pre-training approaches by a significant margin (i.e., +1.3 mIoU on occupancy prediction and +1.0 NDS on 3D detection).

cs.CV

Sun-as-a-star Analysis of the Solar Eruption Source Region Using Ha Spectroscopic Observations of CHASE

Sun-as-a-star analyses serve as a bridge for comparative studies on solar and stellar activities. To investigate the typical Sun-as-a-star Ha temporal spectral characteristics in solar eruption source regions, we analyzed five different types of solar eruptions, using spectroscopic data from the Chinese Ha Solar Explorer (CHASE). Because the spatially-integral Ha spectrum of source region is mainly contributed by emission from heated plasma in flare ribbons and absorption from cold plasma in evolving filaments, we separately analyze the sub-regions of the source region dominated by different dynamical processes. It is revealed that filament eruptions show emission near Ha line center, accompanied by blueshifted/redshifted absorption, while flare ribbons show Ha line center emission with red asymmetry and line broadening. Moreover, a special spectral signature likely associated with coronal mass ejections (CMEs) is identified: prominent blueshifted absorption without a clear deceleration phase, along with redshifted absorption, which can be used as a probe when searching stellar CMEs. Furthermore, in the X9.0 flare (SOL2024-10-03T12:18) accompanied by a violent CME, the expected blueshifted signal is not visible in the spatially-integral Ha spectra. This suggests that filament-lifting signals associated with CMEs in the source region can be obscured by the simultaneous dominant flare-ribbon emission within the integration region, which may explain the relatively small number of confirmed stellar CMEs observed in Ha. We also find that comparison between the Ha and UV spectral observations can effectively reveal the velocity evolution of erupting filaments and potential existence of associated CMEs.

astro-ph.SR

VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving

This paper introduces VisionPAD, a novel self-supervised pre-training paradigm designed for vision-centric algorithms in autonomous driving. In contrast to previous approaches that employ neural rendering with explicit depth supervision, VisionPAD utilizes more efficient 3D Gaussian Splatting to reconstruct multi-view representations using only images as supervision. Specifically, we introduce a self-supervised method for voxel velocity estimation. By warping voxels to adjacent frames and supervising the rendered outputs, the model effectively learns motion cues in the sequential data. Furthermore, we adopt a multi-frame photometric consistency approach to enhance geometric perception. It projects adjacent frames to the current frame based on rendered depths and relative poses, boosting the 3D geometric representation through pure image supervision. Extensive experiments on autonomous driving datasets demonstrate that VisionPAD significantly improves performance in 3D object detection, occupancy prediction and map segmentation, surpassing state-of-the-art pre-training strategies by a considerable margin.

cs.CV

Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models

Large Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate LLMs with an important representation. To effectively encode occupancy as input for the LLM and address the category imbalances associated with occupancy, we propose Motion Separation Variational Autoencoder (MS-VAE). This innovative approach utilizes prior knowledge to distinguish dynamic objects from static scenes before inputting them into a tailored Variational Autoencoder (VAE). This separation enhances the model's capacity to concentrate on dynamic trajectories while effectively reconstructing static scenes. The efficacy of Occ-LLM has been validated across key tasks, including 4D occupancy forecasting, self-ego planning, and occupancy-based scene question answering. Comprehensive evaluations demonstrate that Occ-LLM significantly surpasses existing state-of-the-art methodologies, achieving gains of about 6\% in Intersection over Union (IoU) and 4\% in mean Intersection over Union (mIoU) for the task of 4D occupancy forecasting. These findings highlight the transformative potential of Occ-LLM in reshaping current paradigms within robotic and autonomous driving.

cs.RO

DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation

In the realm of image generation, creating customized images from visual prompt with additional textual instruction emerges as a promising endeavor. However, existing methods, both tuning-based and tuning-free, struggle with interpreting the subject-essential attributes from the visual prompt. This leads to subject-irrelevant attributes infiltrating the generation process, ultimately compromising the personalization quality in both editability and ID preservation. In this paper, we present DisEnvisioner, a novel approach for effectively extracting and enriching the subject-essential features while filtering out -irrelevant information, enabling exceptional customization performance, in a tuning-free manner and using only a single image. Specifically, the feature of the subject and other irrelevant components are effectively separated into distinctive visual tokens, enabling a much more accurate customization. Aiming to further improving the ID consistency, we enrich the disentangled features, sculpting them into more granular representations. Experiments demonstrate the superiority of our approach over existing methods in instruction response (editability), ID consistency, inference speed, and the overall image quality, highlighting the effectiveness and efficiency of DisEnvisioner. Project page: https://disenvisioner.github.io/.

cs.CV

OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction

We present OmniBooth, an image generation framework that enables spatial control with instance-level multi-modal customization. For all instances, the multimodal instruction can be described through text prompts or image references. Given a set of user-defined masks and associated text or image guidance, our objective is to generate an image, where multiple objects are positioned at specified coordinates and their attributes are precisely aligned with the corresponding guidance. This approach significantly expands the scope of text-to-image generation, and elevates it to a more versatile and practical dimension in controllability. In this paper, our core contribution lies in the proposed latent control signals, a high-dimensional spatial feature that provides a unified representation to integrate the spatial, textual, and image conditions seamlessly. The text condition extends ControlNet to provide instance-level open-vocabulary generation. The image condition further enables fine-grained control with personalized identity. In practice, our method empowers users with more flexibility in controllable generation, as users can choose multi-modal conditions from text or images as needed. Furthermore, thorough experiments demonstrate our enhanced performance in image synthesis fidelity and alignment across different tasks and datasets. Project page: https://len-li.github.io/omnibooth-web/

cs.CV

SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs

The advancement of autonomous driving is increasingly reliant on high-quality annotated datasets, especially in the task of 3D occupancy prediction, where the occupancy labels require dense 3D annotation with significant human effort. In this paper, we propose SyntheOcc, which denotes a diffusion model that Synthesize photorealistic and geometric-controlled images by conditioning Occupancy labels in driving scenarios. This yields an unlimited amount of diverse, annotated, and controllable datasets for applications like training perception models and simulation. SyntheOcc addresses the critical challenge of how to efficiently encode 3D geometric information as conditional input to a 2D diffusion model. Our approach innovatively incorporates 3D semantic multi-plane images (MPIs) to provide comprehensive and spatially aligned 3D scene descriptions for conditioning. As a result, SyntheOcc can generate photorealistic multi-view images and videos that faithfully align with the given geometric labels (semantics in 3D voxel space). Extensive qualitative and quantitative evaluations of SyntheOcc on the nuScenes dataset prove its effectiveness in generating controllable occupancy datasets that serve as an effective data augmentation to perception models.

cs.CV

Statistics of Solar White-Light Flares I: Optimization of Identification Methods and Application

White-light flares (WLFs) are energetic activity in stellar atmosphere. However, the observed solar WLF is relatively rare compared to stellar WLFs or solar flares observed at other wavelengths, limiting our further understanding solar/stellar WLFs through statistical studies. By analyzing flare observations from the \emph{Solar Dynamics Observatory (SDO)}, here we improve WLF identification methods for obtaining more solar WLFs and their accurate light curves from two aspects: 1) imposing constraints defined by the typical temporal and spatial distribution characteristics of WLF-induced signals; 2) setting the intrinsic threshold for each pixel in the flare ribbon region according to its inherent background fluctuation rather than a fixed threshold for the whole region. Applying the optimized method to 90 flares (30 C-class ones, 30 M-class ones, and 30 X-class ones) for a statistical study, we identified a total of 9 C-class WLFs, 18 M-class WLFs, and 28 X-class WLFs. The WLF identification rate of C-class flares reported here reaches 30\%, which is the highest to date to our best knowledge. It is also revealed that in each GOES energy level, the proportion of WLFs is higher in confined flares than that in eruptive flares. Moreover, a power-law relation is found between the WLF energy (\emph{E}) and duration ($τ$): $τ\propto {E}^{0.22}$, similar to those of solar hard/soft X-ray flares and other stellar WLFs. These results indicate that we could recognize more solar WLFs through optimizing the identification method, which will lay a base for future statistical and comparison study of solar and stellar WLFs.

astro-ph.SR