SearcharxivSearch

arXiv subjects

Sijia Cai

Publications and source records attributed to Sijia Cai.

16 recordsLinked to original sources

Low Ly$α$ Visibility in Galaxy Overdensities: Reionization Topology and Neutral-Fraction Ceilings from DIVER over $4.8<z<11$

Ly-alpha emission is widely used to trace cosmic reionization, but its interpretation depends on how Ly-alpha visibility varies with galaxy environment. We use deep JWST/NIRSpec observations from Deep Insights into UV Spectroscopy at the Epoch of Reionization (DIVER) in GOODS-N to measure Ly-alpha visibility for 250 galaxies at 4.8 25 A. We combine these measurements with H-alpha and [O III] emitters from JWST/NIRCam wide-field slitless spectroscopy to map the density field around each DIVER galaxy. Galaxies with high Ly-alpha equivalent widths (W_Lyalpha>25 A) or high effective Ly-alpha escape fractions (f_esc,Lyalpha^eff>0.05) tend to lie farther from nearby H-alpha and [O III] emitters than galaxies with lower Ly-alpha visibility. The clearest signal occurs near the prominent GOODS-N overdensity at z~5.2, where fewer than 15% of galaxies show strong Ly-alpha emission. This trend is opposite to the simplest inside-out reionization expectation that overdensities produce larger ionized regions and enhance Ly-alpha visibility. Possible explanations include circumgalactic and local intergalactic opacity, dense absorbers, and gas kinematics. We also derive an empirical upper envelope for f_esc,Lyalpha^eff and calibrate it with reionization simulations. Interpreting this envelope as a limiting IGM-attenuation signal gives neutral-fraction ceilings of _max=0.36, 0.76, 0.74, 0.84, and 1.0 at z~5.2, 5.8, 6.7, 7.7, and 9.8, respectively. The z~8 ceiling disfavors an almost completely neutral IGM at this epoch. These results support patchy reionization already underway by z~8 and show that galaxy Ly-alpha visibility encodes both large-scale ionization topology and near-source gas structure.

astro-ph.GA

EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent content creation. Existing methods, usually utilizing a reference image as the style prior, suffer from content leakage, data scarcity and limited adaptability to long videos, leading to suboptimal results with severe style drift and motion distortion. For these issues, we present EchoStyle, a scalable text-driven framework to achieve high-quality stylization of videos with arbitrary lengths. To start with, we construct a video-to-video architecture to appropriately re-fuse the video content and the text style. To address data scarcity, we pioneer an automatic reverse-synthesis pipeline to establish V-Style20k, a large-scale stylization dataset of 20k high-quality video pairs. To facilitate long video stylization, we devise an init-follow-mode mechanism along with a sliding-window inference strategy. Extensive experiments demonstrate EchoStyle's excellent performance across a wide range of artistic styles, even comparable to leading closed-source solutions.

cs.CV

A massive barred spiral galaxy at z = 5.102 discovered by JWST

We report M1149-BSG-z5, a barred spiral galaxy at $z = 5.102$, identified in the parallel field of MACS J1149+2223 with JWST and HST. M1149-BSG-z5 is the highest redshift barred galaxy candidate to date. Both isophote ellipse fitting and structural modeling support a stellar bar of length $a_\mathrm{bar} \approx 4.5$ kpc, and extended spiral arms peaking at $r \approx 5.5$ kpc. M1149-BSG-z5 is a massive main sequence star-forming galaxy, with a stellar mass of $10^{10.45}\rm M_\odot$ and a star-formation rate of $144\,\rm M_\odot/yr$. A concentrated bulge is embedded in an extended disk with a global Sérsic index $n = 2.37$. With an effective radius of $R_{e} = 2.61\rm \ kpc$, M1149-BSG-z5 is larger than typical galaxies at $z \sim 5$ and comparable to barred galaxies at $2 < z < 4$. M1149-BSG-z5 also hosts a broad-line AGN, with a relatively low black-hole-to-stellar mass ratio of $\rm M_{\rm BH}/M_\ast\sim10^{-3}$. Its metal-enriched emission-line properties indicate that it is already chemically evolved. These properties imply M1149-BSG-z5 as an early-assembled and structurally evolved galaxy. We also find that M1149-BSG-z5 resides in an overdense region with a nearby companion galaxy, suggesting an interaction-driven bar formation mechanism. Its concentrated light, early assembly and main-sequence star formation also suggest baryon-dominated, gas-rich conditions, where gravitational instability can further accelerate the bar formation.

astro-ph.GA

A Strongly Lensed Ultra-faint Arc at $z \approx 10$ with an F200W excess in Abell S1063

Strong gravitational lensing provides a powerful route to probing intrinsically faint galaxies during the first few hundred million years of cosmic history. In this Letter, we report the identification of GAR10, a highly magnified F115W-dropout galaxy at $z\approx10$ in the Abell S1063 cluster field, using deep JWST/NIRCam imaging from the GLIMPSE and GO-1840 programs. The source shows an unusually blue ultraviolet (UV) continuum and a significant F200W excess relative to adjacent bands. Under our high-magnification lensing solution, we infer a median magnification of $μ=43^{+78}_{-20}$, corresponding to an intrinsic UV magnitude of $M_{\rm UV}\approx-15.8$. We use exploratory Prospector SED modeling to examine two physically motivated interpretations of the observed photometry. In Case I, GAR10 is described by an extremely metal-poor, continuum-dominated stellar population at $z=10.75_{-0.34}^{+0.41}$, with a blue UV slope of $β=-2.92\pm0.12$ and a low metallicity of $\log(Z/Z_\odot)=-3.56_{-0.85}^{+0.65}$, consistent with an extremely metal-poor or Pop III-like continuum-dominated interpretation under the adopted priors. In Case II, GAR10 is interpreted as an extremely young (1--3 Myr), high-ionization galaxy at $z=10.45_{-0.21}^{+0.11}$, in which the F200W excess is produced by intense rest-frame UV emission lines, including CIV, HeII, and CIII]. Both cases can partially reproduce the current photometry within the adopted priors, but they imply distinct ionizing sources, enrichment histories, and possible contributions to cosmic reionization. GAR10 therefore represents a rare laboratory for studying ultra-faint galaxy formation at cosmic dawn. Future JWST/NIRSpec spectroscopy will be essential to distinguishing between the steep continuum and emission-line origins of the F200W excess.

astro-ph.GA

Discovery of an Extremely Metal-Poor Galaxy at $z=3.654$ Using JWST Infrared Spectroscopy

We report the discovery of an extremely metal-poor galaxy at a redshift of z = 3.654, identified through infrared spectroscopy using the James Webb Space Telescope (JWST). This galaxy, CAPERS-39810, exhibits a metallicity of 12 + log(O/H) = $6.73\pm0.13$, indicative of its primitive chemical composition, resembling the early stages of galaxy formation in the Universe. We use JWST NIRSpec/MSA for spectroscopic analysis, complemented by photometric data from the COSMOS2025 catalog. Our analysis employs the R3 strong-line diagnostic method to estimate metallicity, due to the lack of auroral lines in the spectrum. The galaxy's emission lines, including Hb, [O III], Ha and He I, are clearly detected. The rest-frame equivalent widths of the strong hydrogen recombination lines are EW_0(Hb) = $184\pm48$ Åand EW_0(Ha) = $1144\pm48$ Å. Furthermore, we perform detailed spectral energy distribution modeling to derive a galaxy logarithmic stellar mass of $8.02^{+0.22}_{-0.34}$ $M_\odot$. This discovery adds to the growing body of evidence for the existence of very low-metallicity galaxies existed at cosmic noon of $z\approx3$, which are crucial for understanding the processes of chemical enrichment and star formation in young galaxies at the cosmic noon.

astro-ph.GA

AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References

Identity-preserving video generation offers powerful tools for creative expression, allowing users to customize videos featuring their beloved characters. However, prevailing methods are typically designed and optimized for a single identity reference. This underlying assumption restricts creative flexibility by inadequately accommodating diverse real-world input formats. Relying on a single source also constitutes an ill-posed scenario, causing an inherently ambiguous setting that makes it difficult for the model to faithfully reproduce an identity across novel contexts. To address these issues, we present AnyID, an ultra-fidelity identity-preservation video generation framework that features two core contributions. First, we introduce a scalable omni-referenced architecture that effectively unifies heterogeneous identity inputs (e.g., faces, portraits, and videos) into a cohesive representation. Second, we propose a primary-referenced generation paradigm, which designates one reference as a canonical anchor and uses a novel differential prompt to enable precise, attribute-level controllability. We conduct training on a large-scale, meticulously curated dataset to ensure robustness and high fidelity, and then perform a final fine-tuning stage using reinforcement learning. This process leverages a preference dataset constructed from human evaluations, where annotators performed pairwise comparisons of videos based on two key criteria: identity fidelity and prompt controllability. Extensive evaluations validate that AnyID achieves ultra-high identity fidelity as well as superior attribute-level controllability across different task settings.

cs.CV

EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer

Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only training objectives, which inherently bias models toward appearance fidelity at the expense of learning underlying kinematic principles. To address this, we introduce EchoMotion, a framework designed to model the joint distribution of appearance and human motion, thereby improving the quality of complex human action video generation. EchoMotion extends the DiT (Diffusion Transformer) framework with a dual-branch architecture that jointly processes tokens concatenated from different modalities. Furthermore, we propose MVS-RoPE (Motion-Video Syncronized RoPE), which offers unified 3D positional encoding for both video and motion tokens. By providing a synchronized coordinate system for the dual-modal latent sequence, MVS-RoPE establishes an inductive bias that fosters temporal alignment between the two modalities. We also propose a Motion-Video Two-Stage Training Strategy. This strategy enables the model to perform both the joint generation of complex human action videos and their corresponding motion sequences, as well as versatile cross-modal conditional generation tasks. To facilitate the training of a model with these capabilities, we construct HuMoVe, a large-scale dataset of approximately 80,000 high-quality, human-centric video-motion pairs. Our findings reveal that explicitly representing human motion is complementary to appearance, significantly boosting the coherence and plausibility of human-centric video generation.

cs.CV

A Metal-Free Galaxy at $z = 3.19$? Evidence of Late Population III Star Formation at Cosmic Noon

Star formation from metal-free gas, the hallmark of the first generation of Population III stars, was long assumed to occur only in the very early Universe. We report the discovery of MPG-CR3 (Metal-Pristine Galaxy COSMOS Redshift 3; hereafter CR3), an extremely metal-poor galaxy at redshift $z= 3.193\pm0.016$. From JWST, VLT, and Subaru observations, CR3 exhibits exceptionally strong Ly$α$, H$α$, and He I $λ$10830 emission. We measure rest-frame equivalent widths of EW$_0$(Ly$α$) $=822\pm101$ Angstrom and EW$_0$(H$α$) $=2814\pm327$ Angstrom, among the highest seen in star-forming systems. No metal lines, e.g. [O III] $λ\lambda4959,5007$, C IV $λ\lambda1548,1550$, have statistically significant detections, placing a 2-$σ$ upper limit on the gas-phase metallicity of 12+log(O/H) < 6.52 ($Z < 7\times10^{-3}\ Z_\odot$) with strong-line calibration established by JWST, making it the most metal-poor galaxy known at cosmic noon. Considering systematic uncertainties of $\gtrsim 0.3$ dex in the calibrations, the most conservative 2-$σ$ upper limit is set to 12+log(O/H) < 6.95. The observed Ly$α$/H$α$ flux ratio is $13.9\pm2.5$, indicating negligible dust attenuation. Spectral energy distribution modeling with Pop III stellar templates indicates a very young ($\sim2$ Myr), low-mass ($M_* \approx 6.1\times 10^5 M_\odot$) stellar population. Further, the photometric redshifts reveal that CR3 could reside in a slightly underdense environment ($δ\approx -0.12$). CR3 provides evidence that first-generation star formation could persist well after the epoch of reionization, challenging the conventional view that pristine star formation ended by $z\gtrsim6$.

astro-ph.GA

Discovery of a Pair of Galaxies with Both Hosting X-ray Binary Candidates at $z=2.544$

Among high-redshift galaxies, aside from active galactic nuclei (AGNs), X-ray binaries (XRBs) can be significant sources of X-ray emission. XRBs play a crucial role in galaxy evolution, reflecting the stellar populations of galaxies and regulating star formation through feedback, thereby shaping galaxy structure. In this study, we report a spectroscopically confirmed X-ray emitting galaxy pair (UDF3 and UDF3-2) at $z = 2.544$. By combining multi-wavelength observations from JWST/NIRSpec MSA spectra, JWST/NIRCam and MIRI imaging, Chandra, HST, VLT, ALMA, and VLA, we analyze the ionized emission lines, which are primarily driven by H II region-like processes. Additionally, we find that the mid-infrared radiation can be fully attributed to dust emission from galaxy themselves. Our results indicate that the X-ray emission from these two galaxies is dominated by high-mass XRBs, with luminosities of $L_X= (1.43\pm0.40) \times 10^{42} \, \text{erg} \, \text{s}^{-1}$ for UDF3, and $(0.40\pm0.12) \times 10^{42} \, \text{erg} \, \text{s}^{-1}$ for UDF3-2. Furthermore, we measure the star formation rate (SFR) of $529_{-88}^{+64}$ $M_\odot$ yr$^{-1}$ for UDF3, placing it $\approx$ 0.5 dex below the $L_X$/SFR-$z$ relation. This offset reflects the redshift-dependent enhancement of $L_X$/SFR-$z$ relation, which is influenced by metallicity and serves as a key observable for XRB evolution. In contrast, UDF3-2, with the SFR of $34_{-6}^{+6}$ $M_\odot$ yr$^{-1}$, aligns well with the $L_X$/SFR-$z$ relation. This galaxy pair represents the highest-redshift non-AGN-dominated galaxies with individual X-ray detections reported to date. This finding suggests that the contribution of XRBs to galaxy X-ray emission at high redshift may be underestimated.

astro-ph.GA

PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Models

Controllable generation is considered a potentially vital approach to address the challenge of annotating 3D data, and the precision of such controllable generation becomes particularly imperative in the context of data production for autonomous driving. Existing methods focus on the integration of diverse generative information into controlling inputs, utilizing frameworks such as GLIGEN or ControlNet, to produce commendable outcomes in controllable generation. However, such approaches intrinsically restrict generation performance to the learning capacities of predefined network architectures. In this paper, we explore the innovative integration of controlling information and introduce PerLDiff (\textbf{Per}spective-\textbf{L}ayout \textbf{Diff}usion Models), a novel method for effective street view image generation that fully leverages perspective 3D geometric information. Our PerLDiff employs 3D geometric priors to guide the generation of street view images with precise object-level control within the network learning process, resulting in a more robust and controllable output. Moreover, it demonstrates superior controllability compared to alternative layout control methods. Empirical results justify that our PerLDiff markedly enhances the precision of controllable generation on the NuScenes and KITTI datasets.

cs.CV

EchoShot: Multi-Shot Portrait Video Generation

Video diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge for multiple shots with identity consistency and flexible content controllability. In this work, we propose EchoShot, a native and scalable multi-shot framework for portrait customization built upon a foundation video diffusion model. To start with, we propose shot-aware position embedding mechanisms within video diffusion transformer architecture to model inter-shot variations and establish intricate correspondence between multi-shot visual content and their textual descriptions. This simple yet effective design enables direct training on multi-shot video data without introducing additional computational overhead. To facilitate model training within multi-shot scenario, we construct PortraitGala, a large-scale and high-fidelity human-centric video dataset featuring cross-shot identity consistency and fine-grained captions such as facial attributes, outfits, and dynamic motions. To further enhance applicability, we extend EchoShot to perform reference image-based personalized multi-shot generation and long video synthesis with infinite shot counts. Extensive evaluations demonstrate that EchoShot achieves superior identity consistency as well as attribute-level controllability in multi-shot portrait video generation. Notably, the proposed framework demonstrates potential as a foundational paradigm for general multi-shot video modeling.

cs.CV

RoScenes: A Large-scale Multi-view 3D Dataset for Roadside Perception

We introduce RoScenes, the largest multi-view roadside perception dataset, which aims to shed light on the development of vision-centric Bird's Eye View (BEV) approaches for more challenging traffic scenes. The highlights of RoScenes include significantly large perception area, full scene coverage and crowded traffic. More specifically, our dataset achieves surprising 21.13M 3D annotations within 64,000 $m^2$. To relieve the expensive costs of roadside 3D labeling, we present a novel BEV-to-3D joint annotation pipeline to efficiently collect such a large volume of data. After that, we organize a comprehensive study for current BEV methods on RoScenes in terms of effectiveness and efficiency. Tested methods suffer from the vast perception area and variation of sensor layout across scenes, resulting in performance levels falling below expectations. To this end, we propose RoBEV that incorporates feature-guided position embedding for effective 2D-3D feature assignment. With its help, our method outperforms state-of-the-art by a large margin without extra computational overhead on validation set. Our dataset and devkit will be made available at https://github.com/xiaosu-zhu/RoScenes.

cs.CV

CT3D++: Improving 3D Object Detection with Keypoint-induced Channel-wise Transformer

The field of 3D object detection from point clouds is rapidly advancing in computer vision, aiming to accurately and efficiently detect and localize objects in three-dimensional space. Current 3D detectors commonly fall short in terms of flexibility and scalability, with ample room for advancements in performance. In this paper, our objective is to address these limitations by introducing two frameworks for 3D object detection with minimal hand-crafted design. Firstly, we propose CT3D, which sequentially performs raw-point-based embedding, a standard Transformer encoder, and a channel-wise decoder for point features within each proposal. Secondly, we present an enhanced network called CT3D++, which incorporates geometric and semantic fusion-based embedding to extract more valuable and comprehensive proposal-aware information. Additionally, CT3D ++ utilizes a point-to-key bidirectional encoder for more efficient feature encoding with reduced computational cost. By replacing the corresponding components of CT3D with these novel modules, CT3D++ achieves state-of-the-art performance on both the KITTI dataset and the large-scale Way\-mo Open Dataset. The source code for our frameworks will be made accessible at https://github.com/hlsheng1/CT3D-plusplus.

cs.CV

Rethinking IoU-based Optimization for Single-stage 3D Object Detection

Since Intersection-over-Union (IoU) based optimization maintains the consistency of the final IoU prediction metric and losses, it has been widely used in both regression and classification branches of single-stage 2D object detectors. Recently, several 3D object detection methods adopt IoU-based optimization and directly replace the 2D IoU with 3D IoU. However, such a direct computation in 3D is very costly due to the complex implementation and inefficient backward operations. Moreover, 3D IoU-based optimization is sub-optimal as it is sensitive to rotation and thus can cause training instability and detection performance deterioration. In this paper, we propose a novel Rotation-Decoupled IoU (RDIoU) method that can mitigate the rotation-sensitivity issue, and produce more efficient optimization objectives compared with 3D IoU during the training stage. Specifically, our RDIoU simplifies the complex interactions of regression parameters by decoupling the rotation variable as an independent term, yet preserving the geometry of 3D IoU. By incorporating RDIoU into both the regression and classification branches, the network is encouraged to learn more precise bounding boxes and concurrently overcome the misalignment issue between classification and regression. Extensive experiments on the benchmark KITTI and Waymo Open Dataset validate that our RDIoU method can bring substantial improvement for the single-stage 3D object detection.

cs.CV

Improving 3D Object Detection with Channel-wise Transformer

Though 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed components such as keypoints sampling, set abstraction and multi-scale feature fusion to produce powerful 3D object representations. Such methods, however, have limited ability to capture rich contextual dependencies among points. In this paper, we leverage the high-quality region proposal network and a Channel-wise Transformer architecture to constitute our two-stage 3D object detection framework (CT3D) with minimal hand-crafted design. The proposed CT3D simultaneously performs proposal-aware embedding and channel-wise context aggregation for the point features within each proposal. Specifically, CT3D uses proposal's keypoints for spatial contextual modelling and learns attention propagation in the encoding module, mapping the proposal to point embeddings. Next, a new channel-wise decoding module enriches the query-key interaction via channel-wise re-weighting to effectively merge multi-level contexts, which contributes to more accurate object predictions. Extensive experiments demonstrate that our CT3D method has superior performance and excellent scalability. Remarkably, CT3D achieves the AP of 81.77% in the moderate car category on the KITTI test 3D detection benchmark, outperforms state-of-the-art 3D detectors.

cs.CV

Efficient Background Modeling Based on Sparse Representation and Outlier Iterative Removal

Background modeling is a critical component for various vision-based applications. Most traditional methods tend to be inefficient when solving large-scale problems. In this paper, we introduce sparse representation into the task of large scale stable background modeling, and reduce the video size by exploring its 'discriminative' frames. A cyclic iteration process is then proposed to extract the background from the discriminative frame set. The two parts combine to form our Sparse Outlier Iterative Removal (SOIR) algorithm. The algorithm operates in tensor space to obey the natural data structure of videos. Experimental results show that a few discriminative frames determine the performance of the background extraction. Further, SOIR can achieve high accuracy and high speed simultaneously when dealing with real video sequences. Thus, SOIR has an advantage in solving large-scale tasks.

cs.CV