SearcharxivSearch

arXiv subjects

Jian Ren

Publications and source records attributed to Jian Ren.

At least 19 recordsLinked to original sources

Mergers Drive Structural Complexity but Not Starbursts in Lyman-$\alpha$ Emitters at $3 < z < 4$: A JWST Spatially Resolved View

Recent observations with the James Webb Space Telescope (JWST) reveal that the merger fraction among Ly$\alpha$ emitters (LAEs) at redshifts $z > 3$ is significantly higher than previously estimated. In this study, we focus on three high signal-to-noise merging LAE systems at $3 < z < 4$, selected from the VLT/MUSE-Deep survey in the GOODS-S field. We combine new \textit{JWST}/NIRCam broadband and medium-band imaging with archival \textit{HST}/ACS data to perform spatially resolved spectral energy distribution (SED) fitting using the \textsc{Bagpipes} software package. Our analysis reveals that two of the systems are minor mergers, while the third is a major merger. The close agreement between spatially resolved and integrated stellar mass estimates indicates that recent star formation does not significantly outshine the light from older stellar populations in these systems. Moreover, both the individual components and the systems as a whole lie on the star-forming main sequence, further supporting the conclusion that these mergers have not yet triggered substantial starburst activity. Furthermore, we detect prominent color gradients and disturbed dust distributions in these merging systems, indicating that the mergers have already induced significant internal structural perturbations. These morphological and dust-related changes may facilitate the escape of Ly$\alpha$ photons -- potentially through mechanisms such as gas redistribution or a reduced covering fraction of neutral hydrogen -- thereby playing a key role in shaping the observed properties of LAEs.

astro-ph.GA

FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents

Scientific document representation learning provides powerful embeddings for various tasks, while current methods face challenges across three approaches. 1) Contrastive training with citation-structural signals underutilizes citation information and still generates single-vector representations. 2) Fine-grained representation learning, which generates multiple vectors at the sentence or aspect level, requires costly integration and lacks domain generalization. 3) Task-aware learning depends on manually predefined task categorization, overlooking nuanced task distinctions and requiring extra training data for task-specific modules. To address these problems, we propose a new method that unifies the three approaches for better representations, namely FLeW. Specifically, we introduce a novel triplet sampling method that leverages citation intent and frequency to enhance citation-structural signals for training. Citation intents (background, method, result), aligned with the general structure of scientific writing, facilitate a domain-generalized facet partition for fine-grained representation learning. Then, we adopt a simple weight search to adaptively integrate three facet-level embeddings into a task-specific document embedding without task-aware fine-tuning. Experiments show the applicability and robustness of FLeW across multiple scientific tasks and fields, compared to prior models.

cs.IR

The Size Evolution and the Size-Mass Relation of Lyman-Alpha Emitters across $3 \lesssim z < 7$ as Observed by JWST

Understanding the morphological structures of Lyman-alpha emitters (LAEs) is crucial for unveiling their formation pathways and the physical origins of Ly$\alpha$ emission. However, the evolution of their sizes and structural scaling relations remains debated. In this study, we analyze a large sample of 876 spectroscopically confirmed LAEs at $3 \lesssim z < 7$, selected from the MUSE, VANDELS, and CANDELSz7 surveys in the GOODS-S, UDS, and COSMOS fields. Utilizing James Webb Space Telescope (JWST) NIRCam imaging data, we measure their rest-frame UV and optical V-band effective radii ($R_{\rm e}$) through two-dimensional S\'{e}rsic profile fitting. Our results show that these LAEs are generally compact, with a median $R_{\rm e,UV}$ of 0.50$^{+0.30}_{-0.24}$ kpc and a median $R_{\rm e,V}$ of 0.57$^{+0.33}_{-0.24}$ kpc. The size evolution follows $R_{\rm e,UV} \propto (1 + z)^{-0.91 \pm 0.10}$ and $R_{\rm e,V} \propto (1 + z)^{-0.93 \pm 0.18}$, respectively. Their UV and optical sizes are statistically comparable, indicating negligible UV-to-optical color gradients. For the first time, we establish the rest-frame optical size-mass relation for LAEs at $z>3$, finding slopes comparable to typical star-forming galaxies (SFGs), but with slightly smaller sizes at a given stellar mass. These results provide important clues for understanding structural evolution of LAEs in the early universe.

astro-ph.GA

The JWST Unveils the Bimodal Nature of Lyman Alpha Emitters at 3 <z<7: Pristine versus Merger-Driven Populations

We present a systematic study of merging galaxies among Lyman-alpha emitters (LAEs) using JWST/NIRCam high-resolution imaging data. From a large sample of 817 spectroscopically confirmed LAEs at $3 8.5$) and bright ($M_{\rm UV}<-19.5$) systems. At fixed $M_*$ and $M_{\rm UV}$, we find negligible differences in the UV slope ($\beta$) between late-stage mergers and isolated LAEs; however, a clear bimodal distribution emerges in the $M_*$-sSFR plane, where isolated LAEs peak at $\log(M_*/M_\odot)\approx7.8$ and $\log({\rm sSFR/yr^{-1}})\approx-7.4$, and late-stage mergers peak at $\log(M_*/M_\odot)\approx8.6$ and $\log({\rm sSFR/yr^{-1}})\approx-7.6$. Our results reveal two evolutionary classes -- Pristine LAEs, low-mass ($M_*<10^{8.5}M_\odot$), isolated systems that represent early-stage galaxies with minimal merger interactions, and Merger-driven LAEs, massive ($M_*>10^{8.5}M_\odot$) systems in which mergers enhance star formation and facilitate the escape of Lyman-alpha photons or accrete pristine LAEs -- both of which are consistent with both observational and theoretical expectations and collectively demonstrate that mergers are a central driver of LAE evolution across the first two billion years.

astro-ph.GA

AI-Assisted NLOS Sensing for RIS-Based Indoor Localization in Smart Factories

In the era of Industry 4.0, precise indoor localization is vital for automation and efficiency in smart factories. Reconfigurable Intelligent Surfaces (RIS) are emerging as key enablers in 6G networks for joint sensing and communication. However, RIS faces significant challenges in Non-Line-of-Sight (NLOS) and multipath propagation, particularly in localization scenarios, where detecting NLOS conditions is crucial for ensuring not only reliable results and increased connectivity but also the safety of smart factory personnel. This study introduces an AI-assisted framework employing a Convolutional Neural Network (CNN) customized for accurate Line-of-Sight (LOS) and Non-Line-of-Sight (NLOS) classification to enhance RIS-based localization using measured, synthetic, mixed-measured, and mixed-synthetic experimental data, that is, original, augmented, slightly noisy, and highly noisy data, respectively. Validated through such data from three different environments, the proposed customized-CNN (cCNN) model achieves {95.0\%-99.0\%} accuracy, outperforming standard pre-trained models like Visual Geometry Group 16 (VGG-16) with an accuracy of {85.5\%-88.0\%}. By addressing RIS limitations in NLOS scenarios, this framework offers scalable and high-precision localization solutions for 6G-enabled smart factories.

eess.SP

Evaluating the Accuracy of Non-parametric Galaxy Morphological Indicator Measurements in the CSST Imaging Survey

The Chinese Space Station Telescope (CSST) is China's upcoming next-generation ultraviolet and optical survey telescope, with imaging resolution capabilities comparable to the Hubble Space Telescope (HST). In this study, we utilized a comprehensive sample of 3,679 CSST realistic mock galaxies constructed from HST CANDELS/GOODS-North deep imaging observations, with stellar masses $\log\left(M_{*} / M_{\odot}\right) > 9.0$ and redshifts $z < 2$. We evaluate the detection capabilities of CSST surveys and the accuracy in measuring the non-parametric morphological indicators ($C$, $A$, $Gini$, $M_{\rm 20}$, $A_{\rm O}$, $D_{\rm O}$) of galaxies. Our findings show that in terms of galaxy detection capabilities, CSST's deep field surveys can achieve the same level as HST's deep field observations; however, in wide-field surveys, CSST exhibits a significant deficiency in detecting high-redshift, low-mass, low-surface-brightness galaxies. Regarding the measurement of galaxy morphology, CSST's deep field surveys achieve high accuracy across all indicators except for the asymmetry indicator ($A$), whereas its wide-field surveys suffer from significant systematic biases. We thus provide simple correction functions to adjust the non-parametric morphological indicators obtained from CSST's wide-field and deep-field observations, thereby aligning CSST measurements with those from HST. This adjustment enables the direct application of non-parametric morphological classification methods originally developed for HST data to galaxies observed by CSST.

astro-ph.GA

The Evolution of Size and Merger Fraction of Submillimeter Galaxies across $1 < z \lesssim 6$ as Observed by JWST

Precise tracking of the growth in galaxy size and the evolution of merger fractions with redshift is vital for understanding the formation history of submillimeter galaxies (SMGs). This study investigates these evolutions over a broad redshift range ($1 < z \lesssim 6$), using a sample of 222 SMGs with a median redshift of $z = 2.61^{+0.89}_{-0.82}$ identified by ALMA and JCMT, enhanced by the advanced imaging capabilities of the JWST/NIRCam and MIRI. We find significant evolution in effective radii ($R_e$) in rest-frame V-band ($R_e \propto (1 + z)^{-0.87 \pm 0.08}$) and near-infrared (NIR) band ($R_e \propto (1 + z)^{-0.88 \pm 0.11}$), with the NIR size evolution resembling that of massive star-forming galaxies at lower redshift. Visual inspections reveal a major merger fraction of $24.3 \pm 3.7\%$ and an interaction fraction of up to $48.4 \pm 11.1\%$. The major merger fraction exhibits an increase from 14.7$\pm9.1$\% at $z = 1$ to 26.6$\pm 8.4$\% at $z = 3$, after which it remains approximately constant across the redshift range $3 < z < 6$. In contrast, the interaction fraction remains relatively stable across the range $2 < z < 5$. Our results indicate that late-stage major mergers are not the primary formation mechanism for SMGs at $z<3$, while interactions appear to play a significant role across the broader redshift range of $1<z<6$. Additionally, HST-based major merger identifications may overestimate the true fraction by a factor of 1.7 at $z \sim 2$. These findings highlight the varying roles of mergers and interactions in driving the formation of massive, dusty star-forming galaxies across different redshifts.

astro-ph.GA

Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach

We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video dataset, PointVid, is then used to fine-tune a latent diffusion model, enabling it to track 2D objects with 3D Cartesian coordinates. Building on this, we regularize the shape and motion of objects in the video to eliminate undesired artifacts, e.g., non-physical deformation. Consequently, we enhance the quality of generated RGB videos and alleviate common issues like object morphing, which are prevalent in current video models due to a lack of shape awareness. With our 3D augmentation and regularization, our model is capable of handling contact-rich scenarios such as task-oriented videos, where 3D information is essential for perceiving shape and motion of interacting solids. Our method can be seamlessly integrated into existing video diffusion models to improve their visual plausibility.

cs.CV

Wonderland: Navigating 3D Scenes from a Single Image

How can one efficiently generate high-quality, wide-scope 3D scenes from arbitrary single images? Existing methods suffer several drawbacks, such as requiring multi-view data, time-consuming per-scene optimization, distorted geometry in occluded areas, and low visual quality in backgrounds. Our novel 3D scene reconstruction pipeline overcomes these limitations to tackle the aforesaid challenge. Specifically, we introduce a large-scale reconstruction model that leverages latents from a video diffusion model to predict 3D Gaussian Splattings of scenes in a feed-forward manner. The video diffusion model is designed to create videos precisely following specified camera trajectories, allowing it to generate compressed video latents that encode multi-view information while maintaining 3D consistency. We train the 3D reconstruction model to operate on the video latent space with a progressive learning strategy, enabling the efficient generation of high-quality, wide-scope, and generic 3D scenes. Extensive evaluations across various datasets affirm that our model significantly outperforms existing single-view 3D scene generation methods, especially with out-of-domain images. Thus, we demonstrate for the first time that a 3D reconstruction model can effectively be built upon the latent space of a diffusion model in order to realize efficient 3D scene generation.

cs.CV

SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image generation, video generation models require more computation and are thus hosted mostly on cloud servers, limiting broader adoption among content creators. In this work, we propose a comprehensive acceleration framework to bring the power of the large-scale video diffusion model to the hands of edge users. From the network architecture scope, we initialize from a compact image backbone and search out the design and arrangement of temporal layers to maximize hardware efficiency. In addition, we propose a dedicated adversarial fine-tuning algorithm for our efficient model and reduce the denoising steps to 4. Our model, with only 0.6B parameters, can generate a 5-second video on an iPhone 16 PM within 5 seconds. Compared to server-side models that take minutes on powerful GPUs to generate a single video, we accelerate the generation by magnitudes while delivering on-par quality.

cs.CV

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training

Existing text-to-image (T2I) diffusion models face several limitations, including large model sizes, slow runtime, and low-quality generation on mobile devices. This paper aims to address all of these challenges by developing an extremely small and fast T2I model that generates high-resolution and high-quality images on mobile platforms. We propose several techniques to achieve this goal. First, we systematically examine the design choices of the network architecture to reduce model parameters and latency, while ensuring high-quality generation. Second, to further improve generation quality, we employ cross-architecture knowledge distillation from a much larger model, using a multi-level approach to guide the training of our model from scratch. Third, we enable a few-step generation by integrating adversarial guidance with knowledge distillation. For the first time, our model SnapGen, demonstrates the generation of 1024x1024 px images on a mobile device around 1.4 seconds. On ImageNet-1K, our model, with only 372M parameters, achieves an FID of 2.06 for 256x256 px generation. On T2I benchmarks (i.e., GenEval and DPG-Bench), our model with merely 379M parameters, surpasses large-scale models with billions of parameters at a significantly smaller size (e.g., 7x smaller than SDXL, 14x smaller than IF-XL).

cs.CV

AsCAN: Asymmetric Convolution-Attention Networks for Efficient Recognition and Generation

Neural network architecture design requires making many crucial decisions. The common desiderata is that similar decisions, with little modifications, can be reused in a variety of tasks and applications. To satisfy that, architectures must provide promising latency and performance trade-offs, support a variety of tasks, scale efficiently with respect to the amounts of data and compute, leverage available data from other tasks, and efficiently support various hardware. To this end, we introduce AsCAN -- a hybrid architecture, combining both convolutional and transformer blocks. We revisit the key design principles of hybrid architectures and propose a simple and effective \emph{asymmetric} architecture, where the distribution of convolutional and transformer blocks is \emph{asymmetric}, containing more convolutional blocks in the earlier stages, followed by more transformer blocks in later stages. AsCAN supports a variety of tasks: recognition, segmentation, class-conditional image generation, and features a superior trade-off between performance and latency. We then scale the same architecture to solve a large-scale text-to-image task and show state-of-the-art performance compared to the most recent public and commercial models. Notably, even without any computation optimization for transformer blocks, our models still yield faster inference speed than existing works featuring efficient attention mechanisms, highlighting the advantages and the value of our approach.

cs.CV

Scalable Ranked Preference Optimization for Text-to-Image Generation

Direct Preference Optimization (DPO) has emerged as a powerful approach to align text-to-image (T2I) models with human feedback. Unfortunately, successful application of DPO to T2I models requires a huge amount of resources to collect and label large-scale datasets, e.g., millions of generated paired images annotated with human preferences. In addition, these human preference datasets can get outdated quickly as the rapid improvements of T2I models lead to higher quality images. In this work, we investigate a scalable approach for collecting large-scale and fully synthetic datasets for DPO training. Specifically, the preferences for paired images are generated using a pre-trained reward function, eliminating the need for involving humans in the annotation process, greatly improving the dataset collection efficiency. Moreover, we demonstrate that such datasets allow averaging predictions across multiple models and collecting ranked preferences as opposed to pairwise preferences. Furthermore, we introduce RankDPO to enhance DPO-based methods using the ranking feedback. Applying RankDPO on SDXL and SD3-Medium models with our synthetically generated preference dataset "Syn-Pic" improves both prompt-following (on benchmarks like T2I-Compbench, GenEval, and DPG-Bench) and visual quality (through user studies). This pipeline presents a practical and scalable solution to develop better preference datasets to enhance the performance of text-to-image models.

cs.CV

MaskControl: Spatio-Temporal Control for Masked Motion Synthesis

Recent advances in motion diffusion models have enabled spatially controllable text-to-motion generation. However, these models struggle to achieve high-precision control while maintaining high-quality motion generation. To address these challenges, we propose MaskControl, the first approach to introduce controllability to the generative masked motion model. Our approach introduces two key innovations. First, \textit{Logits Regularizer} implicitly perturbs logits at training time to align the distribution of motion tokens with the controlled joint positions, while regularizing the categorical token prediction to ensure high-fidelity generation. Second, \textit{Logit Optimization} explicitly optimizes the predicted logits during inference time, directly reshaping the token distribution that forces the generated motion to accurately align with the controlled joint positions. Moreover, we introduce \textit{Differentiable Expectation Sampling (DES)} to combat the non-differential distribution sampling process encountered by logits regularizer and optimization. Extensive experiments demonstrate that MaskControl outperforms state-of-the-art methods, achieving superior motion quality (FID decreases by ~77\%) and higher control precision (average error 0.91 vs. 1.08). Additionally, MaskControl enables diverse applications, including any-joint-any-frame control, body-part timeline control, and zero-shot objective control. Video visualization can be found at https://www.ekkasit.com/ControlMM-page/

cs.CV

Efficient Training with Denoised Neural Weights

Good weight initialization serves as an effective measure to reduce the training cost of a deep neural network (DNN) model. The choice of how to initialize parameters is challenging and may require manual tuning, which can be time-consuming and prone to human error. To overcome such limitations, this work takes a novel step towards building a weight generator to synthesize the neural weights for initialization. We use the image-to-image translation task with generative adversarial networks (GANs) as an example due to the ease of collecting model weights spanning a wide range. Specifically, we first collect a dataset with various image editing concepts and their corresponding trained weights, which are later used for the training of the weight generator. To address the different characteristics among layers and the substantial number of weights to be predicted, we divide the weights into equal-sized blocks and assign each block an index. Subsequently, a diffusion model is trained with such a dataset using both text conditions of the concept and the block indexes. By initializing the image translation model with the denoised weights predicted by our diffusion model, the training requires only 43.3 seconds. Compared to training from scratch (i.e., Pix2pix), we achieve a 15x training time acceleration for a new concept while obtaining even better image generation quality.

cs.CV

Lightweight Predictive 3D Gaussian Splats

Recent approaches representing 3D objects and scenes using Gaussian splats show increased rendering speed across a variety of platforms and devices. While rendering such representations is indeed extremely efficient, storing and transmitting them is often prohibitively expensive. To represent large-scale scenes, one often needs to store millions of 3D Gaussians, occupying gigabytes of disk space. This poses a very practical limitation, prohibiting widespread adoption.Several solutions have been proposed to strike a balance between disk size and rendering quality, noticeably reducing the visual quality. In this work, we propose a new representation that dramatically reduces the hard drive footprint while featuring similar or improved quality when compared to the standard 3D Gaussian splats. When compared to other compact solutions, ours offers higher quality renderings with significantly reduced storage, being able to efficiently run on a mobile device in real-time. Our key observation is that nearby points in the scene can share similar representations. Hence, only a small ratio of 3D points needs to be stored. We introduce an approach to identify such points which are called parent points. The discarded points called children points along with attributes can be efficiently predicted by tiny MLPs.

cs.GR

SF-V: Single Forward Video Generation Model

Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models require multiple denoising steps during sampling, resulting in high computational costs. In this work, we propose a novel approach to obtain single-step video generation models by leveraging adversarial training to fine-tune pre-trained video diffusion models. We show that, through the adversarial training, the multi-steps video diffusion model, i.e., Stable Video Diffusion (SVD), can be trained to perform single forward pass to synthesize high-quality videos, capturing both temporal and spatial dependencies in the video data. Extensive experiments demonstrate that our method achieves competitive generation quality of synthesized videos with significantly reduced computational overhead for the denoising process (i.e., around $23\times$ speedup compared with SVD and $6\times$ speedup compared with existing works, with even better generation quality), paving the way for real-time video synthesis and editing. More visualization results are made publicly available at https://snap-research.github.io/SF-V.

cs.CV

BitsFusion: 1.99 bits Weight Quantization of Diffusion Model

Diffusion-based image generation models have achieved great success in recent years by showing the capability of synthesizing high-quality content. However, these models contain a huge number of parameters, resulting in a significantly large model size. Saving and transferring them is a major bottleneck for various applications, especially those running on resource-constrained devices. In this work, we develop a novel weight quantization method that quantizes the UNet from Stable Diffusion v1.5 to 1.99 bits, achieving a model with 7.9X smaller size while exhibiting even better generation quality than the original one. Our approach includes several novel techniques, such as assigning optimal bits to each layer, initializing the quantized model for better performance, and improving the training strategy to dramatically reduce quantization error. Furthermore, we extensively evaluate our quantized model across various benchmark datasets and through human evaluation to demonstrate its superior generation quality.

cs.CV