SearcharxivSearch

arXiv subjects

Yuan Bian

Publications and source records attributed to Yuan Bian.

16 recordsLinked to original sources

The Lumina Project: Intergalactic Clumping and Recombination Sinks

Recombinations during the Epoch of Reionization are intrinsically inhomogeneous, with different regions of the intergalactic medium contributing unevenly depending on their density, temperature, ionization state, and spatial patchiness. We combine the high- and medium-resolution 95.5 cMpc Thesan-1 andh Thesan-2 runs with the significantly larger 500 cMpc Lumina simulation to measure clumping factors and recombination rates consistently across different resolutions and box sizes. We consider the standard ionized hydrogen clumping factor, $C_{\rm HII} \equiv \langle n_{\rm HII}^2\rangle/\langle n_{\rm HII}\rangle^2$, and a recombination-weighted clumping factor, $C_{\rm rec}$. Despite differences in resolution, volume, and reionization history, the simulations show an approximately universal clumping evolution at the 10-20% level when parametrized by the global ionized fraction $x_{\rm HII}$ rather than by redshift. Across all simulations, $C_{\rm rec}$ remains systematically below $C_{\rm HII}$, with the discrepancy increasing toward lower redshift as photoheating suppresses recombinations. In \lumina, the density-only prescription overpredicts the instantaneous recombination rate by factors of 1.29 at $z\approx8$ and 1.84 at $z\approx5$, and the cumulative recombination count by a factor of 1.45 by $z\approx5$. Mapping the recombination budget in the joint overdensity-temperature plane reveals that the dominant recombination ridges closely follow simple analytic thermal equilibrium bands. Finally, we introduce a phase-space recombination integral and define a phase-space clumping factor, $C_{\rm ps}(\Delta,T)$, which isolates the intrinsic recombination enhancement associated with ionization structure and thermal state at fixed overdensity and temperature.

astro-ph.GA

Shared hidden-factor information framework for multiple behavioral tasks

Understanding cognitive processes in major depressive disorder (MDD) often relies on behavioral tasks, which are typically analyzed separately, overlooking potential correlations and shared latent structure. To address this limitation, we propose the Shared Hidden-factor Information Framework for Multiple Behavioral Tasks (SHIFT), a joint modeling approach that leverages shared information across tasks, allowing each task to benefit from information learned by the others. SHIFT introduces subject-specific latent factors that capture cross-task dependencies while accommodating individual heterogeneity in decision-making, response times (RTs), and strategy switching. To address computational challenges without requiring high-dimensional integration, we develop an expectation-maximization with variational approximation algorithm that preserves both temporal structure and between-task dependencies. Through extensive simulation studies, we demonstrate that SHIFT substantially improves estimation accuracy and efficiency relative to single-task analyses. We then apply SHIFT to a study of MDD to jointly model the Probabilistic Reward Task (PRT) and the Flanker Task (FT). Results indicate that MDD participants show lower engagement in the PRT and reduced focus in the FT compared with healthy controls. Moreover, when individuals are engaged and focused, they exhibit longer RTs. Although observed RTs do not predict treatment response, the shared parameters recovered by SHIFT showed suggestive treatment-modulation patterns, indicating their potential as exploratory behavioral markers for therapeutic outcomes.

stat.ME

Integrative learning of individualized treatment rules from multiple studies with partially overlapping treatments

An individualized treatment rule (ITR) tailors treatments to a patient's specific characteristics. However, randomized controlled trials (RCTs) are often underpowered to detect the treatment effect heterogeneity needed for reliable ITR estimation. To address this limitation, there is growing interest in leveraging information from multiple studies to improve statistical power and support individualized decision-making. A key challenge in this context is that available RCTs may not evaluate the same set of treatments. In this paper, we propose an integrative learning framework that synthesizes evidence across multiple RCTs that share a common comparator but differ in their alternative treatment arms. Our method integrates information through a regularized weighted misclassification risk function and adaptively determines the contribution of each study to the ITRs of the others. We rigorously study the excess risk of the resulting estimator. Simulation studies demonstrate that the proposed approaches improve the estimation of both value and benefit functions. We illustrate the utility of our methodology using data from two landmark studies of major depressive disorder: the Establishing Moderators and Biosignatures of Antidepressant Response in Clinical Care study and the International Study to Predict Optimized Treatment in Depression study, both of which include a selective serotonin reuptake inhibitor as a common treatment arm. We find that the separate learning method outperforms one-size-fits-all methods, and our integrative methods further improve performance.

stat.ME

Sample size and power determination for assessing overall SNP effects in joint modeling of longitudinal and time-to-event data

Longitudinal biomarkers are frequently collected in clinical studies due to their strong association with time-to-event outcomes. While considerable progress has been made in methods for jointly modeling longitudinal and survival data, comparatively little attention has been paid to statistical design considerations, particularly sample size and power calculations, in genetic studies. Yet, appropriate sample size estimation is essential for ensuring adequate power and valid inference. Genetic variants may influence event risk through both direct effects and indirect effects mediated by longitudinal biomarkers. In this paper, we derive a closed-form sample size formula for testing the overall effect of a single nucleotide polymorphism within a joint modeling framework. Simulation studies demonstrate that the proposed formula yields accurate and robust performance in finite samples. We illustrate the practical utility of our method using data from the Diabetes Control and Complications Trial.

stat.ME

Boosting methods for interval-censored data with regression and classification

Boosting has garnered significant interest across both machine learning and statistical communities. Traditional boosting algorithms, designed for fully observed random samples, often struggle with real-world problems, particularly with interval-censored data. This type of data is common in survival analysis and time-to-event studies where exact event times are unobserved but fall within known intervals. Effective handling of such data is crucial in fields like medical research, reliability engineering, and social sciences. In this work, we introduce novel nonparametric boosting methods for regression and classification tasks with interval-censored data. Our approaches leverage censoring unbiased transformations to adjust loss functions and impute transformed responses while maintaining model accuracy. Implemented via functional gradient descent, these methods ensure scalability and adaptability. We rigorously establish their theoretical properties, including optimality and mean squared error trade-offs. Our proposed methods not only offer a robust framework for enhancing predictive accuracy in domains where interval-censored data are common but also complement existing work, expanding the applicability of existing boosting techniques. Empirical studies demonstrate robust performance across various finite-sample scenarios, highlighting the practical utility of our approaches.

stat.ML

Cold Gas Infall onto A Brightest Group Galaxy via A Gas-Rich Minor Merger

Dust and cold gas are not uncommon in nearby early-type galaxies (ETGs), and represent an important aspect of their evolution. However, their origin has been debated for decades. Potential sources include internal processes (e.g., mass loss from evolved stars), external mechanisms (e.g., minor mergers or cooling flows), or a combination of both. Gas-rich minor mergers have long been proposed as an important channel for cold gas fueling in both observations and simulations, but direct evidence of cold gas transportation via gas-rich minor mergers remains elusive, particularly in galaxy groups and clusters where environmental effects are prevalent. In this letter, we present the first unambiguous case of direct cold gas transportation onto a brightest group galaxy (BGG) at $z=0.25$, driven by an ongoing close-separation gas-rich minor merger with a mass ratio of $\sim1:56$. High-resolution JWST imaging reveals a heavily obscured, low-mass satellite that is barely visible at restframe optical wavelengths. Tidal stripping from this satellite deposits gas and dust onto the BGG, forming prominent $\sim$10 kpc dust lanes in situ. Cosmological simulations indicate that such interactions preferentially occur in gas-rich satellites undergoing their first infall in highly eccentric orbits. Our results highlight the pivotal role of gas-rich minor mergers in replenishing cold gas reservoirs and shaping the evolution of central ETGs in galaxy groups.

astro-ph.GA

Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding

Monocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two key limitations. Firstly, they often over-rely on high-certainty keywords that explicitly identify the target object while neglecting critical spatial descriptions. Secondly, generalized textual features contain both 2D and 3D descriptive information, thereby capturing an additional dimension of details compared to singular 2D or 3D visual features. This characteristic leads to cross-dimensional interference when refining visual features under text guidance. To overcome these challenges, we propose Mono3DVG-EnSD, a novel framework that integrates two key components: the CLIP-Guided Lexical Certainty Adapter (CLIP-LCA) and the Dimension-Decoupled Module (D2M). The CLIP-LCA dynamically masks high-certainty keywords while retaining low-certainty implicit spatial descriptions, thereby forcing the model to develop a deeper understanding of spatial relationships in captions for object localization. Meanwhile, the D2M decouples dimension-specific (2D/3D) textual features from generalized textual features to guide corresponding visual features at same dimension, which mitigates cross-dimensional interference by ensuring dimensionally-consistent cross-modal interactions. Through comprehensive comparisons and ablation studies on the Mono3DRefer dataset, our method achieves state-of-the-art (SOTA) performance across all metrics. Notably, it improves the challenging Far(Acc@0.5) scenario by a significant +13.54%.

cs.CV

A Dual-stage Prompt-driven Privacy-preserving Paradigm for Person Re-Identification

With growing concerns over data privacy, researchers have started using virtual data as an alternative to sensitive real-world images for training person re-identification (Re-ID) models. However, existing virtual datasets produced by game engines still face challenges such as complex construction and poor domain generalization, making them difficult to apply in real scenarios. To address these challenges, we propose a Dual-stage Prompt-driven Privacy-preserving Paradigm (DPPP). In the first stage, we generate rich prompts incorporating multi-dimensional attributes such as pedestrian appearance, illumination, and viewpoint that drive the diffusion model to synthesize diverse data end-to-end, building a large-scale virtual dataset named GenePerson with 130,519 images of 6,641 identities. In the second stage, we propose a Prompt-driven Disentanglement Mechanism (PDM) to learn domain-invariant generalization features. With the aid of contrastive learning, we employ two textual inversion networks to map images into pseudo-words representing style and content, respectively, thereby constructing style-disentangled content prompts to guide the model in learning domain-invariant content features at the image level. Experiments demonstrate that models trained on GenePerson with PDM achieve state-of-the-art generalization performance, surpassing those on popular real and virtual Re-ID datasets.

cs.CV

Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability

Multimodal industrial surface defect detection (MISDD) aims to identify and locate defect in industrial products by fusing RGB and 3D modalities. This article focuses on modality-missing problems caused by uncertain sensors availability in MISDD. In this context, the fusion of multiple modalities encounters several troubles, including learning mode transformation and information vacancy. To this end, we first propose cross-modal prompt learning, which includes: i) the cross-modal consistency prompt serves the establishment of information consistency of dual visual modalities; ii) the modality-specific prompt is inserted to adapt different input patterns; iii) the missing-aware prompt is attached to compensate for the information vacancy caused by dynamic modalities-missing. In addition, we propose symmetric contrastive learning, which utilizes text modality as a bridge for fusion of dual vision modalities. Specifically, a paired antithetical text prompt is designed to generate binary text semantics, and triple-modal contrastive pre-training is offered to accomplish multimodal learning. Experiment results show that our proposed method achieves 73.83% I-AUROC and 93.05% P-AUROC with a total missing rate 0.7 for RGB and 3D modalities (exceeding state-of-the-art methods 3.84% and 5.58% respectively), and outperforms existing approaches to varying degrees under different missing types and rates. The source code will be available at https://github.com/SvyJ/MISDD-MM.

cs.CV

Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding

Monocular 3D visual grounding is a novel task that aims to locate 3D objects in RGB images using text descriptions with explicit geometry information. Despite the inclusion of geometry details in the text, we observe that the text embeddings are sensitive to the magnitude of numerical values but largely ignore the associated measurement units. For example, simply equidistant mapping the length with unit "meter" to "decimeters" or "centimeters" leads to severe performance degradation, even though the physical length remains equivalent. This observation signifies the weak 3D comprehension of pre-trained language model, which generates misguiding text features to hinder 3D perception. Therefore, we propose to enhance the 3D perception of model on text embeddings and geometry features with two simple and effective methods. Firstly, we introduce a pre-processing method named 3D-text Enhancement (3DTE), which enhances the comprehension of mapping relationships between different units by augmenting the diversity of distance descriptors in text queries. Next, we propose a Text-Guided Geometry Enhancement (TGE) module to further enhance the 3D-text information by projecting the basic text features into geometrically consistent space. These 3D-enhanced text features are then leveraged to precisely guide the attention of geometry features. We evaluate the proposed method through extensive comparisons and ablation studies on the Mono3DRefer dataset. Experimental results demonstrate substantial improvements over previous methods, achieving new state-of-the-art results with a notable accuracy gain of 11.94\% in the "Far" scenario. Our code will be made publicly available.

cs.CV

Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion

Person re-identification (re-id) models are vital in security surveillance systems, requiring transferable adversarial attacks to explore the vulnerabilities of them. Recently, vision-language models (VLM) based attacks have shown superior transferability by attacking generalized image and textual features of VLM, but they lack comprehensive feature disruption due to the overemphasis on discriminative semantics in integral representation. In this paper, we introduce the Attribute-aware Prompt Attack (AP-Attack), a novel method that leverages VLM's image-text alignment capability to explicitly disrupt fine-grained semantic features of pedestrian images by destroying attribute-specific textual embeddings. To obtain personalized textual descriptions for individual attributes, textual inversion networks are designed to map pedestrian images to pseudo tokens that represent semantic embeddings, trained in the contrastive learning manner with images and a predefined prompt template that explicitly describes the pedestrian attributes. Inverted benign and adversarial fine-grained textual semantics facilitate attacker in effectively conducting thorough disruptions, enhancing the transferability of adversarial examples. Extensive experiments show that AP-Attack achieves state-of-the-art transferability, significantly outperforming previous methods by 22.9% on mean Drop Rate in cross-model&dataset attack scenarios.

cs.CV

Joint modeling for learning decision-making dynamics in behavioral experiments

Major depressive disorder (MDD), a leading cause of disability and mortality, is associated with reward-processing abnormalities and concentration issues. Motivated by the probabilistic reward task from the Establishing Moderators and Biosignatures of Antidepressant Response in Clinical Care (EMBARC) study, we propose a novel framework that integrates the reinforcement learning (RL) model and drift-diffusion model (DDM) to jointly analyze reward-based decision-making with response times. To account for emerging evidence suggesting that decision-making may alternate between multiple interleaved strategies, we model latent state switching using a hidden Markov model (HMM). In the ''engaged'' state, decisions follow an RL-DDM, simultaneously capturing reward processing, decision dynamics, and temporal structure. In contrast, in the ''lapsed'' state, decision-making is modeled using a simplified DDM, where specific parameters are fixed to approximate random guessing with equal probability. The proposed method is implemented using a computationally efficient generalized expectation-maximization (EM) algorithm with forward-backward procedures. Through extensive numerical studies, we demonstrate that our proposed method outperforms competing approaches across various reward-generating distributions, under both strategy-switching and non-switching scenarios, as well as in the presence of input perturbations. When applied to the EMBARC study, our framework reveals that MDD patients exhibit lower overall engagement than healthy controls and experience longer decision times when they do engage. Additionally, we show that neuroimaging measures of brain activities are associated with decision-making characteristics in the ''engaged'' state but not in the ''lapsed'' state, providing evidence of brain-behavior association specific to the ''engaged'' state.

stat.ME

Modality Unified Attack for Omni-Modality Person Re-Identification

Deep learning based person re-identification (re-id) models have been widely employed in surveillance systems. Recent studies have demonstrated that black-box single-modality and cross-modality re-id models are vulnerable to adversarial examples (AEs), leaving the robustness of multi-modality re-id models unexplored. Due to the lack of knowledge about the specific type of model deployed in the target black-box surveillance system, we aim to generate modality unified AEs for omni-modality (single-, cross- and multi-modality) re-id models. Specifically, we propose a novel Modality Unified Attack method to train modality-specific adversarial generators to generate AEs that effectively attack different omni-modality models. A multi-modality model is adopted as the surrogate model, wherein the features of each modality are perturbed by metric disruption loss before fusion. To collapse the common features of omni-modality models, Cross Modality Simulated Disruption approach is introduced to mimic the cross-modality feature embeddings by intentionally feeding images to non-corresponding modality-specific subnetworks of the surrogate model. Moreover, Multi Modality Collaborative Disruption strategy is devised to facilitate the attacker to comprehensively corrupt the informative content of person images by leveraging a multi modality feature collaborative metric disruption loss. Extensive experiments show that our MUA method can effectively attack the omni-modality re-id models, achieving 55.9%, 24.4%, 49.0% and 62.7% mean mAP Drop Rate, respectively.

cs.CV

Learning to Learn Transferable Generative Attack for Person Re-Identification

Deep learning-based person re-identification (re-id) models are widely employed in surveillance systems and inevitably inherit the vulnerability of deep networks to adversarial attacks. Existing attacks merely consider cross-dataset and cross-model transferability, ignoring the cross-test capability to perturb models trained in different domains. To powerfully examine the robustness of real-world re-id models, the Meta Transferable Generative Attack (MTGA) method is proposed, which adopts meta-learning optimization to promote the generative attacker producing highly transferable adversarial examples by learning comprehensively simulated transfer-based cross-model\&dataset\&test black-box meta attack tasks. Specifically, cross-model\&dataset black-box attack tasks are first mimicked by selecting different re-id models and datasets for meta-train and meta-test attack processes. As different models may focus on different feature regions, the Perturbation Random Erasing module is further devised to prevent the attacker from learning to only corrupt model-specific features. To boost the attacker learning to possess cross-test transferability, the Normalization Mix strategy is introduced to imitate diverse feature embedding spaces by mixing multi-domain statistics of target models. Extensive experiments show the superiority of MTGA, especially in cross-model\&dataset and cross-model\&dataset\&test attacks, our MTGA outperforms the SOTA methods by 21.5\% and 11.3\% on mean mAP drop rate, respectively. The code of MTGA will be released after the paper is accepted.

cs.CV

Boosting prediction with data missing not at random

Boosting has emerged as a useful machine learning technique over the past three decades, attracting increased attention. Most advancements in this area, however, have primarily focused on numerical implementation procedures, often lacking rigorous theoretical justifications. Moreover, these approaches are generally designed for datasets with fully observed data, and their validity can be compromised by the presence of missing observations. In this paper, we employ semiparametric estimation approaches to develop boosting prediction methods for data with missing responses. We explore two strategies for adjusting the loss functions to account for missingness effects. The proposed methods are implemented using a functional gradient descent algorithm, and their theoretical properties, including algorithm convergence and estimator consistency, are rigorously established. Numerical studies demonstrate that the proposed methods perform well in finite sample settings.

stat.ME

Two Channels of Metal-Rich Compact Stellar System Formation: Starbursts under High Ram Pressure versus Tidal Stripping

Most galaxies follow well-defined scaling relations of metallicity and stellar mass; however, some outliers at the low mass end of the observed galaxy population exhibit unusually high metallicity for their mass. Understanding how these objects get to be so metal-rich is vital for understanding the role of feedback in galaxy formation. Using the TNG50 simulation, we explore the origins of this phenomenon. We identify 227 metal-rich, compact stellar systems (CSSs) that deviate significantly from this scaling relation. These CSSs are satellites located in the vicinity of massive host galaxies, with stellar masses ranging from $10^{8} M_{\odot}$ to $10^{10}\ M_{\odot}$ (including six systems that are close analogs of the M31-M32 system). Contrary to the previously assumed scenario that such objects are predominantly products of tidal stripping, our results suggest a more prevalent role for ram pressure in their formation. Indeed, 76% (173) of these CSSs are formed through a burst of star formation occurring around the time of the first pericentric passage, typically at redshifts $z\lesssim1$, aided by strong ram pressure and tidal forces. The high ram pressure, resulting from the CSSs' rapid motion near the halo center, facilitates metal enrichment, producing high-metallicity CSSs by confining the metal-rich gas from bursty star formation, which leads to distinct stellar populations characterized by enhanced metallicity as well as high $α$-abundance. Only the remaining 24% (54) of metal-rich CSSs are generated through the tidal stripping of massive progenitors. Our results further indicate that M32 is more likely to have formed through intense star formation events rather than through gradual, tidal stripping, thereby providing crucial insights into the nature of low mass, compact galaxy formation.

astro-ph.GA