SearcharxivSearch

arXiv subjects

Zhao Chen

Publications and source records attributed to Zhao Chen.

At least 19 recordsLinked to original sources

Improved Cosmological Constraints from Morphology-Based Marked Correlation Functions

The cosmic web contains morphology-dependent information that is not fully captured by standard two-point statistics. We construct morphology-based marked correlation functions (MCFs) by assigning marks to halos according to the cosmic-web morphology identified with the \textsc{Nexus} algorithm. Using the \textsc{Kun} simulation suite, which spans 129 $w_0w_a$CDM cosmologies, we build Gaussian-process emulators for the MCFs as functions of cosmological parameters and tracer bias. We then apply the emulators to mock halo catalogues from the independent \textsc{Jiutian} simulation and perform a joint likelihood analysis to quantify the resulting cosmological constraints. We consider two marker choices: a discrete morphology marker and a continuous morphology strength marker. The continuous marker improves the Figure of Merit (FoM) by a factor of $\sim 8.6$ relative to the standard 2PCF and reduces the $1\sigma$ uncertainty on $\sigma_8$ by a factor of $\sim 5$. The discrete marker gives a more modest FoM improvement of $\sim 17\%$. We further test the impact of tracer selection by varying the halo mass threshold by a factor of $\sim 4.5$. Even for the lowest mass threshold, the continuous marker remains unbiased and achieves a FoM about $\sim 3.4$ times higher than that of the 2PCF alone. These results show that morphology-based MCFs, combined with simulation-based emulation, provide a useful framework for extracting additional cosmological information from large-scale structure surveys.

astro-ph.CO

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. The core insight is to select a subset of samples whose empirical distribution maximizes its Wasserstein distance to the fully contaminated empirical distribution, thereby preferentially isolating and removing geometrically influential outliers. To render this optimization computationally tractable, we introduce three algorithms: a marginal screening scheme, SinkMarg, and two joint optimization algorithms, SinkWF and SlicedWF, leveraging entropic optimal transport and sliced Wasserstein approximations, respectively. On the theoretical front, we introduce the Far Exclusion and Local Projection (FELP) contamination model, which characterizes corruptions consisting of well-separated outliers and locally indistinguishable perturbations. Under this model, we prove that the WF estimator achieves minimax optimality over distribution families with bounded covariance. Extensive numerical experiments on synthetic datasets, benchmark anomaly detection suites, and robust generative learning with diffusion models demonstrate that WF serves as a highly practical, model-agnostic preprocessing tool. It delivers competitive outlier detection performance and provides substantial downstream benefits for generative modeling under heavy contamination.

stat.ML

Evaluating Language Models in Realistic Conversational Contexts

As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summarization, translation, or short-form QA tasks fall short of adequately measuring the consistency of human-scale dialogue, especially when derivation and validation of these metrics themselves often rely on synthetic rather than human sources. We fill the gap by introducing UPHELD (UPwork Human-Scale Evaluated Long Dialogues), a large, reference-full benchmark for evaluating human-scale conversational ability beyond factual correctness. UPHELD consists of hundreds of complete human-to-human dialogues authored by professional script writers, with realistic turn densities and 36,000+ per-turn human annotations across 30,000+ expert-generated dialogue turns. Using UPHELD, we systematically evaluate classical automatic metrics and reference-free LLM-as-a-judge approaches, and find them unreliable when correlated with expert human judgment. Building off this analysis, we use UPHELD to develop a Mixture-of-Judges framework that combines multiple evaluative signals and improves correlation with human assessments by approximately 30%. Overall, UPHELD provides a robust, human-grounded foundation for evaluating human-scale conversational intelligence that fills a crucial gap in the pre-existing LLM dataset landscape.

cs.CL

Cosmological Constraints from Bias-Robust Wavelet Scattering Statistics for Stage-IV Galaxy Surveys

A central challenge in precision cosmology with galaxy surveys is to extract non-Gaussian information from large-scale structure while controlling systematic uncertainties such as tracer bias. Conventional clustering statistics, such as the two-point correlation function (2PCF), capture limited nonlinear information and typically require explicit bias modeling, which can introduce systematic errors if the adopted bias prescription is inaccurate. To address this problem, we introduce $R^{\rm wst}$, a bias-robust statistic constructed from $m$-mode ratios of the wavelet scattering transform (WST). Using simulation-based inference, we train a Gaussian-process-regression emulator on the \texttt{Kun} simulation suite and use \texttt{JiuTian} simulations for covariance estimation and validation. The emulator achieves percent-level accuracy, sufficient for the expected observational uncertainties. We show that $R^{\rm wst}$ yields unbiased constraints on $\Omega_m$, $\sigma_8$, $n_s$, and $w_0$, and improves the breaking of the $\Omega_m$--$\sigma_8$ degeneracy by about a factor of two compared with 2PCF. Its constraining power remains stable across a broad range of tracer-bias scenarios, demonstrating that $R^{\rm wst}$ can mitigate bias-induced systematics without explicit bias modeling. These results establish $R^{\rm wst}$ as a powerful and robust statistic for precision cosmology with Stage-IV surveys.

astro-ph.CO

Cosmological constraints from neighbor-density-weighted marked correlation functions

We investigate whether neighbor-density-weighted marked correlation functions (MCFs) can extract cosmological information beyond the standard redshift-space two-point correlation function (2PCF). Using the Kun suite of 129 $w_0w_a$CDM$+\sum m_\nu$ simulations in $1~h^{-1}{\rm Gpc}$ boxes, we construct Gaussian-process emulators for the normalized scale statistic $\widehat{W}^{\alpha}(s)$ and the angular statistic $\widehat{W}^{\alpha}_{\Delta s}(\mu)$. We perform joint analyses combining multiple mark parameters $\alpha$ and quantify the information gain using the FoM in the $\Omega_m$--$\sigma_8$ plane. Relative to the 2PCF case, three-mark combinations improve the FoM by factors of $1.7$--$2.5$, while five-mark combinations increase the gain to $1.9$--$2.4$, depending on the statistic and mark definition. We further compare density and normalized-gradient marks, finding that they are nearly redundant for isotropic statistics but complementary for angular statistics, where their combination improves the FoM by up to $43\%$. Tests of scale range and halo selection show that the marked statistics remain robust under changes in analysis choices, with the angular statistic retaining additional cosmological information that is less sensitive to tracer selection. Our results demonstrate that MCFs substantially enhance cosmological constraints beyond the standard 2PCF and provide a robust probe for next-generation galaxy surveys.

astro-ph.CO

Factor-Adjusted Multiple Testing for High-Dimensional Individual Mediation Effects

Identifying individual mediators is a central goal of high-dimensional mediation analysis, yet pervasive dependence among mediators can invalidate standard debiased inference and lead to substantial false discovery rate (FDR) inflation. We propose a Factor-Adjusted Debiased Mediation Testing (FADMT) framework that enables large-scale inference for individual mediation effects with FDR control under complex dependence structures. Our approach posits an approximate factor structure on the unobserved errors of the mediator model, extracts common latent factors, and constructs decorrelated pseudo-mediators for the subsequent inferential procedure. We establish the asymptotic normality of the debiased estimator and develop a multiple testing procedure with theoretical FDR control under mild high-dimensional conditions. By adjusting for latent factor induced dependence, FADMT also improves robustness to spurious associations driven by shared latent variation in observational studies. Extensive simulations demonstrate the superior finite-sample performance across a wide range of correlation structures. Applications to TCGA-BRCA multi-omics data and to China's stock connect study further illustrate the practical utility of the proposed method.

stat.ME

ELUCID-DESI I: A Parallel MPI Implementation of the Initial Condition Solver for Large-Scale Reconstruction Simulations

We present a highly scalable, MPI-parallelized framework for reconstructing the initial cosmic density field, designed to meet the computational demands of next-generation cosmological simulations, particularly the upcoming ELUCID-DESI simulation based on DESI BGS data. Building upon the Hamiltonian Monte Carlo approach and the FastPM solver, our code employs domain decomposition to efficiently distribute memory between nodes. Although communication overhead increases the per-step runtime of the MPI version by roughly a factor of eight relative to the shared-memory implementation, our scaling tests-spanning different particle numbers, core counts, and node layouts-show nearly linear scaling with respect to both the number of particles and the number of CPU cores. Furthermore, to significantly reduce computational costs during the initial burn-in phase, we introduce a novel ``guess'' module that rapidly generates a high-quality initial density field. The results of the simulation test confirm substantial efficiency gains: for $256^3$ particles, 53 steps ($\sim$ 54 core hours) are saved, accelerating convergence by a factor of $\sim$ 18; for $1024^3$, 106 steps ($\sim$7500 core hours), achieving a speedup factor of $\sim$ 3. The total core hour gain grows with the number of particles, rendering large-volume reconstructions computationally practical for upcoming surveys, including our planned ELUCID-DESI reconstruction simulation with $4096^3$ particles. We estimate that achieving convergence for this scale (targeting DESI-BGS data) requires about 800 HMCMC steps ($\sim$ 5 million core hours). Our initial guess module will save approximately 360 steps ($\sim$2.3 million core hours), reducing the total computational time by about 45\%.

astro-ph.GA

CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis

High-quality annotated datasets are crucial for advancing machine learning in medical image analysis. However, a critical gap exists: most datasets either offer a single, clean ground truth, which hides real-world expert disagreement, or they provide multiple annotations without a separate gold standard for objective evaluation. To bridge this gap, we introduce CytoCrowd, a new public benchmark for cytology analysis. The dataset features 446 high-resolution images, each with two key components: (1) raw, conflicting annotations from four independent pathologists, and (2) a separate, high-quality gold-standard ground truth established by a senior expert. This dual structure makes CytoCrowd a versatile resource. It serves as a benchmark for standard computer vision tasks, such as object detection and classification, using the ground truth. Simultaneously, it provides a realistic testbed for evaluating annotation aggregation algorithms that must resolve expert disagreements. We provide comprehensive baseline results for both tasks. Our experiments demonstrate the challenges presented by CytoCrowd and establish its value as a resource for developing the next generation of models for medical image analysis.

cs.CV

MIU2Net: weak-lensing mass inversion using deep learning with nested U-structures

One of the primary goals of next-generation gravitational lensing surveys is to measure the large-scale distribution of dark matter, which requires accurate mass inversion to convert weak-lensing shear maps into convergence (kappa) fields. This work develops a mass inversion method tailored for upcoming space missions such as CSST and Euclid, aiming to recover both the mass distribution and the convergence power spectrum with high fidelity. We introduce MIU2Net, a versatile deep-learning framework for kappa-map reconstruction based on the U2-Net architecture. A new loss function is constructed to jointly estimate the convergence field and its frequency-domain energy distribution, effectively balancing optimal mean squared error and optimal power-spectrum recovery. The method incorporates realistic observational effects into shear fields, including shape noise, reduced shear, and complex masks. Under noise levels anticipated for future space-based lensing surveys, MIU2Net recovers the convergence power spectrum with 4% uncertainties up to l approximately 500, significantly outperforming Wiener filtering and MCALens. Beyond two-point statistics, the method accurately reconstructs the convergence distribution, peak centroid, and peak amplitude. Compared to other learning-based approaches such as DeepMass, MIU2Net reduces the root-mean-square error by 5% without smoothing and by 38% with a 1-arcmin smoothing scale. MIU2Net represents a substantial advancement in mass inversion methodology, offering improved accuracy in both RMSE and power-spectrum reconstruction. It provides a promising tool for mapping dark matter environments and large-scale structures in the era of next-generation space lensing surveys.

astro-ph.CO

Goodness-of-fit Tests for Heavy-tailed Random Fields

We develop goodness-of-fit tests for max-stable random fields, which are used to model heavy-tailed spatial data. The test statistics are constructed based on the Fourier transforms of the indicators of extreme values in the heavy-tailed spatial data, whose asymptotic distribution is a Gaussian random field under a hypothesized max-stable random field. Since the covariance structure of the limiting Gaussian random field lacks an explicit expression, we propose a stationary bootstrap procedure for spatial fields to approximate critical values. Simulation studies confirm the theoretical distributional results, and applications to PM2.5 and temperature data illustrate the practical utility of the proposed method for model assessment.

stat.ME

The first AKRA mass map reconstruction from HSC Y1 data

Weak lensing mass-mapping from shear catalogs faces systematic challenges from survey masks and spatially varying noise. To overcome these issues and reconstruct unbiased convergence $\kappa$ maps, we have constructed the AKRA (Accurate Kappa Reconstruction Algorithm), a prior-free and maximum-likelihood based analytical method. It has been validated for mock shear catalogs with a variety of survey masks. In this work, we present the first real-data application of the AKRA on the Subaru Hyper Suprime-Cam Year 1 (HSC Y1) data. We first validate AKRA using mock shear catalogs from the \texttt{Kun} simulation suite, with masks corresponding to the six HSC Y1 regions (\texttt{GAMA09H}, \texttt{GAMA15H}, \texttt{HECTOMAP}, \texttt{VVDS}, \texttt{WIDE12H}, and \texttt{XMMLSS}). The investigated statistics, including the lensing power spectrum, $\langle \kappa^2\rangle$, $\langle \kappa^3\rangle$, and the one-point probability distribution function of $\kappa$, are all unbiased. We then apply AKRA to the HSC Y1 shear catalog and provide reconstructed $\kappa$ maps ready for subsequent scientific analyses.

astro-ph.CO

Extending CSST Emulator to post-DESI era

The recent DESI BAO measurements have revealed a potential deviation from a cosmological constant, suggesting a dynamic nature of dark energy. To rigorously test this result, complementary probes such as weak gravitational lensing are crucial, demanding highly accurate and efficient predictions of the nonlinear matter power spectrum within the $w_0w_a$CDM framework. However, most existing emulators fail to cover the full parameter posterior from DESI DR2+CMB constraints in the $w_0\mbox{-}w_a$ plane. In this work, we extend the spectral equivalence method outlined in Casarini et al. 2016 to use auxiliary $w_0w_a$CDM models for approximating the power spectrum of a target $w_0w_a$CDM cosmology, moving beyond the previous use of $w$CDM auxiliaries. Incorporating this enhanced module, the extended CSST Emulator achieves a prediction accuracy of $\leq1\%$ over the $1\sigma$ confidence region from DESI DR2+CMB constraints for $z\leq3$, with a mild degradation in accuracy outside this posterior region. This performance is rigorously validated by additional simulations of dynamic dark energy cosmologies. The emulator's applicable parameter space has been generalized to fully encompass the $2\sigma$ region, greatly enhancing its utility for cosmological analysis in the post-DESI era.

astro-ph.CO

SCAR: A Characterization Scheme for Multi-Modal Dataset

Foundation models exhibit remarkable generalization across diverse tasks, largely driven by the characteristics of their training data. Recent data-centric methods like pruning and compression aim to optimize training but offer limited theoretical insight into how data properties affect generalization, especially the data characteristics in sample scaling. Traditional perspectives further constrain progress by focusing predominantly on data quantity and training efficiency, often overlooking structural aspects of data quality. In this study, we introduce SCAR, a principled scheme for characterizing the intrinsic structural properties of datasets across four key measures: Scale, Coverage, Authenticity, and Richness. Unlike prior data-centric measures, SCAR captures stable characteristics that remain invariant under dataset scaling, providing a robust and general foundation for data understanding. Leveraging these structural properties, we introduce Foundation Data-a minimal subset that preserves the generalization behavior of the full dataset without requiring model-specific retraining. We model single-modality tasks as step functions and estimate the distribution of the foundation data size to capture step-wise generalization bias across modalities in the target multi-modal dataset. Finally, we develop a SCAR-guided data completion strategy based on this generalization bias, which enables efficient, modality-aware expansion of modality-specific characteristics in multimodal datasets. Experiments across diverse multi-modal datasets and model architectures validate the effectiveness of SCAR in predicting data utility and guiding data acquisition. Code is available at https://github.com/McAloma/SCAR.

cs.LG

Artifacts in Halo Shapes: Imprints of the Initial Condition

Grid type pre-initial conditions are commonly used to initialize particle positions in cosmological simulations. While these conditions are known to produce noticeable numerical artifacts in void regions, their impact on halo properties has generally been assumed to be negligible. In this work, we employ multiple simulations to demonstrate that grid initialization induces statistically significant artifacts in halo shapes, despite the modest absolute amplitude ($\sim 1\%$) making them unimportant for most cosmological studies. We identify a redshift-dependent artificial alignment pattern: at low redshifts ($z<2$), halo shapes preferentially orient away from the simulation box's Cartesian axes, whereas their constituent particles initially exhibit alignment with these axes. We propose a mathematical hypothesis to explain this flipping behavior.

astro-ph.CO

Wireless AI Evolution: From Statistical Learners to Electromagnetic-Guided Foundation Models

While initial applications of artificial intelligence (AI) in wireless communications over the past decade have demonstrated considerable potential using specialized models for targeted communication tasks, the revolutionary demands of sixth-generation (6G) networks for holographic communications, ubiquitous sensing, and native intelligence are propelling a necessary evolution towards AI-native wireless networks. The arrival of large AI models paves the way for the next phase of Wireless AI, driven by wireless foundation models (WFMs). In particular, pre-training on universal electromagnetic (EM) principles equips WFMs with the essential adaptability for a multitude of demanding 6G applications. However, existing large AI models face critical limitations, including pre-training strategies disconnected from EM-compliant constraints leading to physically inconsistent predictions, a lack of embedded understanding of wave propagation physics, and the inaccessibility of massive labeled datasets for comprehensive EM-aware training. To address these challenges, this article presents an electromagnetic information theory-guided self-supervised pre-training (EIT-SPT) framework designed to systematically inject EM physics into WFMs. The EIT-SPT framework aims to infuse WFMs with intrinsic EM knowledge, thereby enhancing their physical consistency, generalization capabilities across varied EM landscapes, and overall data efficiency. Building upon the proposed EIT-SPT framework, this article first elaborates on diverse potential applications in 6G scenarios of WFMs, then validates the efficacy of the proposed framework through illustrative case studies, and finally summarizes critical open research challenges and future directions for WFMs.

cs.IT

CSST Cosmological Emulator II: Generalized Accurate Halo Mass Function Emulation

Accurate theoretical prediction for halo mass function across a broad cosmological space is crucial for the forthcoming China Space Station Telescope (CSST) observations, which will capture cosmological information from multiple probes, e.g., cluster abundance, and weak lensing. In this work, we quantify the percent-level impact of different mass binning schemes when measuring the differential halo mass function from simulations, and demonstrate that the cumulative form of the halo mass function is independent of the binning scheme. Through the recently finished Kun simulation suite, we propose a generalized framework to construct multiple accurate halo mass function emulators for different halo mass definitions, including $M_{200m}$, $M_{vir}$, and $M_{200c}$. This extends our CSST Emulator to provide fast and accurate halo mass function predictions for halo mass $M\geq 10^{12}\,h^{-1}M_{\odot}$ up to $z=3.0$. For redshifts $z\leq 1.0$, the accuracy is within $2\%$ for $M\leq 10^{13}\,h^{-1}M_{\odot}$, $5\%$ for $M\leq 10^{14}\,h^{-1}M_{\odot}$, and $10\%$ for $M\leq 10^{15}\,h^{-1}M_{\odot}$, which is comparable with the statistical errors of training simulations. This tool is integrated in CSST Emulator and publicly available at https://github.com/czymh/csstemu, providing a fast and accurate theoretical tool to obtain unbiased cosmological constraints of the upcoming CSST survey.

astro-ph.CO

Uni2D: A Universal Machine Learning Interatomic Potential for Two-Dimensional Materials

Accurate interatomic potentials (IAPs) are essential for modeling the potential energy surfaces (PES) that govern atomic interactions in materials. However, most existing IAPs are developed for bulk materials and often struggle to accurately and efficiently capture the diverse chemical environments of two-dimensional (2D) materials, which limits large-scale simulation and design of emerging 2D systems. To address this challenge, we develop Uni2D, an interatomic potential tailored for 2D materials. The Uni2D model is trained on a dataset comprising approximately 327,000 structure-energy-force-stress mappings derived from about 20,000 distinct 2D materials, covering 89 chemical elements. The model demonstrates reliable predictive performance for energies, forces, and stresses, and demonstrates quantitatively robust accuracy in tasks such as structural relaxation, equation-of-state calculations, and molecular dynamics simulations, making the model suitable for high-throughput screening of 2D materials. For derived properties, including elastic properties, lattice dynamics, and other screening-related metrics, the model provides qualitative to semi-quantitative predictions that remain useful for trend analysis and preliminary evaluation. To enhance usability, we further introduce an intelligent agent powered by a large language model (LLM), enabling automated workflows and natural language interaction for 2D materials simulations. Our work provides an efficient and accessible framework for high-throughput screening and computational exploration of 2D materials.

cond-mat.mtrl-sci

Asymptotic Theory for Regularized Estimation in Functional Time Series Models

Functional autoregressive (FAR) models provide a fundamental framework for analyzing temporally dependent functional data. However, the infinite-dimensional nature of the underlying Hilbert space introduces intrinsic ill-posedness, as the autocovariance operators are compact and lack bounded inverses. This paper develops a new theoretical framework for the regularized estimation and asymptotic analysis of FAR models. Leveraging Hilbert space theory, we rigorously characterize the distinction between finite- and infinite-dimensional time series analysis and formalize the necessity of regularization. To stabilize the estimation of autoregressive operators, we introduce a Tikhonov regularization scheme and derive Yule-Walker-type estimators in a general Hilbert space, and further specialize to the $L^2$ space for explicit forms. Within this unified framework, we establish the consistency and asymptotic normality of the regularized estimators and reveal that asymptotic normality can be achieved only for the predictors rather than the operator estimates themselves. Furthermore, we derive the mean squared prediction error (MSPE) and decompose its bias-variance structure. A comprehensive simulation study and an application to high-frequency functional data from wearable devices demonstrate the practical validity of the theory and the ability of FAR models to capture dynamic functional patterns.

stat.ME