SearcharxivSearch

arXiv subjects

Yu Yu

Publications and source records attributed to Yu Yu.

At least 37 records · Page 2Linked to original sources

A semi-analytical mock galaxy catalog for the CSST extragalactic surveys from the Jiutian simulations

We introduce a mock galaxy catalog built for the CSST extragalactic surveys using the primary runs of the Jiutian $N$-body simulation suites. The catalogs are built by coupling the GAlaxy Evolution and Assembly (GAEA) semi-analytical model of galaxy formation with merger trees extracted from the simulations using the Hierarchical Bound-Tracing (HBT+) algorithm. The spectral energy distributions (SEDs) and broadband magnitudes are computed using the neural-network-based stellar population synthesizer StarDuster, which is trained on radiative transfer simulations to account for detailed galaxy geometry in modeling dust obscuration. Galaxy light-cones up to $z=5$ are subsequently generated with the BLiC light-cone builder which interpolates the properties of galaxies over time using an optimized interpolation scheme. The resulting catalogs exhibit good convergence in many statistical properties of the galaxy population produced from two different resolution simulations. The catalogs reproduce a number of observed galaxy properties across a range of galaxy mass and redshift, including the stellar mass functions, the luminosity function, gas mass fraction, galaxy size-mass relation and galaxy clustering. We also present the photometric and redshift distributions of galaxies expected to be observed in the CSST surveys.

astro-ph.GA

SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion

Large Language Models (LLMs) excel at code-related tasks but often struggle in realistic software repositories, where project-specific APIs and cross-file dependencies are crucial. Retrieval-augmented methods mitigate this by injecting repository context at inference time. The low inference-time latency budget affects either retrieval quality or the added latency adversely impacts user experience. We address this limitation with SpecAgent, an agent that improves both latency and code-generation quality by proactively exploring repository files during indexing and constructing speculative context that anticipates future edits in each file. This indexing-time asynchrony allows thorough context computation, masking latency, and the speculative nature of the context improves code-generation quality. Additionally, we identify the problem of future context leakage in existing benchmarks, which can inflate reported performance. To address this, we construct a synthetic, leakage-free benchmark that enables a more realistic evaluation of our agent against baselines. Experiments show that SpecAgent consistently achieves absolute gains of 9-11% (48-58% relative) compared to the best-performing baselines, while significantly reducing inference latency.

cs.SE

ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer

Patient face images provide a convenient mean for evaluating eye diseases, while also raising privacy concerns. Here, we introduce ROFI, a deep learning-based privacy protection framework for ophthalmology. Using weakly supervised learning and neural identity translation, ROFI anonymizes facial features while retaining disease features (over 98\% accuracy, $\kappa > 0.90$). It achieves 100\% diagnostic sensitivity and high agreement ($\kappa > 0.90$) across eleven eye diseases in three cohorts, anonymizing over 95\% of images. ROFI works with AI systems, maintaining original diagnoses ($\kappa > 0.80$), and supports secure image reversal (over 98\% similarity), enabling audits and long-term care. These results show ROFI's effectiveness of protecting patient privacy in the digital medicine era.

cs.CV

Extending CSST Emulator to post-DESI era

The recent DESI BAO measurements have revealed a potential deviation from a cosmological constant, suggesting a dynamic nature of dark energy. To rigorously test this result, complementary probes such as weak gravitational lensing are crucial, demanding highly accurate and efficient predictions of the nonlinear matter power spectrum within the $w_0w_a$CDM framework. However, most existing emulators fail to cover the full parameter posterior from DESI DR2+CMB constraints in the $w_0\mbox{-}w_a$ plane. In this work, we extend the spectral equivalence method outlined in Casarini et al. 2016 to use auxiliary $w_0w_a$CDM models for approximating the power spectrum of a target $w_0w_a$CDM cosmology, moving beyond the previous use of $w$CDM auxiliaries. Incorporating this enhanced module, the extended CSST Emulator achieves a prediction accuracy of $\leq1\%$ over the $1\sigma$ confidence region from DESI DR2+CMB constraints for $z\leq3$, with a mild degradation in accuracy outside this posterior region. This performance is rigorously validated by additional simulations of dynamic dark energy cosmologies. The emulator's applicable parameter space has been generalized to fully encompass the $2\sigma$ region, greatly enhancing its utility for cosmological analysis in the post-DESI era.

astro-ph.CO

From Text to Alpha: Can LLMs Track Evolving Signals in Corporate Disclosures?

Natural language processing (NLP) has been widely used in quantitative finance, but traditional methods often struggle to capture rich narratives in corporate disclosures, leaving potentially informative signals under-explored. Large language models (LLMs) offer a promising alternative due to their ability to extract nuanced semantics. In this paper, we ask whether semantic signals extracted by LLMs from corporate disclosures predict alpha, defined as abnormal returns beyond broad market movements and common risk factors. We introduce a simple framework, LLM as extractor, embedding as ruler, which extracts context-aware, metric-focused textual spans and quantifies semantic changes across consecutive disclosure periods using embedding-based similarity. This allows us to measure the degree of metric shifting -- how much firms move away from previously emphasized metrics, referred as moving targets. In experiments with portfolio and cross-sectional regression tests against a recent NER-based baseline, our method achieves more than twice the risk-adjusted alpha and shows significantly stronger predictive power. Qualitative analysis suggests that these gains stem from preserving contextual qualifiers and filtering out non-metric terms that keyword-based approaches often miss.

cs.CE

Artifacts in Halo Shapes: Imprints of the Initial Condition

Grid type pre-initial conditions are commonly used to initialize particle positions in cosmological simulations. While these conditions are known to produce noticeable numerical artifacts in void regions, their impact on halo properties has generally been assumed to be negligible. In this work, we employ multiple simulations to demonstrate that grid initialization induces statistically significant artifacts in halo shapes, despite the modest absolute amplitude ($\sim 1\%$) making them unimportant for most cosmological studies. We identify a redshift-dependent artificial alignment pattern: at low redshifts ($z<2$), halo shapes preferentially orient away from the simulation box's Cartesian axes, whereas their constituent particles initially exhibit alignment with these axes. We propose a mathematical hypothesis to explain this flipping behavior.

astro-ph.CO

Validation of Fast Mocks Generation for CSST Photometric Survey

Weak lensing has become a powerful tool for probing the matter distribution in the Universe and constraining cosmological parameters. This paper aims to explore the fast mock generation pipeline to obtain the covariance matrix of the $3\times 2$pt analysis for the upcoming China Space Station Telescope (CSST). We adopt the $N_\mathrm{side}$ pipeline, which generates matter distribution with lognormal assumptions, to create full-sky galaxy mocks with certain two-point statistics. We also employ the Markov-Chain Monte Carlo simulation to test the accuracy of the covariance matrix from the mock-generated galaxy catalogue. Our work validates the accuracy of the $3\times 2$pt statistics in both spherical harmonic space and real space. The critical scale below which the fractional error of correlation exceeds 1$\%$ can decrease as the resolution parameter $N_\mathrm{side}$ increases. After excluding certain scales, the covariance matrix from the mock-generated galaxy catalogue can constrain the cosmological parameters with 0.1$\%$ accuracy. This work demonstrates the potential of $\texttt{GLASS}$ for real-space cosmological measurements and highlights the importance of discarding appropriate scales.

astro-ph.CO

CSST Cosmological Emulator II: Generalized Accurate Halo Mass Function Emulation

Accurate theoretical prediction for halo mass function across a broad cosmological space is crucial for the forthcoming China Space Station Telescope (CSST) observations, which will capture cosmological information from multiple probes, e.g., cluster abundance, and weak lensing. In this work, we quantify the percent-level impact of different mass binning schemes when measuring the differential halo mass function from simulations, and demonstrate that the cumulative form of the halo mass function is independent of the binning scheme. Through the recently finished Kun simulation suite, we propose a generalized framework to construct multiple accurate halo mass function emulators for different halo mass definitions, including $M_{200m}$, $M_{vir}$, and $M_{200c}$. This extends our CSST Emulator to provide fast and accurate halo mass function predictions for halo mass $M\geq 10^{12}\,h^{-1}M_{\odot}$ up to $z=3.0$. For redshifts $z\leq 1.0$, the accuracy is within $2\%$ for $M\leq 10^{13}\,h^{-1}M_{\odot}$, $5\%$ for $M\leq 10^{14}\,h^{-1}M_{\odot}$, and $10\%$ for $M\leq 10^{15}\,h^{-1}M_{\odot}$, which is comparable with the statistical errors of training simulations. This tool is integrated in CSST Emulator and publicly available at https://github.com/czymh/csstemu, providing a fast and accurate theoretical tool to obtain unbiased cosmological constraints of the upcoming CSST survey.

astro-ph.CO

CSST Cosmological Emulator III: Hybrid Lagrangian Bias Expansion Emulation of Galaxy Clustering

Galaxy clustering is an important probe in the upcoming China Space Station Telescope (CSST) survey to understand the structure growth and reveal the nature of the dark sector. However, it is a long-term challenge to model this biased tracer and connect the observable to the underlying physics. In this work, we present a hybrid Lagrangian bias expansion emulator, combining the Lagrangian bias expansion and the accurate dynamical evolution from $N$-body simulation, to predict the power spectrum of the biased tracer in real space. We employ the Kun simulation suite to construct the emulator, emulating across the space of 8 cosmological parameters including dynamic dark energy $w_0$, $w_a$, and total neutrino mass $\sum m_{\nu}$. The sample variance due to the finite simulation box is further reduced using the Zel'dovich variance control, and it enables the precise measurement of the Lagrangian basis spectra up to the quadratic order. The emulation of basis spectra realizes 1% level accuracy, covering wavelength $ k \leq 1 \,{\rm Mpc}^{-1}h$ and redshift $0\leq z\leq 3$ up to the quadratic order field. To validate the emulator, we perform a joint fit to the halo auto power spectrum and the halo-matter cross power spectrum measured from 46 independent simulations. Depending on the choice of counterterm, the joint fit is unbiased up to $k_{\rm max}\simeq 0.7\,{\rm Mpc}^{-1}h$ within $1\sim 2$ percent accuracy, for all the redshift and halo mass samples. As part of the CSST cosmological emulator series, this emulator is expected to provide accurate theoretical predictions for the galaxy power spectrum in upcoming CSST survey.

astro-ph.CO

Xinyu AI Search: Enhanced Relevance and Comprehensive Results with Rich Answer Presentations

Traditional search engines struggle to synthesize fragmented information for complex queries, while generative AI search engines face challenges in relevance, comprehensiveness, and presentation. To address these limitations, we introduce Xinyu AI Search, a novel system that incorporates a query-decomposition graph to dynamically break down complex queries into sub-queries, enabling stepwise retrieval and generation. Our retrieval pipeline enhances diversity through multi-source aggregation and query expansion, while filtering and re-ranking strategies optimize passage relevance. Additionally, Xinyu AI Search introduces a novel approach for fine-grained, precise built-in citation and innovates in result presentation by integrating timeline visualization and textual-visual choreography. Evaluated on recent real-world queries, Xinyu AI Search outperforms eight existing technologies in human assessments, excelling in relevance, comprehensiveness, and insightfulness. Ablation studies validate the necessity of its key sub-modules. Our work presents the first comprehensive framework for generative AI search engines, bridging retrieval, generation, and user-centric presentation.

cs.IR

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models

With the rapid development of Large Language Models (LLMs), aligning these models with human preferences and values is critical to ensuring ethical and safe applications. However, existing alignment techniques such as RLHF or DPO often require direct fine-tuning on LLMs with billions of parameters, resulting in substantial computational costs and inefficiencies. To address this, we propose Micro token-level Accept-Reject Aligning (MARA) approach designed to operate independently of the language models. MARA simplifies the alignment process by decomposing sentence-level preference learning into token-level binary classification, where a compact three-layer fully-connected network determines whether candidate tokens are "Accepted" or "Rejected" as part of the response. Extensive experiments across seven different LLMs and three open-source datasets show that MARA achieves significant improvements in alignment performance while reducing computational costs. The source code and implementation details are publicly available at https://github.com/IAAR-Shanghai/MARA, and the trained models are released at https://huggingface.co/IAAR-Shanghai/MARA_AGENTS.

cs.CL

Meta-Calibration of the Cosmic Magnification Coefficient: Toward Unbiased Weak Lensing Reconstruction by Counting Galaxies

Weak lensing alters galaxy sizes and fluxes, influencing the clustering patterns of galaxies through cosmic magnification. This effect enables the reconstruction of weak lensing convergence $\hat{\kappa}$ maps for DES and DECaLS by linearly combining galaxy overdensities across magnitude bins in the $g$, $r$, and $z$ photometry bands \citep{Qin+,Qin2+}. In this study, we enhance the lensing reconstruction method by addressing biases in the magnification coefficient estimation, which arise from incomplete consideration of selection effects, especially those induced by photometric redshift (photo-$z$) selection. Using a Random Forest-based photo-$z$ estimation for DECaLS and DES galaxies, we quantify the impact of photo-$z$ induced selection on magnification coefficient estimation. Our results show that neglecting photo-$z$ selection introduces significant biases in the magnification coefficient, leading to deviations in the reconstructed convergence map amplitude $A$, with values ranging from 0.4 to 3.5 depending on the survey, redshift, and magnitude cuts. By incorporating an improved magnification coefficient estimation that accounts for photo-$z$ selection, these biases are significantly reduced, with $A$ converging to $\sim 1$ as the magnitude cuts approach optimal values. This improvement is consistently observed across DES and DECaLS datasets and redshift bins, despite differences in survey strategies and depths. Our findings highlight the importance of addressing photo-$z$ induced selection to achieve unbiased weak lensing reconstructions and accurate cosmic magnification measurements.

astro-ph.CO

An Efficient Private GPT Never Autoregressively Decodes

The wide deployment of the generative pre-trained transformer (GPT) has raised privacy concerns for both clients and servers. While cryptographic primitives can be employed for secure GPT inference to protect the privacy of both parties, they introduce considerable performance overhead.To accelerate secure inference, this study proposes a public decoding and secure verification approach that utilizes public GPT models, motivated by the observation that securely decoding one and multiple tokens takes a similar latency. The client uses the public model to generate a set of tokens, which are then securely verified by the private model for acceptance. The efficiency of our approach depends on the acceptance ratio of tokens proposed by the public model, which we improve from two aspects: (1) a private sampling protocol optimized for cryptographic primitives and (2) model alignment using knowledge distillation. Our approach improves the efficiency of secure decoding while maintaining the same level of privacy and generation quality as standard secure decoding. Experiments demonstrate a $2.1\times \sim 6.0\times$ speedup compared to standard decoding across three pairs of public-private models and different network conditions.

cs.CR

The Jiutian simulations for the CSST extra-galactic surveys

We provide an overview of the Jiutian simulations, a hybrid simulation suite for the China Space Survey Telescope (CSST) extragalactic surveys. It consists of four complementary modules: the primary runs with high resolutions with the fiducial concordance cosmology, the emulator runs exploring the parameter uncertainties around the fiducial cosmology, the reconstruction runs intended for recovering the observed Universe position by position, and the extension runs employing extended cosmologies beyond the standard model. For the primary runs, two independent pipelines are adopted to construct subhaloes and merger trees. On top of them, four sets of mock galaxy light-cone catalogs are produced from semi-analytical models and subhalo abundance matching, providing a variety of observational properties including galaxy SED, emission lines, lensing distortions, and mock images. The 129 emulator runs are used to train the CSST emulator, achieving one percent accuracy in predicting the matter power spectrum over $k\leq 10h{\rm Mpc}^{-1}$ and $z\leq 2$. The reconstruction runs employ a number of subgrid baryonic models to predict the evolution and galaxy population resembling certain regions in the real Universe with constrained initial conditions, enabling controlled investigation of galaxy formation on top of structure formation. The extension runs cover models with warm dark matter, $f(R)$ gravity, interacting dark energy, and nonzero neutrino masses, revealing differences in the cosmic structure under alternative cosmological models. We introduce the specifications for each run, the data products derived from them, the corresponding pipeline developments, and present some main tests. Using the primary runs, we also show that the subhalo peak mass functions of different levels are approximately universal. These simulations form a comprehensive and open library for CSST surveys and beyond.

astro-ph.CO

Redshift Distributions of Fast Radio Bursts Inferred Using Clustering in Dispersion Measure Space

Fast radio bursts (FRBs), millisecond-duration radio transient events, possess the potential to serve as excellent cosmological probes. The FRB redshift distribution contains information about the FRB sources, providing key constraints on the types of engines. However, it is quite challenging to obtain the FRB redshifts due to the poor localization and the faintness of the host galaxies. This reality severely restricts the application prospects and study of the physical origins of FRBs. We propose that the clustering of observed FRBs can be an effective approach to address this issue without needing to accurately model dispersion measure (DM) contributions from the host galaxy and the immediate environment of the source. Using the clustering of $5\times 10^7$ simulated FRBs from future observations with sensitivity similar to the second phase of the Square Kilometre Array, we show that in extragalactic DM space, the redshift distributions can be accurately reconstructed, and the mean redshift for FRBs between 384.8 and 1450.3 $\rm pc\,cm^{-3}$ can be constrained to $\sim\!0.001\pm0.003 (1+z)$. The results demonstrate the potential of FRB clustering to constrain redshift distributions and provide valuable insights into FRB source models and cosmological applications.

astro-ph.CO

Semi-Gradient SARSA Routing with Theoretical Guarantee on Traffic Stability and Weight Convergence

We consider the traffic control problem of dynamic routing over parallel servers, which arises in a variety of engineering systems such as transportation and data transmission. We propose a semi-gradient, on-policy algorithm that learns an approximate optimal routing policy. The algorithm uses generic basis functions with flexible weights to approximate the value function across the unbounded state space. Consequently, the training process lacks Lipschitz continuity of the gradient, boundedness of the temporal-difference error, and a prior guarantee on ergodicity, which are the standard prerequisites in existing literature on reinforcement learning theory. To address this, we combine a Lyapunov approach and an ordinary differential equation-based method to jointly characterize the behavior of traffic state and approximation weights. Our theoretical analysis proves that the training scheme guarantees traffic state stability and ensures almost surely convergence of the weights to the approximate optimum. We also demonstrate via simulations that our algorithm attains significantly faster convergence than neural network-based methods with an insignificant approximation error.

cs.LG

CSST Cosmological Emulator I: Matter Power Spectrum Emulation with one percent accuracy to k = 10 h/Mpc

In the near future, the China Space Station Telescope (CSST) will obtain unprecedented imaging and spectroscopic data. The statistical errors in the cosmological parameter constraints will be reduced significantly. The corresponding theoretical tools must meet the percent-level accuracy required to extract as much cosmological information as possible from the observations. We present the CSST Emulator to provide nonlinear power spectrum predictions in the eight cosmological parameter space $\Omega_\mathrm{cb},\Omega_\mathrm{b},H_{0},n_{s},A_{s},w_{0}, w_{a}$, and $m_\nu$. It is constructed based on the Kun simulation suite, consisting of 129 high-resolution simulations with box size $L=1\,h^{-1} {\rm Gpc}$ and evolving $3072^3$ particles. The determinations of parameter ranges, sampling method, and emulation strategy in the whole construction have been optimized exquisitely. This enables our prediction for $k\leq 10\,h {\rm Mpc}^{-1}$ and $z\leq 2.0$ to reach $1\%$ accuracy validated through internal and external simulations. We also compare our results with recent BACCO, EuclidEmulator2, and Mira-Titan IV emulators, which demonstrate the CSST Emulator's excellent performance across a wide cosmological parameter range in the nonlinear regime. CSST Emulator is publicly released at https://github.com/czymh/csstemu, and provides a fundamental theoretical tool for accurate cosmological inference with future CSST observations.

astro-ph.CO

Weak Lensing Reconstruction by Counting Galaxies: Improvement with DES Y3 Galaxies

In \citep{Qin+}, we attempted to reconstruct the weak lensing convergence map $\hat{\kappa}$ from cosmic magnification by linearly weighting the DECaLS galaxy overdensities in different magnitude bins of $grz$ photometry bands. The $\hat{\kappa}$ map is correlated with cosmic shear at 20-$\sigma$ significance. However, the low galaxy number density in the DECaLS survey prohibits the measurement of $\hat{\kappa}$ auto-correlation. In this paper, we apply the reconstruction method to the Dark Energy Survey Year 3 (DES Y3) galaxies from the DES Data Release 2 (DR2). With greater survey depth and higher galaxy number density, convergence-shear cross-correlation signals are detected with $S/N\approx 9,16,20$ at $0.4<z_\kappa<0.6,0.6<z_\kappa<0.8$ and $0.8<z_\kappa<1.0$ respectively. More remarkably, the $\hat{\kappa}-\hat{\kappa}$ correlations of the $0.4<z_\kappa<0.6$ and $0.6<z_\kappa<0.8$ bins show reasonably good agreement with predictions based on theoretical interpretation of $\hat{\kappa}-\gamma$ measurement. This result takes a step further towards the cosmological application of our lensing reconstruction method.

astro-ph.CO