SearcharxivSearch

arXiv subjects

Sicheng Lin

Publications and source records attributed to Sicheng Lin.

7 recordsLinked to original sources

RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning

Rare diseases represent one of the most challenging settings for clinical decision-making, where heterogeneous presentations, sparse evidence and limited expertise create persistent uncertainty throughout the care pathway. Although artificial intelligence could help, existing systems largely address isolated tasks, particularly diagnosis, and usually rely on downstream investigations rather than information available at initial presentation. Here we show that clinical AI performance under uncertainty can be improved not by scaling a single model, but by exploiting the diversity of multiple imperfect reasoning systems. Across heterogeneous large language models, we identify divergent reasoning trajectories with complementary error patterns and develop RareLens, which learns to reconcile these perspectives into actionable decisions across four stages of rare disease care: risk screening, diagnosis, treatment planning and prognosis prediction. Built on RarelensBench, a real-world dataset of 157,525 cases spanning all 33 Orphanet categories and more than 7,000 conditions, RareLens outperformed every frontier model tested, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet and Gemini-2.5-Pro, across all stages. It achieved an area under the curve of 0.917 for screening and top-1 accuracies of 65.5% and 89.8% for diagnosis and treatment. In an external evaluation involving 1,287 cases and 23 physicians, autonomous RareLens and physicians assisted by RareLens both outperformed unaided physicians, while demonstrating that effective human-AI collaboration requires more than simply providing model outputs. These findings establish divergent model reasoning as an exploitable source of information and suggest a general strategy for building AI systems that operate reliably under high clinical uncertainty.

cs.AI

Fourth order correlation of baryon number and electric charge as a better magnetometer of QCD

This work focuses on the fourth order correlations $\chi^{BQ}_{31}$, $\chi^{QB}_{31}$, $\chi^{BQ}_{22}$, $\chi^{BS}_{31}$, $\chi^{SB}_{31}$, $\chi^{BS}_{22}$, $\chi^{QS}_{31}$, $\chi^{SQ}_{31}$, $\chi^{QS}_{22}$, $\chi^{BQS}_{211}$, $\chi^{QBS}_{211}$, $\chi^{SBQ}_{211}$ of baryon number $B$, electric charge $Q$ and strangeness $S$ at finite temperature, magnetic field and vanishing quark chemical potential. The study is carried out in the framework of a three-flavor PNJL model, considering both cases with and without inverse magnetic catalysis effect. We find that, fourth order correlations $\chi^{BQ}_{31}$ at chiral restoration phase transition is more sensitive to the magnetic field than other second order and fourth order correlations and fluctuations, and can be served as a more effective magnetometer of QCD.

nucl-th

Abundance matching analysis of the emission line galaxy sample in the extended Baryon Oscillation Spectroscopic Survey

We present the measurements of the small-scale clustering for the emission line galaxy (ELG) sample from the extended Baryon Oscillation Spectroscopic Survey (eBOSS) in the Sloan Digital Sky Survey IV (SDSS-IV). We use conditional abundance matching method to interpret the clustering measurements from $0.34h^{-1}\textrm{Mpc}$ to $70h^{-1}\textrm{Mpc}$. In order to account for the correlation between properties of emission line galaxies and their environment, we add a secondary connection between star formation rate of ELGs and halo accretion rate. Three parameters are introduced to model the ELG [OII] luminosity and to mimic the target selection of eBOSS ELGs. The parameters in our models are optimized using Markov Chain Monte Carlo (MCMC) method. We find that by conditionally matching star formation rate of galaxies and the halo accretion rate, we are able to reproduce the eBOSS ELG small scale clustering within 1$\sigma$ error level. Our best fit model shows that the eBOSS ELG sample only consists of $\sim 12\%$ of all star-forming galaxies, and the satellite fraction of eBOSS ELG sample is 19.3\%. We show that the effect of assembly bias is $\sim20\%$ on the two-point correlation function and $\sim5\%$ on the void probability function at scale of $r\sim 20 h^{-1}\rm Mpc$.

astro-ph.CO

Balancing Approach for Causal Inference at Scale

With the modern software and online platforms to collect massive amount of data, there is an increasing demand of applying causal inference methods at large scale when randomized experimentation is not viable. Weighting methods that directly incorporate covariate balancing have recently gained popularity for estimating causal effects in observational studies. These methods reduce the manual efforts required by researchers to iterate between propensity score modeling and balance checking until a satisfied covariate balance result. However, conventional solvers for determining weights lack the scalability to apply such methods on large scale datasets in companies like Snap Inc. To address the limitations and improve computational efficiency, in this paper we present scalable algorithms, DistEB and DistMS, for two balancing approaches: entropy balancing and MicroSynth. The solvers have linear time complexity and can be conveniently implemented in distributed computing frameworks such as Spark, Hive, etc. We study the properties of balancing approaches at different scales up to 1 million treated units and 487 covariates. We find that with larger sample size, both bias and variance in the causal effect estimation are significantly reduced. The results emphasize the importance of applying balancing approaches on large scale datasets. We combine the balancing approach with a synthetic control framework and deploy an end-to-end system for causal impact estimation at Snap Inc.

stat.ME

The Completed SDSS-IV Extended Baryon Oscillation Spectroscopic Survey: GLAM-QPM mock galaxy catalogs for the Emission Line Galaxy Sample

We present 2000 mock galaxy catalogs for the analysis of baryon acoustic oscillations in the Emission Line Galaxy (ELG) sample of the Extended Baryon Oscillation Spectroscopic Survey Data Release 16 (eBOSS DR16). Each mock catalog has a number density of $6.7 \times 10^{-4} h^3 \rm Mpc^{-3}$, covering a redshift range from 0.6 to 1.1. The mocks are calibrated to small-scale eBOSS ELG clustering measurements at scales of around 10 $h^{-1}$Mpc. The mock catalogs are generated using a combination of GaLAxy Mocks (GLAM) simulations and the Quick Particle-Mesh (QPM) method. GLAM simulations are used to generate the density field, which is then assigned dark matter halos using the QPM method. Halos are populated with galaxies using a halo occupation distribution (HOD). The resulting mocks match the survey geometry and selection function of the data, and have slightly higher number density which allows room for systematic analysis. The large-scale clustering of mocks at the baryon acoustic oscillation (BAO) scale is consistent with data and we present the correlation matrix of the mocks.

astro-ph.CO

The Completed SDSS-IV extended Baryon Oscillation Spectroscopic Survey: measurement of the BAO and growth rate of structure of the emission line galaxy sample from the anisotropic power spectrum between redshift 0.6 and 1.1

We analyse the large-scale clustering in Fourier space of emission line galaxies (ELG) from the Data Release 16 of the Sloan Digital Sky Survey IV extended Baryon Oscillation Spectroscopic Survey. The ELG sample contains 173,736 galaxies covering 1,170 square degrees in the redshift range $0.6 < z < 1.1$. We perform a BAO measurement from the post-reconstruction power spectrum monopole, and study redshift space distortions (RSD) in the first three even multipoles. Photometric variations yield fluctuations of both the angular and radial survey selection functions. Those are directly inferred from data, imposing integral constraints which we model consistently. The full data set has only a weak preference for a BAO feature ($1.4\sigma$). At the effective redshift $z_{\rm eff} = 0.845$ we measure $D_{\rm V}(z_{\rm eff})/r_{\rm drag} = 18.33_{-0.62}^{+0.57}$, with $D_{\rm V}$ the volume-averaged distance and $r_{\rm drag}$ the comoving sound horizon at the drag epoch. In combination with the RSD measurement, at $z_{\rm eff} = 0.85$ we find $f\sigma_8(z_{\rm eff}) = 0.289_{-0.096}^{+0.085}$, with $f$ the growth rate of structure and $\sigma_8$ the normalisation of the linear power spectrum, $D_{\rm H}(z_{\rm eff})/r_{\rm drag} = 20.0_{-2.2}^{+2.4}$ and $D_{\rm M}(z_{\rm eff})/r_{\rm drag} = 19.17 \pm 0.99$ with $D_{\rm H}$ and $D_{\rm M}$ the Hubble and comoving angular distances, respectively. These results are in agreement with those obtained in configuration space, thus allowing a consensus measurement of $f\sigma_8(z_{\rm eff}) = 0.315 \pm 0.095$, $D_{\rm H}(z_{\rm eff})/r_{\rm drag} = 19.6_{-2.1}^{+2.2}$ and $D_{\rm M}(z_{\rm eff})/r_{\rm drag} = 19.5 \pm 1.0$. This measurement is consistent with a flat $\Lambda$CDM model with Planck parameters.

astro-ph.CO

The Completed SDSS-IV extended Baryon Oscillation Spectroscopic Survey: Cosmological Implications from two Decades of Spectroscopic Surveys at the Apache Point observatory

We present the cosmological implications from final measurements of clustering using galaxies, quasars, and Ly$\alpha$ forests from the completed Sloan Digital Sky Survey (SDSS) lineage of experiments in large-scale structure. These experiments, composed of data from SDSS, SDSS-II, BOSS, and eBOSS, offer independent measurements of baryon acoustic oscillation (BAO) measurements of angular-diameter distances and Hubble distances relative to the sound horizon, $r_d$, from eight different samples and six measurements of the growth rate parameter, $f\sigma_8$, from redshift-space distortions (RSD). This composite sample is the most constraining of its kind and allows us to perform a comprehensive assessment of the cosmological model after two decades of dedicated spectroscopic observation. We show that the BAO data alone are able to rule out dark-energy-free models at more than eight standard deviations in an extension to the flat, $\Lambda$CDM model that allows for curvature. When combined with Planck Cosmic Microwave Background (CMB) measurements of temperature and polarization the BAO data provide nearly an order of magnitude improvement on curvature constraints. The RSD measurements indicate a growth rate that is consistent with predictions from Planck primary data and with General Relativity. When combining the results of SDSS BAO and RSD with external data, all multiple-parameter extensions remain consistent with a $\Lambda$CDM model. Regardless of cosmological model, the precision on $\Omega_\Lambda$, $H_0$, and $\sigma_8$, remains at roughly 1\%, showing changes of less than 0.6\% in the central values between models. The inverse distance ladder measurement under a o$w_0w_a$CDM yields $H_0= 68.20 \pm 0.81 \, \rm km\, s^{-1} Mpc^{-1}$, remaining in tension with several direct determination methods. (abridged)

astro-ph.CO