SearcharxivSearch

arXiv subjects

Zhou Fan

Publications and source records attributed to Zhou Fan.

At least 19 recordsLinked to original sources

Empirical Bayes linear regression in high dimensions: Method of moments and sub-linear sample complexity

We study empirical Bayes estimation of the prior in high-dimensional linear regression $\mathbf{y}=\mathbf{X}\mathbf{\beta}+\mathbf{\varepsilon}$, where the regression coefficients are drawn independently from an unknown sub-Gaussian prior. In contrast to the sequence model, the design matrix couples the latent coefficients, so that recovering the prior requires deconvolving it from both the noise and copies of itself. We introduce the \emph{Empirical Bayes Method of Moments} (EBMoM), a computationally efficient procedure for general designs that recursively estimates the prior moments through a lower-triangular system of estimating equations and runs in time $O(np^2)$. Under mild design conditions, satisfied in particular by a broad class of correlated random designs, we show that EBMoM consistently estimates a growing number of moments and hence the prior itself, provided that $n\geq p^{1-o(1)}$. A matching information-theoretic lower bound, valid for a broad class of designs, shows that this sub-linear sample complexity is optimal for nonparametric prior estimation. This improves on existing results for likelihood-based methods whose consistency requires a linear sample size $n=\Omega(p)$.

math.ST

Dynamical mean-field limit and replica-symmetric free energy for the orthogonally-invariant SK model

We study a class of diffusion processes on $\mathbb{R}^n$ interacting through a symmetric matrix $X\in\mathbb{R}^{n\times n}$. When eigenvectors of $X$ are Haar-uniform on the orthogonal group, we derive a dynamical mean-field limit for the empirical law of sample paths, extending the classical Sompolinsky--Zippelius characterization for $X\sim\mathrm{GOE}$. The limit takes the form of a generalized Langevin equation with correlated Gaussian noise and memory, whose correlation and response kernels relate to those of the original dynamics through convolution equations involving the free cumulants of the eigenvalue distribution of $X$. For the overdamped Langevin diffusion associated with $\mu(\boldsymbol{\theta})\propto \exp\!\big(\frac12\boldsymbol{\theta}^{\top}X\boldsymbol{\theta}\big)\prod_{i=1}^n\nu(\mathrm{d}\theta_i)$, we analyze the mean-field limit under a rapid-mixing assumption. The correlation and response kernels admit time-translation-invariant approximants satisfying a fluctuation-dissipation relation. The generalized Langevin equation admits a Markovian approximation coupled to an auxiliary multivariate OU process and converges to a replica-symmetric prediction for the empirical coordinate law under $\mu$. This auxiliary correlation structure is characterized through the infinitesimal generator of a Markov semigroup for the lifted path-history process. Consequently, the free energy converges to a replica-symmetric limit under an explicit high-temperature condition, which for an Ising model is $\|X\|_{\mathrm{op}}<1/2$. By recent dynamical universality results, the same free-energy characterization holds for deterministic models without random disorder when $X$ satisfies a set of deterministic delocalization conditions.

math.PR

Geometric planted matchings in high dimensions: The power of multiple views

We study the problem of recovering the correspondence between a collection of $n$ points in $\mathbb{R}^d$ and a noisy, permuted version of those points. In the high-dimensional regime $d=\omega(\log n)$, under a Gaussian model with noise variance $\sigma^2=d/(b\log n)$, prior work identifies $b=2$ as the threshold for almost exact recovery. We prove that this threshold is all-or-nothing: for every fixed $b<2$, no estimator recovers a positive fraction of the matching, and even estimating the matched point cloud in Euclidean distance is asymptotically no better than ignoring the correspondence. On the other hand, we consider a multi-view generalization of the problem where $K$ noisy, independently permuted copies of the same latent point cloud are observed. Here we show that a simple polynomial-time procedure recovers all relative matchings up to $o(n)$ errors whenever $b>K/(K-1)$. Thus multiple views can break the impossibility barrier $b=2$ for the original matching problem: in particular, for $3/2 < b < 2$, the two-view model has no nontrivial recovery, but a third view makes all latent correspondences efficiently recoverable.

math.ST

X-rays breaking out of pre-explosion ejecta mark a supernova's first light

Massive stars die as core-collapse supernovae, whose optical light emerges days after the implosion. Theory predicts that the initial collapse-driven shock, upon breaking through the star and dense circumstellar medium, emits a brief thermal flash of soft X-rays and ultraviolet. Yet these elusive first signals have remained largely undetected, owing to limited wide-field soft X-ray monitoring. Here we report the discovery of a soft X-ray flash, EP260321a, followed days later by a broad-lined supernova from an envelope-stripped progenitor. Its X-ray spectrum, best modeled with blackbody, establishes it as the long-sought archetypal shock breakout. The burst's duration and energetics place the breakout at a radius of 300 solar radii, tracing a dense surrounding shell and revealing abrupt mass ejection within the final month before collapse.

astro-ph.HE

Impact of Satellite Constellations on Observations with the 80-cm Telescope and the Mini-SiTian at the Xinglong Observatory, NAOC

The rapid development of mega-constellations in low Earth orbit (LEO) severely impacts ground-based optical astronomical observations. By combining WorldWide Telescope (WWT) simulations with 2019 and 2023 observational data from the Xinglong Observatory 80-cm telescope and 2023 data from the Mini-SiTian (MST), we find that satellite visibility increases with deployment, particularly during the summer. For the 80-cm telescope, the fraction of images containing satellite trails increased from an average of 0.34% in 2019 to 0.7% in 2023; meanwhile, for the MST in 2023, the fraction rose from 5% in January to 12% by December, peaking at 19% in the summer. Through stratified analysis of solar elevation and local time, we find that observations during twilight and summer are particularly susceptible to satellite trail interference. Photometric analysis reveals that the interference intensity increases for fainter sources and those closer to the trails. Furthermore, a comparative analysis across different seeing conditions shows that the deviation of median standardized residuals ({\sigma}) is significantly greater under poor seeing than under good seeing conditions.

astro-ph.IM

The Stellar Abundances and Galactic Evolution Survey (SAGES). V. The First Data Release of the DDO51 Band

We present the first public data release of DDO51 band from the Stellar Abundances and Galactic Evolution Survey (SAGES), based on Nanshan One-meter Wide-field Telescope (NOWT) observations obtained between 2023 September and 2024 January. This release initiates the DDO51-band component of the survey, covering $\sim$ 2,500 deg$^2$ of the northern sky and including more than 10 million sources. The DDO51 filter is centered near the \ion{Mg}{1}~$b$ triplet and the adjacent MgH feature, offering sensitivity to stellar surface gravity. The data reduction pipeline incorporates an improved astrometric solution anchored to Gaia DR3 and a photometric calibration strategy tied to synthetic photometry from Gaia XP spectra. These procedures yield a point-source depth of $\sim$18.9 mag at S/N$\sim$10 and an internal photometric precision $\approx$6-7 mmag at the bright end. A preliminary color--color analysis using Gaia broadband photometry confirms the expected sensitivity of the DDO51 band to stellar surface gravity, demonstrating a clear photometric separation between dwarf and giant sequences for late-type stars. This dataset, when combined with existing SAGES photometry in other bands, provides a crucial tool for disentangling the substructures of the Milky Way. All data products from this release upon publication will be available.

astro-ph.SR

The multiple corrugations in the Galactic disk derived from the LAMOST and Gaia survey data

Large spectroscopic and astrometric surveys have revealed complex wave-like features in the Milky Way disk, suggesting that its kinematic and chemical structures are shaped by time-dependent perturbations. Recent studies have reported oscillatory patterns in the Rg-Vphi-VR space, hinting at a possible structural transition in the outer disk. We aim to characterise the transition between the inner and outer Galactic thin disk and to investigate whether radial corrugations can provide a plausible physical interpretation of the observed features. We analysed two large stellar samples from LAMOST DR8 and Gaia DR3, combining spatial, kinematic, and chemical diagnostics. A simplified corrugation model consisting of two radial waves propagating in opposite directions was constructed and fitted to the observed VR pattern. We further validated the model using N-body simulations. Both LAMOST and Gaia samples reproduce the previously reported wave-like pattern in the Rg-Vphi-VR plane. We identify a clear transition between the inner and outer disks via the variations in rotational velocity and metallicities. The corrugation model naturally reproduces the periodic variation of VR with galactocentric radius, and the superposition of the inward and outward propagating modes gives rise to a comparable oscillatory pattern in both observations and simulations. Our modelling suggests that radial corrugations can provide a plausible interpretation of the observed kinematic signatures. The results highlight the complex, multi-perturber nature of the Galactic disk and motivate further investigation with upcoming surveys.

astro-ph.GA

A fast X-ray transient with chromatic flares: signatures of violent collisions induced by late-time central engine reactivation

Extragalactic Fast X-ray Transients (EFXTs) represent an emerging class of high-energy phenomena characterized by X-ray outbursts lasting from tens to hundreds of seconds. However, for more than half of the EFXTs, their physical origins remain elusive. In this Letter, we report the discovery of EP250302a, a luminous EFXT detected by the Einstein Probe (EP) at a redshift of $z = 1.131$. The multi-wavelength light curves of EP250302a reveal remarkable temporal features that distinguish it from the previously known EP-detected EFXT population, most notably a needle-like X-ray flare accompanied by smooth optical rebrightening during the afterglow phase. We suggest that the distinct X-ray and optical behaviors constitute the first observed instance of late-time violent collision of two relativistic shells in an EFXT. Drawing on insights from GRB studies, such a collision process strongly indicates the reactivation of a central engine, making EP250302a-like transients a unique laboratory for probing the late-time activity and jet physics of EFXT central engines.

astro-ph.HE

The Spectroscopic and Photometric Study of a Star Cluster Sample in Andromeda Halo

Halo star clusters serve as vital tracers for the formation and evolution of the Andromeda galaxy. In this work, we present physical parameters for 29 M31 halo star clusters, derived from a combination of spectroscopic and photometric data. Low-resolution spectra were acquired using the BFOSC spectrograph on the NAOC Xinglong 2.16-m telescope. For the photometric analysis, we utilized uSC and vSAGE bands from the SAGE survey, complemented by archival data from GALEX(NUV, FUV), PAN-STARRS(grizy) and the 2MASS(JHK). Ages and metallicities were determined via ULySS (Vazdekis et al. and pegase-hr) SSP model and the Bruzual & Charlot (2003) (BC03) stellar population synthesis models. The derived parameters show good agreement with literature values. Notably, for three of these clusters, this study represents the first combined photometric and spectroscopic analysis.

astro-ph.GA

Spectral Dataset of Stripped-Envelope Supernovae from the Tsinghua Supernova Group

The extent of envelope stripping in the progenitor stars is directly reflected in the diversity of spectral features observed in stripped-envelope supernovae (SESNe). Through extensive spectral observation and analysis, we aim to clarify the statistical differences between the subclasses of SESNe. The Tsinghua Supernova group obtained 249 optical spectra of 62 SESNe during the years from 2010 to 2020, covering phases from $-$16 to over 190 days relative to maximum light. Most spectra were obtained during the photospheric phases after the supernova explosion. For each spectrum, the pseudo-equivalent widths (pEWs) and blueshift velocities of principal lines were measured. We further investigated the common spectral features by analysing their velocity and strength correlations across all subtypes. We identify the feature near 6200~\AA\ in SNe Ib as H$\mathrm{\alpha}$ through comparison with SNe IIb and Ic, which resolves inconsistent literature interpretations. Our finding reveals prevalent residual hydrogen in SNe Ib, further supporting a continuous stripping sequence from SNe IIb to Ib. We observe a trend in increasing velocity among different subtypes of stripped-envelope SNe, with SNe IIb exhibiting the lowest line velocities, followed by Ib, Ic, and Ic-BL. Typically, the O~I lines in SNe Ic/Ic-BL are stronger than those seen in SNe IIb/Ib. In nebular phases, the [Ca II] emission dominates over [O I] in SNe IIb/Ib while [O I] is stronger in SNe Ic, including the He-rich SN 2016coi. This spectral dichotomy implies that progenitors of SNe Ic (BL) have more massive CO cores and hence higher initial masses.

astro-ph.HE

Bayesian inference of planted matchings: Local posterior approximation and infinite-volume limit

We study Bayesian inference of an unknown matching $\pi^*$ between two correlated random point sets $\{X_i\}_{i=1}^n$ and $\{Y_i\}_{i=1}^n$ in $[0,1]^d$, under a critical scaling $\|X_i-Y_{\pi^*(i)}\|_2 \asymp n^{-1/d}$, in both an exact matching model where all points are observed and a partial matching model where a fraction of points may be missing. Restricting to the simplest setting of $d=1$, in this work, we address the questions of (1) whether the posterior distribution over matchings is approximable by a local algorithm, and (2) whether marginal statistics of this posterior have a well-defined limit as $n \to \infty$. We answer both questions affirmatively for partial matching, where a decay-of-correlations arises for large $n$. For exact matching, we show that the posterior is approximable locally only after a global sorting of the points, and that defining a large-$n$ limit of marginal statistics requires a careful indexing of points in the Poisson point process limit of the data, based on a notion of flow. We leave as an open question the extensions of such results to dimensions $d \geq 2$.

math.ST

Anisotropic local law for non-separable sample covariance matrices

We establish local laws for sample covariance matrices $K = N^{-1}\sum_{i=1}^N \g_i\g_i^*$ where the random vectors $\g_1, \ldots, \g_N \in \R^n$ are independent with common covariance $\Sigma$. Previous work has largely focused on the separable model $\g = \Sigma^{1/2}\w$ with $\w$ having independent entries, but this structure is rarely present in statistical applications involving dependent or nonlinearly transformed data. Under a concentration assumption for quadratic forms $\g^*A\g$, we prove an optimal averaged local law showing that the Stieltjes transform of $K$ converges to its deterministic limit uniformly down to the optimal scale $\eta \geq N^{-1+\eps}$. Under an additional structural assumption on the cumulant tensors of $\g$ -- which interpolates between the highly structured case of independent entries and generic dependence -- we establish the full anisotropic local law, providing entrywise control of the resolvent $(K-zI)^{-1}$ in arbitrary directions. We discuss several classes of non-separable examples satisfying our assumptions, including conditionally mean-zero distributions, the random features model $\g = \sigma(X\w)$ arising in machine learning, and Gaussian measures with nonlinear tilting. The proofs introduce a tensor network framework for analyzing fluctuation averaging in the presence of higher-order cumulant structure.

math.PR

SN 2024abfl: A Low-Luminosity Type IIP Supernova in NGC 2146 from a Low-Mass Red Supergiant Progenitor

Type IIP supernovae (SNe IIP) exhibit a significant diversity in their explosion properties, yet the physical mechanisms driving this diversity remain unknown. In this work, we present photometric and spectroscopic observations of SN 2024abfl, a SN IIP in NGC 2146 with a directly detected red supergiant (RSG) progenitor. We find it has a low plateau luminosity ($M_V \sim -15$ mag) and a relatively long plateau length ($\sim 126.5$ days). By fitting a semi-analytical model, we estimated a $^{56}$Ni mass of $\sim 0.009 M_\odot$, an initial kinetic energy of $\sim 0.42$ foe, an initial thermal energy of $\sim 0.03$ foe and an ejecta mass of $\sim 8.3 M_\odot$. The spectral evolution of SN 2024abfl is similar to those of other SNe IIP, except for much lower ejecta velocities at similar epochs. At later epochs, we find a relatively high-velocity H$\alpha$ absorption feature at $\sim -4000$ km s$^{-1}$, possibly due to a fast-moving plume of matter in the inner ejecta, and two emission features at $\pm 2000$ km s$^{-1}$, possibly caused by CSM interaction. We estimate the progenitor mass to be $\le 15 M_\odot$ based on nebular spectra. We conclude that SN 2024abfl is a low-luminosity SN IIP originating from a low-mass RSG progenitor.

astro-ph.HE

High-dimensional learning dynamics of multi-pass Stochastic Gradient Descent in multi-index models

We study the learning dynamics of a multi-pass, mini-batch Stochastic Gradient Descent (SGD) procedure for empirical risk minimization in high-dimensional multi-index models with isotropic random data. In an asymptotic regime where the sample size $n$ and data dimension $d$ increase proportionally, for any sub-linear batch size $\kappa \asymp n^\alpha$ where $\alpha \in [0,1)$, and for a commensurate ``critical'' scaling of the learning rate, we provide an asymptotically exact characterization of the coordinate-wise dynamics of SGD. This characterization takes the form of a system of dynamical mean-field equations, driven by a scalar Poisson jump process that represents the asymptotic limit of SGD sampling noise. We develop an analogous characterization of the Stochastic Modified Equation (SME) which provides a Gaussian diffusion approximation to SGD. Our analyses imply that the limiting dynamics for SGD are the same for any batch size scaling $\alpha \in [0,1)$, and that under a commensurate scaling of the learning rate, dynamics of SGD, SME, and gradient flow are mutually distinct, with those of SGD and SME coinciding in the special case of a linear model. We recover a known dynamical mean-field characterization of gradient flow in a limit of small learning rate, and of one-pass/online SGD in a limit of increasing sample size $n/d \to \infty$.

stat.ML

When does Gaussian equivalence fail and how to fix it: Non-universal behavior of random features with quadratic scaling

A major effort in modern high-dimensional statistics has been devoted to the analysis of linear predictors trained on nonlinear feature embeddings via empirical risk minimization (ERM). Gaussian equivalence theory (GET) has emerged as a powerful universality principle in this context: it states that the behavior of high-dimensional, complex features can be captured by Gaussian surrogates, which are more amenable to analysis. Despite its remarkable successes, numerical experiments show that this equivalence can fail even for simple embeddings -- such as polynomial maps -- under general scaling regimes. We investigate this breakdown in the setting of random feature (RF) models in the quadratic scaling regime, where both the number of features and the sample size grow quadratically with the data dimension. We show that when the target function depends on a low-dimensional projection of the data, such as generalized linear models, GET yields incorrect predictions. To capture the correct asymptotics, we introduce a Conditional Gaussian Equivalent (CGE) model, which can be viewed as appending a low-dimensional non-Gaussian component to an otherwise high-dimensional Gaussian model. This hybrid model retains the tractability of the Gaussian framework and accurately describes RF models in the quadratic scaling regime. We derive sharp asymptotics for the training and test errors in this setting, which continue to agree with numerical simulations even when GET fails. Our analysis combines general results on CLT for Wiener chaos expansions and a careful two-phase Lindeberg swapping argument. Beyond RF models and quadratic scaling, our work hints at a rich landscape of universality phenomena in high-dimensional ERM.

math.ST

Slitless Spectroscopy Source Detection Using YOLO Deep Neural Network

Slitless spectroscopy eliminates the need for slits, allowing light to pass directly through a prism or grism to generate a spectral dispersion image that encompasses all celestial objects within a specified area. This technique enables highly efficient spectral acquisition. However, when processing CSST slitless spectroscopy data, the unique design of its focal plane introduces a challenge: photometric and slitless spectroscopic images do not have a one-to-one correspondence. As a result, it becomes essential to first identify and count the sources in the slitless spectroscopic images before extracting spectra. To address this challenge, we employed the You Only Look Once (YOLO) object detection algorithm to develop a model for detecting targets in slitless spectroscopy images. This model was trained on 1,560 simulated CSST slitless spectroscopic images. These simulations were generated from the CSST Cycle 6 and Cycle 9 main survey data products, representing the Galactic and nearby galaxy regions and the high galactic latitude regions, respectively. On the validation set, the model achieved a precision of 88.6% and recall of 90.4% for spectral lines, and 87.0% and 80.8% for zeroth-order images. In testing, it maintained a detection rate >80% for targets brighter than 21 mag (medium-density regions) and 20 mag (low-density regions) in the Galactic and nearby galaxies regions, and >70% for targets brighter than 18 mag in high galactic latitude regions.

astro-ph.IM

A$^2$Search: Ambiguity-Aware Question Answering with Reinforcement Learning

Recent advances in Large Language Models (LLMs) and Reinforcement Learning (RL) have led to strong performance in open-domain question answering (QA). However, existing models still struggle with questions that admit multiple valid answers. Standard QA benchmarks, which typically assume a single gold answer, overlook this reality and thus produce inappropriate training signals. Existing attempts to handle ambiguity often rely on costly manual annotation, which is difficult to scale to multi-hop datasets such as HotpotQA and MuSiQue. In this paper, we present A$^2$Search, an annotation-free, end-to-end training framework to recognize and handle ambiguity. At its core is an automated pipeline that detects ambiguous questions and gathers alternative answers via trajectory sampling and evidence verification. The model is then optimized with RL using a carefully designed $\mathrm{AnsF1}$ reward, which naturally accommodates multiple answers. Experiments on eight open-domain QA benchmarks demonstrate that A$^2$Search achieves new state-of-the-art performance. With only a single rollout, A$^2$Search-7B yields an average $\mathrm{AnsF1}@1$ score of $48.4\%$ across four multi-hop benchmarks, outperforming all strong baselines, including the substantially larger ReSearch-32B ($46.2\%$). Extensive analyses further show that A$^2$Search resolves ambiguity and generalizes across benchmarks, highlighting that embracing ambiguity is essential for building more reliable QA systems. Our code, data, and model weights can be found at https://github.com/zfj1998/A2Search

cs.CL

A fast powerful X-ray transient from possible tidal disruption of a white dwarf

Stars captured by black holes (BHs) can be torn apart by strong tidal forces, producing electromagnetic flares. To date, more than 100 tidal disruption events (TDEs) have been observed, each involving invariably normal gaseous stars whose debris falls onto the BH, sustaining the flares over years. White dwarfs (WDs), which are the most prevalent compact stars and a million times denser--and therefore tougher--than gaseous stars, can only be disrupted by intermediate-mass black holes (IMBHs) of 10^2--10^5 solar masses. WD-TDEs are considered to generate more powerful and short-lived flares, but their evidence has been lacking. Here we report observations of a fast and luminous X-ray transient EP250702a detected by Einstein Probe. Its one-day-long X-ray peak as luminous as 10^(47-49) erg/s showed strong recurrent flares with hard spectra extending to several tens of MeV gamma-rays, as detected by Fermi/GBM and Konus-Wind, indicating relativistic jet emission. The jet's X-ray dropped sharply from 3 x 10^49 erg/s to around 10^44 erg/s within 20 days (10 days in the source rest frame). These characteristics are inconsistent with any known transient phenomena other than a jetted-TDE evolving over an unprecedentedly short timescale, indicating the disruption of a WD by an IMBH. At late times, a new soft component progressively dominates the X-ray spectrum, exhibiting an extreme super-Eddington luminosity, which possibly originates from an accretion disc. WD-TDEs open a new window for investigating the elusive IMBHs and their surrounding stellar environments, and they are prime sources of gravitational waves in the band of space-based interferometers.

astro-ph.HE