SearcharxivSearch

arXiv subjects

Fengnan Gao

Publications and source records attributed to Fengnan Gao.

9 recordsLinked to original sources

Conservative adventurers have more future academic impact

Some scientists explore unfamiliar topics, while others exploit existing ones, yet the link between these choices and academic performance remains unclear. Prior studies offer conflicting evidence, often relying on single metrics and overlooking confounding factors. To address this, we complement the traditional switch frequency metric with switch distances, control for confounders, and establish a clear connection between past switching behaviors and future performance. We identify a group, 'conservative adventurers', who frequently switch topics within close domains, excelling in future performance compared to others (up to 19% more citations per future paper). This rare behavioral pattern suggests an effective balance between exploration and exploitation. Beyond correlations, we question whether conservative adventuring can be intentionally adopted as a strategy. While proving intentionality is challenging, individuals who drastically transition to conservative adventurers likely do so purposefully, achieving significant performance gains. Our findings, based on three datasets covering 31,780,857 papers in physics, biomedicine, and chemistry primarily since the twentieth century (1976-2015 for physics; 1900-2021 for biomedicine and chemistry), provide insights for scientific career understanding and planning, particularly for junior scientists.

cs.DL

Sparse change detection in high-dimensional linear regression

We introduce a new methodology 'charcoal' for estimating the location of sparse changes in high-dimensional linear regression coefficients, without assuming that those coefficients are individually sparse. The procedure works by constructing different sketches (projections) of the design matrix at each time point, where consecutive projection matrices differ in sign in exactly one column. The sequence of sketched design matrices is then compared against a single sketched response vector to form a sequence of test statistics whose behaviour shows a surprising link to the well-known CUSUM statistics of univariate changepoint analysis. The procedure is computationally attractive, and strong theoretical guarantees are derived for its estimation accuracy. Simulations confirm that our methods perform well in extensive settings, and a real-world application to a large single-cell RNA sequencing dataset showcases the practical relevance.

math.ST

Pitfalls of amateur regression: The Dutch New Herring controversies

Applying simple linear regression models, an economist analysed a published dataset from an influential annual ranking in 2016 and 2017 of consumer outlets for Dutch New Herring and concluded that the ranking was manipulated. His finding was promoted by his university in national and international media, and this led to public outrage and ensuing discontinuation of the survey. We reconstitute the dataset, correcting errors and exposing features already important in a descriptive analysis of the data. The economist has continued his investigations, and in a follow-up publication repeats the same accusations. We point out errors in his reasoning and show that alleged evidence for deliberate manipulation of the ranking could easily be an artefact of specification errors. Temporal and spatial factors are both important and complex, and their effects cannot be captured using simple models, given the small sample sizes and many factors determining perceived taste of a food product.

stat.AP

Statistical Inference in Parametric Preferential Attachment Trees

The preferential attachment (PA) model is a popular way of modeling dynamic social networks, such as collaboration networks. Assuming that the PA function takes a parametric form, we propose and study the maximum likelihood estimator of the parameter. Using a supercritical continuous-time branching process framework, we prove the almost sure consistency and asymptotic normality of this estimator. We also provide an estimator that only depends on the final snapshot of the network and prove its consistency, and its asymptotic normality under general conditions. We compare the performance of the estimators to a nonparametric estimator in a small simulation study.

math.ST

Two-sample testing of high-dimensional linear regression coefficients via complementary sketching

We introduce a new method for two-sample testing of high-dimensional linear regression coefficients without assuming that those coefficients are individually estimable. The procedure works by first projecting the matrices of covariates and response vectors along directions that are complementary in sign in a subset of the coordinates, a process which we call 'complementary sketching'. The resulting projected covariates and responses are aggregated to form two test statistics, which are shown to have essentially optimal asymptotic power under a Gaussian design when the difference between the two regression coefficients is sparse and dense respectively. Simulations confirm that our methods perform well in a broad class of settings and an application to a large single-cell RNA sequencing dataset demonstrates its utility in the real world.

math.ST

Community detection in sparse latent space models

We show that a simple community detection algorithm originated from stochastic blockmodel literature achieves consistency, and even optimality, for a broad and flexible class of sparse latent space models. The class of models includes latent eigenmodels (arXiv:0711.1146). The community detection algorithm is based on spectral clustering followed by local refinement via normalized edge counting.

stat.ML

Consistent Estimation in General Sublinear Preferential Attachment Trees

We propose an empirical estimator of the preferential attachment function $f$ in the setting of general preferential attachment trees. Using a supercritical continuous-time branching process framework, we prove the almost sure consistency of the proposed estimator. We perform simulations to study the empirical properties of our estimators.

math.ST

On the Asymptotic Normality of Estimating the Affine Preferential Attachment Network Models with Random Initial Degrees

We consider the estimation of the affine parameter (and power-law exponent) in the preferential attachment model with random initial degrees. We derive the likelihood, and show that the maximum likelihood estimator (MLE) is asymptotically normal and efficient. We also propose a quasi-maximum-likelihood estimator (QMLE) to overcome the MLE's dependence on the history of the initial degrees. To demonstrate the power of our idea, we present numerical simulations.

math.ST

Posterior contraction rates for deconvolution of Dirichlet-Laplace mixtures

We study nonparametric Bayesian inference with location mixtures of the Laplace density and a Dirichlet process prior on the mixing distribution. We derive a contraction rate of the corresponding posterior distribution, both for the mixing distribution relative to the Wasserstein metric and for the mixed density relative to the Hellinger and $L_q$ metrics.

math.ST