SearcharxivSearch

arXiv subjects

Yichen Qin

Publications and source records attributed to Yichen Qin.

At least 19 recordsLinked to original sources

On the invariance of irregular Hodge numbers under crepant birational equivalences

The Batyrev--Kontsevich theorem asserts that birational Calabi--Yau varieties have the same Hodge numbers. In this article, we prove an analogue for Landau--Ginzburg models $(U,f)$, where $U$ is a smooth quasi-projective complex variety and $f$ is a regular function on $U$. Under a natural non-degeneracy assumption, we show that the classes of such models in the localized Grothendieck ring of complex algebraic varieties with exponentials are invariant under crepant birational equivalences. Consequently, the irregular Hodge numbers of the twisted de Rham cohomology $\mathrm{H}^k_{\mathrm{dR}}(U,f)$ are invariant as well.

math.AG

Classical and irregular Hodge numbers

Let $U$ be a smooth quasi-projective complex variety with a regular function $f$. The twisted de Rham cohomology groups $\mathrm{H}^k_{\mathrm{dR}}(U, f)$ carry the decreasing irregular Hodge filtration, whose graded pieces have dimensions known as the irregular Hodge numbers. In this paper, we prove that the irregular Hodge numbers admit an explicit characterization in terms of classical Hodge numbers, closely related to Hodge-theoretic numbers constructed by Katzarkov, Kontsevich, and Pantev for Landau--Ginzburg models. As direct applications, we show that irregular Hodge numbers of non-degenerate functions are independent of the choice of non-degenerate functions, and we give a concrete formula for irregular Hodge numbers for unipotent non-degenerate functions.

math.AG

Perfect Clustering for Sparse Directed Stochastic Block Models

Exact recovery in stochastic block models (SBMs) is well understood in undirected settings, but remains considerably less developed for directed and sparse networks, particularly when the number of communities diverges. Spectral methods for directed SBMs often lack stability in asymmetric, low-degree regimes, and existing non-spectral approaches focus primarily on undirected or dense settings. We propose a fully non-spectral, two-stage procedure for community detection in sparse directed SBMs with potentially growing numbers of communities. The method first estimates the directed probability matrix using a neighborhood-smoothing scheme tailored to the asymmetric setting, and then applies $K$-means clustering to the estimated rows, thereby avoiding the limitations of eigen- or singular value decompositions in sparse, asymmetric networks. Our main theoretical contribution is a uniform row-wise concentration bound for the smoothed estimator, obtained through new arguments that control asymmetric neighborhoods and separate in- and out-degree effects. These results imply the exact recovery of all community labels with probability tending to one, under mild sparsity and separation conditions that allow both $\gamma_n \to 0$ and $K_n \to \infty$. Simulation studies, including highly directed, sparse, and non-symmetric block structures, demonstrate that the proposed procedure performs reliably in regimes where directed spectral and score-based methods deteriorate. To the best of our knowledge, this provides the first exact recovery guarantee for this class of non-spectral, neighborhood-smoothing methods in the sparse, directed setting.

stat.ML

On Siegel's problem and Dwork's conjecture for $G$-functions

We answer in the negative Siegel's problem for $G$-functions, as formulated by Fischler and Rivoal. Roughly, we prove that there are $G$-functions that cannot be written as polynomial expressions in algebraic pullbacks of hypergeometric functions; our examples satisfy differential equations of order two, which is the smallest possible. In fact, we construct infinitely many non-equivalent rank-two local systems of geometric origin which are not algebraic pullbacks of hypergeometric local systems, thereby providing further counterexamples to Dwork's conjecture and answering a question by Krammer. The main ingredients of the proof are a Lie algebra version of Goursat's lemma, the monodromy computations of hypergeometric local systems due to Beukers and Heckman, as well as results on invariant trace fields of Fuchsian groups.

math.NT

Irregular Hodge numbers of Frenkel--Gross connections

Frenkel and Gross constructed a family of connections on $\mathbb{P}^1\backslash\{0,\infty\}$, for almost simple groups $\check{G}$ and their representations. In this article, we calculate the irregular Hodge numbers of these Frenkel--Gross connections, and, as an application, we prove a conjecture of Katzarkov--Kontsevich--Pantev for mirror Landau-Ginzburg models of minuscule homogeneous spaces.

math.AG

Alpha Mining and Enhancing via Warm Start Genetic Programming for Quantitative Investment

Traditional genetic programming (GP) often struggles in stock alpha factor discovery due to its vast search space, overwhelming computational burden, and sporadic effective alphas. We find that GP performs better when focusing on promising regions rather than random searching. This paper proposes a new GP framework with carefully chosen initialization and structural constraints to enhance search performance and improve the interpretability of the alpha factors. This approach is motivated by and mimics the alpha searching practice and aims to boost the efficiency of such a process. Analysis of 2020-2024 Chinese stock market data shows that our method yields superior out-of-sample prediction results and higher portfolio returns than the benchmark.

q-fin.ST

Rigid $G$-connections and nilpotency of $p$-curvatures

Motivated by Simpson's conjecture on the motivicity of rigid irreducible connections, Esnault and Groechenig demonstrated that the mod-$p$ reductions of such connections on smooth projective varieties have nilpotent $p$-curvatures. In this paper, we extend their result to integrable $G$-connections.

math.AG

Adaptive Weighted Random Isolation (AWRI): a simple design to estimate causal effects under network interference

Recently, causal inference under interference has gained increasing attention in the literature. In this paper, we focus on randomized designs for estimating the total treatment effect (TTE), defined as the average difference in potential outcomes between fully treated and fully controlled groups. We propose a simple design called weighted random isolation (WRI) along with a restricted difference-in-means estimator (RDIM) for TTE estimation. Additionally, we derive a novel mean squared error surrogate for the RDIM estimator, supported by a network-adaptive weight selection algorithm. This can help us determine a fair weight for the WRI design, thereby effectively reducing the bias. Our method accommodates directed networks, extending previous frameworks. Extensive simulations demonstrate that the proposed method outperforms nine established methods across a wide range of scenarios.

stat.ME

Assessing Estimation Uncertainty under Model Misspecification

Model misspecification is ubiquitous in data analysis because the data-generating process is often complex and mathematically intractable. Therefore, assessing estimation uncertainty and conducting statistical inference under a possibly misspecified working model is unavoidable. In such a case, classical methods such as bootstrap and asymptotic theory-based inference frequently fail since they rely heavily on the model assumptions. In this article, we provide a new bootstrap procedure, termed local residual bootstrap, to assess estimation uncertainty under model misspecification for generalized linear models. By resampling the residuals from the neighboring observations, we can approximate the sampling distribution of the statistic of interest accurately. Instead of relying on the score equations, the proposed method directly recreates the response variables so that we can easily conduct standard error estimation, confidence interval construction, hypothesis testing, and model evaluation and selection. It performs similarly to classical bootstrap when the model is correctly specified and provides a more accurate assessment of uncertainty under model misspecification, offering data analysts an easy way to guard against the impact of misspecified models. We establish desirable theoretical properties, such as the bootstrap validity, for the proposed method using the surrogate residuals. Numerical results and real data analysis further demonstrate the superiority of the proposed method.

stat.ME

Irregular Hodge filtration of hypergeometric differential equations

Fedorov and Sabbah--Yu calculated the (irregular) Hodge numbers of hypergeometric connections. In this paper, we study the irregular Hodge filtrations on hypergeometric connections defined by rational parameters, and provide a new proof of the aforementioned results. Our approach is based on a geometric interpretation of hypergeometric connections, which enables us to show that certain hypergeometric sums are everywhere ordinary on $|\mathbb{G}_{m,\mathbb{F}_p}|$ (i.e. "Frobenius Newton polygon equals to irregular Hodge polygon").

math.AG

Sparsified Simultaneous Confidence Intervals for High-Dimensional Linear Models

Statistical inference of the high-dimensional regression coefficients is challenging because the uncertainty introduced by the model selection procedure is hard to account for. A critical question remains unsettled; that is, is it possible and how to embed the inference of the model into the simultaneous inference of the coefficients? To this end, we propose a notion of simultaneous confidence intervals called the sparsified simultaneous confidence intervals. Our intervals are sparse in the sense that some of the intervals' upper and lower bounds are shrunken to zero (i.e., $[0,0]$), indicating the unimportance of the corresponding covariates. These covariates should be excluded from the final model. The rest of the intervals, either containing zero (e.g., $[-1,1]$ or $[0,1]$) or not containing zero (e.g., $[2,3]$), indicate the plausible and significant covariates, respectively. The proposed method can be coupled with various selection procedures, making it ideal for comparing their uncertainty. For the proposed method, we establish desirable asymptotic properties, develop intuitive graphical tools for visualization, and justify its superior performance through simulation and real data analysis.

stat.ME

$L$-functions of Kloosterman sheaves

In this article, we study a family of motives $\mathrm{M}_{n+1}^k$ associated with the symmetric power of Kloosterman sheaves, as constructed by Fres\'an, Sabbah, and Yu. They demonstrated that for $n=1$, the motivic $L$-functions of $\mathrm{M}_{2}^k$ extend meromorphically to $\mathbb{C}$ and satisfy the functional equations conjectured by Broadhurst and Roberts. Our work aims to extend these results to the motivic $L$-functions of some of the motives $\mathrm{M}_{n+1}^k$, with $n>1$, as well as other related $2$-dimensional motives. In particular, we prove several conjectures of Evans type, which relate traces of Kloosterman sheaves and Fourier coefficients of modular forms.

math.AG

Hodge numbers of motives attached to Kloosterman and Airy moments

Fres\'an, Sabbah, and Yu constructed motives $\mathrm{M}_{n+1}^k(\mathrm{Kl})$ over $\mathbb{Q}$ encoding symmetric power moments of Kloosterman sums in $n$ variables. When $n=1$, they use the irregular Hodge filtration on the exponential mixed Hodge structure associated with $\mathrm{M}_{2}^k(\mathrm{Kl})$ to compute the Hodge numbers of $\mathrm{M}_{2}^k(\mathrm{Kl})$, which turn out to be either $0$ or $1$. In this article, I explain how to compute the (irregular) Hodge numbers of $\mathrm{M}_{n+1}^k(\mathrm{Kl})$ for $n=2$ or for general values of $n$ such that $\gcd(k,n+1)=1$. I will also discuss related motives attached to Airy moments constructed by Sabbah and Yu. In particular, the computation shows that there are Hodge numbers bigger than $1$ in most cases.

math.AG

Optimal integrating learning for split questionnaire design type data

In the era of data science, it is common to encounter data with different subsets of variables obtained for different cases. An example is the split questionnaire design (SQD), which is adopted to reduce respondent fatigue and improve response rates by assigning different subsets of the questionnaire to different sampled respondents. A general question then is how to estimate the regression function based on such block-wise observed data. Currently, this is often carried out with the aid of missing data methods, which may unfortunately suffer intensive computational cost, high variability, and possible large modeling biases in real applications. In this article, we develop a novel approach for estimating the regression function for SQD-type data. We first construct a list of candidate models using available data-blocks separately, and then combine the estimates properly to make an efficient use of all the information. We show the resulting averaged model is asymptotically optimal in the sense that the squared loss and risk are asymptotically equivalent to those of the best but infeasible averaged estimator. Both simulated examples and an application to the SQD dataset from the European Social Survey show the promise of the proposed method.

stat.ME

Testing for Treatment Effect in Covariate-Adaptive Randomized Clinical Trials with Generalized Linear Models and Omitted Covariates

Concerns have been expressed over the validity of statistical inference under covariate-adaptive randomization despite the extensive use in clinical trials. In the literature, the inferential properties under covariate-adaptive randomization have been mainly studied for continuous responses; in particular, it is well known that the usual two sample t-test for treatment effect is typically conservative, in the sense that the actual test size is smaller than the nominal level. This phenomenon of invalid tests has also been found for generalized linear models without adjusting for the covariates and are sometimes more worrisome due to inflated Type I error. The purpose of this study is to examine the unadjusted test for treatment effect under generalized linear models and covariate-adaptive randomization. For a large class of covariate-adaptive randomization methods, we obtain the asymptotic distribution of the test statistic under the null hypothesis and derive the conditions under which the test is conservative, valid, or anti-conservative. Several commonly used generalized linear models, such as logistic regression and Poisson regression, are discussed in detail. An adjustment method is also proposed to achieve a valid size based on the asymptotic results. Numerical studies confirm the theoretical findings and demonstrate the effectiveness of the proposed adjustment method.

stat.ME

LqRT: Robust Hypothesis Testing of Location Parameters using Lq-Likelihood-Ratio-Type Test in Python

A t-test is considered a standard procedure for inference on population means and is widely used in scientific discovery. However, as a special case of a likelihood-ratio test, t-test often shows drastic performance degradation due to the deviations from its hard-to-verify distributional assumptions. Alternatively, in this article, we propose a new two-sample Lq-likelihood-ratio-type test (LqRT) along with an easy-to-use Python package for implementation. LqRT preserves high power when the distributional assumption is violated, and maintains the satisfactory performance when the assumption is valid. As numerical studies suggest, LqRT dominates many other robust tests in power, such as Wilcoxon test and sign test, while maintaining a valid size. To the extent that the robustness of the Wilcoxon test (minimum asymptotic relative efficiency (ARE) of the Wilcoxon test vs the t-test is 0.864) suggests that the Wilcoxon test should be the default test of choice (rather than "use Wilcoxon if there is evidence of non-normality", the default position should be "use Wilcoxon unless there is good reason to believe the normality assumption"), the results in this article suggest that the LqRT is potentially the new default go-to test for practitioners.

stat.ME

Statistical Inference for Covariate-Adaptive Randomization Procedures

Covariate-adaptive randomization (CAR) procedures are frequently used in comparative studies to increase the covariate balance across treatment groups. However, because randomization inevitably uses the covariate information when forming balanced treatment groups, the validity of classical statistical methods after such randomization is often unclear. In this article, we derive the theoretical properties of statistical methods based on general CAR under the linear model framework. More importantly, we explicitly unveil the relationship between covariate-adaptive and inference properties by deriving the asymptotic representations of the corresponding estimators. We apply the proposed general theory to various randomization procedures such as complete randomization, rerandomization, pairwise sequential randomization, and Atkinson's $D_A$-biased coin design and compare their performance analytically. Based on the theoretical results, we then propose a new approach to obtain valid and more powerful tests. These results open a door to understand and analyze experiments based on CAR. Simulation studies provide further evidence of the advantages of the proposed framework and the theoretical results. Supplementary materials for this article are available online.

math.ST

Statistical inference on random dot product graphs: a survey

The random dot product graph (RDPG) is an independent-edge random graph that is analytically tractable and, simultaneously, either encompasses or can successfully approximate a wide range of random graphs, from relatively simple stochastic block models to complex latent position graphs. In this survey paper, we describe a comprehensive paradigm for statistical inference on random dot product graphs, a paradigm centered on spectral embeddings of adjacency and Laplacian matrices. We examine the analogues, in graph inference, of several canonical tenets of classical Euclidean inference: in particular, we summarize a body of existing results on the consistency and asymptotic normality of the adjacency and Laplacian spectral embeddings, and the role these spectral embeddings can play in the construction of single- and multi-sample hypothesis tests for graph data. We investigate several real-world applications, including community detection and classification in large social networks and the determination of functional and biologically relevant network properties from an exploratory data analysis of the Drosophila connectome. We outline requisite background and current open problems in spectral graph inference.

stat.ME