SearcharxivSearch

arXiv subjects

Fang Yao

Publications and source records attributed to Fang Yao.

At least 19 recordsLinked to original sources

Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

Functional data analysis is an important statistical field that treats data as random functions. In practice, the random functions are often not fully observed but instead measured at discrete times. While simpler problems, such as mean and covariance estimation, have been widely studied for discretely observed data, optimal estimation of linear regression for this data type has remained unsolved for over two decades. To tackle this fundamental challenge, we propose a novel approach, referred to as pooling ridge estimation, which combines the advantages of pooling strategy and RKHS-based method by incorporating the unbiased estimation of operators based on discretely observed measurements from all subjects. This unified estimation framework enables us to achieve minimax optimality in prediction risk in arbitrary sampling schemes ranging from sparse to dense designs, for both scalar-on-function and function-on-function regression models. Such methodological and theoretical advances are obtained for the first time and accurately reveal the influence of discrete sampling. For scalar-on-function regression, the phase transition occurs once, separating the convergence behavior into two distinct regimes. Remarkably, for function-on-function regression, up to three phase transitions may occur, determined by the sampling frequencies of the predictor/response functions. Finally, simulation experiments and two real data examples provide empirical support for the proposed methods.

stat.ME

Structural Change Detection in Dynamic Systems

Structural changes often arise in real-world dynamic systems due to external interventions or environmental shifts, such as policy changes in epidemiology or climate forcing in environmental science. In this paper, we propose a unified framework for detecting and localizing structural changes in dynamic systems governed by ordinary differential equations. Unlike existing methods that assume mean or linear trend changes, our approach accommodates complex, nonlinear dynamics with both stable and diverging trajectories. We develop a new test statistic that combines residual-based discrepancy and normalized parameter contrast, capturing evidence for structural changes from both model fit and parameter shifts. Candidate structural changes are efficiently screened using a multiscale seeded-narrowest-over-threshold algorithm with a data-driven thresholding strategy. To refine selections and control false discoveries, we introduce a false discovery rate control procedure that leverages order-preserved sample splitting and symmetric contrast calibration. Theoretical guarantees are established, including detection consistency, near-minimax localization accuracy, and valid FDR control under weak dependence. Extensive simulations demonstrate superior performance over existing methods in both accuracy and FDR control. Applications to real-world data sets, including COVID-19 dynamics and global temperature trends, highlight the practical relevance and broad applicability of our method.

stat.ME

Deconfounding via Profiled Transfer Learning

Unmeasured confounders are a major source of bias in regression-based effect estimation and causal inference. In this paper, we propose a new profiled transfer learning framework, ProTrans, to address confounding effects in the target dataset, when additional source datasets with similar confounding structures are available. We introduce the concept of profiled residuals to characterize the shared confounding patterns between source and target datasets. By incorporating these profiled residuals into the target debiasing step, we effectively mitigate the latent confounding effects. We also propose a source selection strategy to enhance the robustness of ProTrans to noninformative sources. As a byproduct, ProTrans can also be used to estimate treatment effects in the presence of potential confounders, without the use of auxiliary features such as instrumental or proxy variables, which are often challenging to select in practice. Theoretically, we prove that the resulting estimated model shift from the sources to the target is confounding-free without imposing specific assumptions on the true confounding structure, and that the target parameter estimation achieves the minimax optimal rate under mild conditions. Simulated and real-world experiments validate the effectiveness of ProTrans and support the theoretical findings.

stat.ME

Online Sparse Regression with Expanding Observables

Online high-dimensional regression has gained increasing attention in recent years, yet existing methods typically assume that all candidate features, including important ones, are observed from the outset of data collection. This assumption is often violated in real-world scenarios, where new variables become available gradually as data accumulate. To address this gap, we introduce a novel framework, Recurrent Adaptive Variable Selection (RAVAS), for online regression with expanding observability. RAVAS employs a recurrent procedure that dynamically updates feature selection as both the sample size and the observable feature set grow. The algorithm is designed to be computationally efficient and memory-light, relying only on low-dimensional sufficient statistics that are updated online. A key advantage of the method lies in its ability to detect and incorporate important variables that emerge later, thereby mitigating the effect of early-stage missingness. We establish theoretical guarantees on model selection, estimation error, and feature coverage, and develop an adaptive online tuning strategy. Extensive simulations and real-world experiments verify the effectiveness of RAVAS for high-dimensional streaming data.

math.ST

Covariance test for discretely observed functional data: when and how it works?

For covariance test in functional data analysis, existing methods are developed only for fully observed curves, whereas in practice, trajectories are typically observed discretely and with noise. To bridge this gap, we employ a pool-smoothing strategy to construct an FPC-based test statistic, allowing the number of estimated eigenfunctions to grow with the sample size. This yields a consistently nonparametric test, while the challenge arises from the concurrence of diverging truncation and discretized observations. Facilitated by advancing perturbation bounds of estimated eigenfunctions, we establish that the asymptotic null distribution remains valid across permissable truncation levels. Moreover, when the sampling frequency (i.e., the number of measurements per subject) reaches certain magnitude of sample size, the test behaves as if the functions were fully observed. This phase transition phenomenon differs from the well-known result of the pooling mean/covariance estimation, reflecting the elevated difficulty in covariance test due to eigen-decomposition. The numerical studies, including simulations and real data examples, yield favorable performance compared to existing methods.

stat.ME

Dynamic Matrix Recovery

Matrix recovery from sparse observations is an extensively studied topic emerging in various applications, such as recommendation system and signal processing, which includes the matrix completion and compressed sensing models as special cases. In this work we propose a general framework for dynamic matrix recovery of low-rank matrices that evolve smoothly over time. We start from the setting that the observations are independent across time, then extend to the setting that both the design matrix and noise possess certain temporal correlation via modified concentration inequalities. By pooling neighboring observations, we obtain sharp estimation error bounds of both settings, showing the influence of the underlying smoothness, the dependence and effective samples. We propose a dynamic fast iterative shrinkage thresholding algorithm that is computationally efficient, and characterize the interplay between algorithmic and statistical convergence. Simulated and real data examples are provided to support such findings.

stat.ME

Nonparametric Prior Learning in Differential Equation Modeling

This paper addresses Bayesian inference related to partial differential equations (PDEs), particularly nonparametric regression constrained by PDEs. To effectively encode prior information, we propose a novel framework that learns a prediction function of the prior distribution from historical training datasets. We introduce hyper-prior and hyper-posterior distributions and derive a generalization error estimate, which accommodates data-dependent priors by extending the concept of differential privacy. Some mild conditions are given to validate the error estimate, where various typical PDEs such as diffusion and Darcy flow equations can be integrated. We thus formulate an infinite-dimensional optimization problem to obtain the point estimate of the hyper-posterior. Numerical examples demonstrate the performance of our proposed method in learning the prediction function of priors.

math.ST

Green's matching: an efficient approach to parameter estimation in complex dynamic systems

Parameters of differential equations are essential to characterize intrinsic behaviors of dynamic systems. Numerous methods for estimating parameters in dynamic systems are computationally and/or statistically inadequate, especially for complex systems with general-order differential operators, such as motion dynamics. This article presents Green's matching, a computationally tractable and statistically efficient two-step method, which only needs to approximate trajectories in dynamic systems but not their derivatives due to the inverse of differential operators by Green's function. This yields a statistically optimal guarantee for parameter estimation in general-order equations, a feature not shared by existing methods, and provides an efficient framework for broad statistical inferences in complex dynamic systems.

stat.ME

Spatially Randomized Designs Can Enhance Policy Evaluation

This article studies the benefits of using spatially randomized experimental designs which partition the experimental area into distinct, non-overlapping units with treatments assigned randomly. Such designs offer improved policy evaluation in online experiments by providing more precise policy value estimators and more effective A/B testing algorithms than traditional global designs, which apply the same treatment across all units simultaneously. We examine both parametric and nonparametric methods for estimating and inferring policy values based on this randomized approach. Our analysis includes evaluating the mean squared error of the treatment effect estimator and the statistical power of the associated tests. Additionally, we extend our findings to experiments with spatio-temporal dependencies, where treatments are allocated sequentially over time, and account for potential temporal carryover effects. Our theoretical insights are supported by comprehensive numerical experiments.

math.ST

Deep Semiparametric Partial Differential Equation Models

In many scientific fields, the generation and evolution of data are governed by partial differential equations (PDEs) which are typically informed by established physical laws at the macroscopic level to describe general and predictable dynamics. However, some complex influences may not be fully captured by these laws at the microscopic level due to limited scientific understanding. This work proposes a unified framework to model, estimate, and infer the mechanisms underlying data dynamics. We introduce a general semiparametric PDE (SemiPDE) model that combines interpretable mechanisms based on physical laws with flexible data-driven components to account for unknown effects. The physical mechanisms enhance the SemiPDE model's stability and interpretability, while the data-driven components improve adaptivity to complex real-world scenarios. A deep profiling M-estimation approach is proposed to decouple the solutions of PDEs in the estimation procedure, leveraging both the accuracy of numerical methods for solving PDEs and the expressive power of neural networks. For the first time, we establish a semiparametric inference method and theory for deep M-estimation, considering both training dynamics and complex PDE models. We analyze how the PDE structure affects the convergence rate of the nonparametric estimator, and consequently, the parametric efficiency and inference procedure enable the identification of interpretable mechanisms governing data dynamics. Simulated and real-world examples demonstrate the effectiveness of the proposed methodology and support the theoretical findings.

stat.ME

A General Test for Independent and Identically Distributed Hypothesis

We propose a simple and intuitive test for arguably the most prevailing hypothesis in statistics that data are independent and identically distributed (IID), based on a newly introduced off-diagonal sequential U-process. This IID test is fully nonparametric and applicable to random objects in general spaces, while requiring no specific alternatives such as structural breaks or serial dependence, which allows for detecting general types of violations of the IID assumption. An easy-to-implement jackknife multiplier bootstrap is tailored to produce critical values of the test. Under mild conditions, we establish Gaussian approximation for the proposed U-processes, and derive non-asymptotic coupling and Kolmogorov distance bounds for its maximum and the bootstrapped version, providing rigorous theoretical guarantees. Simulations and real data applications are conducted to demonstrate the usefulness and versatility compared with existing methods.

stat.ME

Functional Tensor Regression

Tensor regression has attracted significant attention in statistical research. This study tackles the challenge of handling covariates with smooth varying structures. We introduce a novel framework, termed functional tensor regression, which incorporates both the tensor and functional aspects of the covariate. To address the high dimensionality and functional continuity of the regression coefficient, we employ a low Tucker rank decomposition along with smooth regularization for the functional mode. We develop a functional Riemannian Gauss--Newton algorithm that demonstrates a provable quadratic convergence rate, while the estimation error bound is based on the tensor covariate dimension. Simulations and a neuroimaging analysis illustrate the finite sample performance of the proposed method.

stat.ME

Function-on-function Differential Regression

Function-on-function regression has been a topic of substantial interest due to its broad applicability, where the relation between functional predictor and response is concerned. In this article, we propose a new framework for modeling the regression mapping that extends beyond integral type, motivated by the prevalence of physical phenomena governed by differential relations, which is referred to as function-on-function differential regression. However, a key challenge lies in representing the differential regression operator, unlike functions that can be expressed by expansions. As a main contribution, we introduce a new notion of model identification involving differential operators, defined through their action on functions. Based on this action-aware identification, we are able to develop a regularization method for estimation using operator reproducing kernel Hilbert spaces. Then a goodness-of-fit test is constructed, which facilitates model checking for differential regression relations. We establish a Bahadur representation for the regression estimator with various theoretical implications, such as the minimax optimality of the proposed estimator, and the validity and consistency of the proposed test. To illustrate the effectiveness of our method, we conduct simulation studies and an application to a real data example on the thermodynamic energy equation.

stat.ME

Semiparametric M-estimation with overparameterized neural networks

We focus on semiparametric regression that has played a central role in statistics, and exploit the powerful learning ability of deep neural networks (DNNs) while enabling statistical inference on parameters of interest that offers interpretability. Despite the success of classical semiparametric method/theory, establishing the $\sqrt{n}$-consistency and asymptotic normality of the finite-dimensional parameter estimator in this context remains challenging, mainly due to nonlinearity and potential tangent space degeneration in DNNs. In this work, we introduce a foundational framework for semiparametric $M$-estimation, leveraging the approximation ability of overparameterized neural networks that circumvent tangent degeneration and align better with training practice nowadays. The optimization properties of general loss functions are analyzed, and the global convergence is guaranteed. Instead of studying the ``ideal'' solution to minimization of an objective function in most literature, we analyze the statistical properties of algorithmic estimators, and establish nonparametric convergence and parametric asymptotic normality for a broad class of loss functions. These results hold without assuming the boundedness of the network output and even when the true function lies outside the specified function space. To illustrate the applicability of the framework, we also provide examples from regression and classification, and the numerical experiments provide empirical support to the theoretical findings.

math.ST

Universal Bootstrap for Spectral Statistics: Beyond Gaussian Approximation

Spectral analysis plays a crucial role in high-dimensional statistics, where determining the asymptotic distribution of various spectral statistics remains a challenging task. Due to the difficulties of deriving the analytic form, recent advances have explored data-driven bootstrap methods for this purpose. However, widely used Gaussian approximation-based bootstrap methods, such as the empirical bootstrap and multiplier bootstrap, have been shown to be inconsistent in approximating the distributions of spectral statistics in high-dimensional settings. To address this issue, we propose a universal bootstrap procedure based on the concept of universality from random matrix theory. Our method consistently approximates a broad class of spectral statistics across both high- and ultra-high-dimensional regimes, accommodating scenarios where the dimension-to-sample-size ratio $p/n$ converges to a nonzero constant or diverges to infinity without requiring structural assumptions on the population covariance matrix, such as eigenvalue decay or low effective rank. We showcase this universal bootstrap method for high-dimensional covariance inference. Extensive simulations and a real-world data study support our findings, highlighting the favorable finite sample performance of the proposed universal bootstrap procedure.

math.ST

Deep Regression for Repeated Measurements

Nonparametric mean function regression with repeated measurements serves as a cornerstone for many statistical branches, such as longitudinal/panel/functional data analysis. In this work, we investigate this problem using fully connected deep neural network (DNN) estimators with flexible shapes. A novel theoretical framework allowing arbitrary sampling frequency is established by adopting empirical process techniques to tackle clustered dependence. We then consider the DNN estimators for Hölder target function and illustrate a key phenomenon, the phase transition in the convergence rate, inherent to repeated measurements and its connection to the curse of dimensionality. Furthermore, we study several examples with low intrinsic dimensions, including the hierarchical composition model, low-dimensional support set and anisotropic Hölder smoothness. We also obtain new approximation results and matching lower bounds to demonstrate the adaptivity of the DNN estimators for circumventing the curse of dimensionality. Simulations and real data examples are provided to support our theoretical findings and practical implications.

math.ST

Computationally Efficient Whole-Genome Signal Region Detection for Quantitative and Binary Traits

The identification of genetic signal regions in the human genome is critical for understanding the genetic architecture of complex traits and diseases. Numerous methods based on scan algorithms (i.e. QSCAN, SCANG, SCANG-STARR) have been developed to allow dynamic window sizes in whole-genome association studies. Beyond scan algorithms, we have recently developed the binary and re-search (BiRS) algorithm, which is more computationally efficient than scan-based methods and exhibits superior statistical power. However, the BiRS algorithm is based on two-sample mean test for binary traits, not accounting for multidimensional covariates or handling test statistics for non-binary outcomes. In this work, we present a distributed version of the BiRS algorithm (dBiRS) that incorporate a new infinity-norm test statistic based on summary statistics computed from a generalized linear model. The dBiRS algorithm accommodates regression-based statistics, allowing for the adjustment of covariates and the testing of both continuous and binary outcomes. This new framework enables parallel computing of block-wise results by aggregation through a central machine to ensure both detection accuracy and computational efficiency, and has theoretical guarantees for controlling family-wise error rates and false discovery rates while maintaining the power advantages of the original algorithm. Applying dBiRS to detect genetic regions associated with fluid intelligence and prospective memory using whole-exome sequencing data from the UK Biobank, we validate previous findings and identify numerous novel rare variants near newly implicated genes. These discoveries offer valuable insights into the genetic basis of cognitive performance and neurodegenerative disorders, highlighting the potential of dBiRS as a scalable and powerful tool for whole-genome signal region detection.

stat.AP

Matrix Completion via Residual Spectral Matching

Noisy matrix completion has attracted significant attention due to its applications in recommendation systems, signal processing and image restoration. Most existing works rely on (weighted) least squares methods under various low-rank constraints. However, minimizing the sum of squared residuals is not always efficient, as it may ignore the potential structural information in the residuals. In this study, we propose a novel residual spectral matching criterion that incorporates not only the numerical but also locational information of residuals. This criterion is the first in noisy matrix completion to adopt the perspective of low-rank perturbation of random matrices and exploit the spectral properties of sparse random matrices. We derive optimal statistical properties by analyzing the spectral properties of sparse random matrices and bounding the effects of low-rank perturbations and partial observations. Additionally, we propose algorithms that efficiently approximate solutions by constructing easily computable pseudo-gradients. The iterative process of the proposed algorithms ensures convergence at a rate consistent with the optimal statistical error bound. Our method and algorithms demonstrate improved numerical performance in both simulated and real data examples, particularly in environments with high noise levels.

stat.ML