SearcharxivSearch

arXiv subjects

Tailen Hsing

Publications and source records attributed to Tailen Hsing.

15 recordsLinked to original sources

Function-On-Function Regression Through Separable Neural Operators

This paper investigates the estimation of the regression operator in function-on-function regression models. While traditional research has predominantly focused on linear models or their immediate nonlinear extensions, we propose a neural operator approach to accommodate general regression operators under mild smoothness assumptions. Operator learning has emerged as an active area of machine learning, particularly for solving physical models governed by partial differential equations. Using this paradigm, our methodology introduces the separable neural operator, a neural-operator architecture that represents the regression operator through input-dependent coefficient functions and output-dependent basis functions. Beyond adapting this architecture to the regression operator estimation problem, we establish the consistency of the estimator under relatively mild smoothness and sampling conditions, allowing functional data to be observed on dense, possibly irregular, discrete grids. We also apply the proposed approach to the BGC Argo data and demonstrate its potential for oceanographic research.

math.ST

funFMC: Overlapping Clustering for Functional Data

In applications such as neuroscience and environmental science, data are naturally modeled as multivariate functional data and often exhibit overlapping cluster structure. Existing clustering methods for functional data typically impose mutually exclusive memberships and therefore fail to capture such structure. We propose a latent factor model based approach with functional factors and a real-valued loading matrix that encodes potentially overlapping cluster memberships. Under mild conditions, we establish identifiability of the loading matrix up to permutation, ensuring that the overlapping cluster structure is recoverable up to label switching. We develop a procedure for estimating both the number of clusters and the associated cluster memberships. This involves solving an infinite-dimensional regression problem in operators, whose solution is characterized using the inner product on the space of Hilbert-Schmidt operators and expressed in terms of real-valued matrices. This formulation enables rigorous asymptotic analysis, and we establish a central limit theorem to facilitate statistical inference on overlapping cluster memberships. We demonstrate the performance of our method using numerical studies and an application to functional magnetic resonance imaging data.

stat.ME

Generative and Nonparametric Approaches for Conditional Distribution Estimation: Methods, Perspectives, and Comparative Evaluations

The inference of conditional distributions is a fundamental problem in statistics, essential for prediction, uncertainty quantification, and probabilistic modeling. A wide range of methodologies have been developed for this task. This article reviews and compares several representative approaches spanning classical nonparametric methods and modern generative models. We begin with the single-index method of Hall and Yao (2005), which estimates the conditional distribution through a dimension-reducing index and nonparametric smoothing of the resulting one-dimensional cumulative conditional distribution function. We then examine the basis-expansion approaches, including FlexCode (Izbicki and Lee, 2017) and DeepCDE (Dalmasso et al., 2020), which convert conditional density estimation into a set of nonparametric regression problems. In addition, we discuss two recent generative simulation-based methods that leverage modern deep generative architectures: the generative conditional distribution sampler (Zhou et al., 2023) and the conditional denoising diffusion probabilistic model (Fu et al., 2024; Yang et al., 2025). A systematic numerical comparison of these approaches is provided using a unified evaluation framework that ensures fairness and reproducibility. The performance metrics used for the estimated conditional distribution include the mean-squared errors of conditional mean and standard deviation, as well as the Wasserstein distance. We also discuss their flexibility and computational costs, highlighting the distinct advantages and limitations of each approach.

stat.ML

Variable Selection and Minimax Prediction in High-dimensional Functional Linear Model

High-dimensional functional data have become increasingly prevalent in modern applications such as high-frequency financial data and neuroimaging data analysis. We investigate a class of high-dimensional linear regression models, where each predictor is a random element in an infinite-dimensional function space, and the number of functional predictors p can potentially be ultra-high. Assuming that each of the unknown coefficient functions belongs to some reproducing kernel Hilbert space (RKHS), we regularize the fitting of the model by imposing a group elastic-net type of penalty on the RKHS norms of the coefficient functions. We show that our loss function is Gateaux sub-differentiable, and our functional elastic-net estimator exists uniquely in the product RKHS. Under suitable sparsity assumptions and a functional version of the irrepresentable condition, we derive a non-asymptotic tail bound for variable selection consistency of our method. Allowing the number of true functional predictors $q$ to diverge with the sample size, we also show a post-selection refined estimator can achieve the oracle minimax optimal prediction rate. The proposed methods are illustrated through simulation studies and a real-data application from the Human Connectome Project.

stat.ME

Multivariate Mat\'ern Models -- A Spectral Approach

The classical Mat\'ern model has been a staple in spatial statistics. Novel data-rich applications in environmental and physical sciences, however, call for new, flexible vector-valued spatial and space-time models. Therefore, the extension of the classical Mat\'ern model has been a problem of active theoretical and methodological interest. In this paper, we offer a new perspective to extending the Mat\'ern covariance model to the vector-valued setting. We adopt a spectral, stochastic integral approach, which allows us to address challenging issues on the validity of the covariance structure and at the same time to obtain new, flexible, and interpretable models. In particular, our multivariate extensions of the Mat\'ern model allow for asymmetric covariance structures. Moreover, the spectral approach provides an essentially complete flexibility in modeling the local structure of the process. We establish closed-form representations of the cross-covariances when available, compare them with existing models, simulate Gaussian instances of these new processes, and demonstrate estimation of the model's parameters through maximum likelihood. An application of the new class of multivariate Mat\'ern models to environmental data indicate their success in capturing inherent covariance-asymmetry phenomena.

stat.ME

Spectral Density Estimation of Function-Valued Spatial Processes

The spectral density function describes the second-order properties of a stationary stochastic process on $\mathbb{R}^d$. This paper considers the nonparametric estimation of the spectral density of a continuous-time stochastic process taking values in a separable Hilbert space. Our estimator is based on kernel smoothing and can be applied to a wide variety of spatial sampling schemes including those in which data are observed at irregular spatial locations. Thus, it finds immediate applications in Spatial Statistics, where irregularly sampled data naturally arise. The rates for the bias and variance of the estimator are obtained under general conditions in a mixed-domain asymptotic setting. When the data are observed on a regular grid, the optimal rate of the estimator matches the minimax rate for the class of covariance functions that decay according to a power law. The asymptotic normality of the spectral density estimator is also established under general conditions for Gaussian Hilbert-space valued processes. Finally, with a view towards practical applications the asymptotic results are specialized to the case of discretely-sampled functional data in a reproducing kernel Hilbert space.

math.ST

Large Deviation Probabilities for Sums of Random Variables with Heavy or Subexponential Tails

Let $S_n$ be the sum of independent random variables with distribution $F$. Under the assumption that $-\log(1-F(x))$ is slowly varying, conditions for $$ \lim_{n\to\infty}\sup_{s\ge t_n}\left|{P[S_n>s]\over n(1-F(s))}-1\right| =0 $$ are given. These conditions extend and strengthen a series of previous results. Additionally, a connection with subexponential distributions is demonstrated. That is, $F$ is subexponential if and only if the condition above holds for some $t_n$ and $$ \lim_{t\to\infty}{1-F(t+x)\over 1-F(t)} = 1 \quad\text{for each real $x$.}$$

math.PR

A functional regression model for heterogeneous BioGeoChemical Argo data in the Southern Ocean

Leveraging available measurements of our environment can help us understand complex processes. One example is Argo Biogeochemical data, which aims to collect measurements of oxygen, nitrate, pH, and other variables at varying depths in the ocean. We focus on the oxygen data in the Southern Ocean, which has implications for ocean biology and the Earth's carbon cycle. Systematic monitoring of such data has only recently begun to be established, and the data is sparse. In contrast, Argo measurements of temperature and salinity are much more abundant. In this work, we introduce and estimate a functional regression model describing dependence in oxygen, temperature, and salinity data at all depths covered by the Argo data simultaneously. Our model elucidates important aspects of the joint distribution of temperature, salinity, and oxygen. Due to fronts that establish distinct spatial zones in the Southern Ocean, we augment this functional regression model with a mixture component. By modelling spatial dependence in the mixture component and in the data itself, we provide predictions onto a grid and improve location estimates of fronts. Our approach is scalable to the size of the Argo data, and we demonstrate its success in cross-validation and a comprehensive interpretation of the model.

stat.ME

Tangent fields, intrinsic stationarity, and self-similarity (with a supplement on Matheron Theory)

This paper studies the local structure of continuous random fields on $\mathbb R^d$ taking values in a complete separable linear metric space ${\mathbb V}$. Extending seminal work of Falconer, we show that the generalized $(1+k)$-th order increment tangent fields are self-similar and almost everywhere intrinsically stationary in the sense of Matheron. These results motivate the further study of the structure of ${\mathbb V}$-valued intrinsic random functions of order $k$ (IRF$_k$,\ $k=0,1,\cdots$). To this end, we focus on the special case where ${\mathbb V}$ is a Hilbert space. Building on the work of Sasvari and Berschneider, we establish the spectral characterization of all second order ${\mathbb V}$-valued IRF$_k$'s, extending the classical Matheron theory. Using these results, we further characterize the class of Gaussian, operator self-similar ${\mathbb V}$-valued IRF$_k$'s, generalizing results of Dobrushin and Didier, Meerschaert and Pipiras, among others. These processes are the Hilbert-space-valued versions of the general $k$-th order operator fractional Brownian fields and are characterized by their self-similarity operator exponent as well as a finite trace class operator valued spectral measure. We conclude with several examples motivating future applications to probability and statistics. In a technical Supplement of independent interest, we provide a unified treatment of the Matheron spectral theory for second-order stationary and intrinsically stationary processes taking values in a separable Hilbert space. We give the proofs of the Bochner-Neeb and Bochner-Schwartz theorems.

math.PR

A functional-data approach to the Argo data

The Argo data is a modern oceanography dataset that provides unprecedented global coverage of temperature and salinity measurements in the upper 2,000 meters of depth of the ocean. We study the Argo data from the perspective of functional data analysis (FDA). We develop spatio-temporal functional kriging methodology for mean and covariance estimation to predict temperature and salinity at a fixed location as a smooth function of depth. By combining tools from FDA and spatial statistics, including smoothing splines, local regression, and multivariate spatial modeling and prediction, our approach provides advantages over current methodology that consider pointwise estimation at fixed depths. Our approach naturally leverages the irregularly-sampled data in space, time, and depth to fit a space-time functional model for temperature and salinity. The developed framework provides new tools to address fundamental scientific problems involving the entire upper water column of the oceans such as the estimation of ocean heat content, stratification, and thermohaline oscillation. For example, we show that our functional approach yields more accurate ocean heat content estimates than ones based on discrete integral approximations in pressure. Further, using the derivative function estimates, we obtain a new product of a global map of the mixed layer depth, a key component in the study of heat absorption and nutrient circulation in the oceans. The derivative estimates also reveal evidence for density inversions in areas distinguished by mixing of particularly different water masses.

stat.AP

Uniform convergence rates for nonparametric regression and principal component analysis in functional/longitudinal data

We consider nonparametric estimation of the mean and covariance functions for functional/longitudinal data. Strong uniform convergence rates are developed for estimators that are local-linear smoothers. Our results are obtained in a unified framework in which the number of observations within each curve/cluster can be of any rate relative to the sample size. We show that the convergence rates for the procedures depend on both the number of sample curves and the number of observations on each curve. For sparse functional data, these rates are equivalent to the optimal rates in nonparametric regression. For dense functional data, root-n rates of convergence can be achieved with proper choices of bandwidths. We further derive almost sure rates of convergence for principal component analysis using the estimated covariance function. The results are illustrated with simulation studies.

math.ST

Deciding the dimension of effective dimension reduction space for functional and high-dimensional data

In this paper, we consider regression models with a Hilbert-space-valued predictor and a scalar response, where the response depends on the predictor only through a finite number of projections. The linear subspace spanned by these projections is called the effective dimension reduction (EDR) space. To determine the dimensionality of the EDR space, we focus on the leading principal component scores of the predictor, and propose two sequential $χ^2$ testing procedures under the assumption that the predictor has an elliptically contoured distribution. We further extend these procedures and introduce a test that simultaneously takes into account a large number of principal component scores. The proposed procedures are supported by theory, validated by simulation studies, and illustrated by a real-data example. Our methods and theory are applicable to functional data and high-dimensional multivariate data.

math.ST

An RKHS formulation of the inverse regression dimension-reduction problem

Suppose that $Y$ is a scalar and $X$ is a second-order stochastic process, where $Y$ and $X$ are conditionally independent given the random variables $ξ_1,...,ξ_p$ which belong to the closed span $L_X^2$ of $X$. This paper investigates a unified framework for the inverse regression dimension-reduction problem. It is found that the identification of $L_X^2$ with the reproducing kernel Hilbert space of $X$ provides a platform for a seamless extension from the finite- to infinite-dimensional settings. It also facilitates convenient computational algorithms that can be applied to a variety of models.

math.ST

Extremes on Trees

This paper considers the asymptotic distribution of the longest edge of the minimal spanning tree and nearest neighbor graph on X_1,...,X_{N_n} where X_1,X_2,... are i.i.d. in \Re^2 with distribution F and N_n is independent of the X_i and satisfies N_n/n\to_p1. A new approach based on spatial blocking and a locally orthogonal coordinate system is developed to treat cases for which F has unbounded support. The general results are applied to a number of special cases, including elliptically contoured distributions, distributions with independent Weibull-like margins and distributions with parallel level curves.

math.PR

On weighted U-statistics for stationary processes

A weighted U-statistic based on a random sample X_1,...,X_n has the form U_n=\sum_{1\le i,j\le n}w_{i-j}K(X_i,X_j), where K is a fixed symmetric measurable function and the w_i are symmetric weights. A large class of statistics can be expressed as weighted U-statistics or variations thereof. This paper establishes the asymptotic normality of U_n when the sample observations come from a nonlinear time series and linear processes.

math.PR