SearcharxivSearch

arXiv subjects

F. Jay Breidt

Publications and source records attributed to F. Jay Breidt.

6 recordsLinked to original sources

A flexible class of priors for orthonormal matrices with basis function-specific structure

Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system's internal variability.

stat.ME

Uniform convergence of the empirical cumulative distribution function under informative selection from a finite population

Consider informative selection of a sample from a finite population. Responses are realized as independent and identically distributed (i.i.d.) random variables with a probability density function (p.d.f.) f, referred to as the superpopulation model. The selection is informative in the sense that the sample responses, given that they were selected, are not i.i.d. f. In general, the informative selection mechanism may induce dependence among the selected observations. The impact of such dependence on the empirical cumulative distribution function (c.d.f.) is studied. An asymptotic framework and weak conditions on the informative selection mechanism are developed under which the (unweighted) empirical c.d.f. converges uniformly, in $L_2$ and almost surely, to a weighted version of the superpopulation c.d.f. This yields an analogue of the Glivenko-Cantelli theorem. A series of examples, motivated by real problems in surveys and other observational studies, shows that the conditions are verifiable for specified designs.

math.ST

Endogenous post-stratification in surveys: classifying with a sample-fitted model

Post-stratification is frequently used to improve the precision of survey estimators when categorical auxiliary information is available from sources outside the survey. In natural resource surveys, such information is often obtained from remote sensing data, classified into categories and displayed as pixel-based maps. These maps may be constructed based on classification models fitted to the sample data. Post-stratification of the sample data based on categories derived from the sample data (``endogenous post-stratification'') violates the standard post-stratification assumptions that observations are classified without error into post-strata, and post-stratum population counts are known. Properties of the endogenous post-stratification estimator are derived for the case of a sample-fitted generalized linear model, from which the post-strata are constructed by dividing the range of the model predictions into predetermined intervals. Design consistency of the endogenous post-stratification estimator is established under mild conditions. Under a superpopulation model, consistency and asymptotic normality of the endogenous post-stratification estimator are established, showing that it has the same asymptotic variance as the traditional post-stratified estimator with fixed strata. Simulation experiments demonstrate that the practical effect of first fitting a model to the survey data before post-stratifying is small, even for relatively small sample sizes.

math.ST

Rank-based estimation for all-pass time series models

An autoregressive-moving average model in which all roots of the autoregressive polynomial are reciprocals of roots of the moving average polynomial and vice versa is called an all-pass time series model. All-pass models are useful for identifying and modeling noncausal and noninvertible autoregressive-moving average processes. We establish asymptotic normality and consistency for rank-based estimators of all-pass model parameters. The estimators are obtained by minimizing the rank-based residual dispersion function given by Jaeckel [Ann. Math. Statist. 43 (1972) 1449--1458]. These estimators can have the same asymptotic efficiency as maximum likelihood estimators and are robust. The behavior of the estimators for finite samples is studied via simulation and rank estimation is used in the deconvolution of a simulated water gun seismogram.

math.ST

Pile-up probabilities for the Laplace likelihood estimator of a non-invertible first order moving average

The first-order moving average model or MA(1) is given by $X_t=Z_t-θ_0Z_{t-1}$, with independent and identically distributed $\{Z_t\}$. This is arguably the simplest time series model that one can write down. The MA(1) with unit root ($θ_0=1$) arises naturally in a variety of time series applications. For example, if an underlying time series consists of a linear trend plus white noise errors, then the differenced series is an MA(1) with unit root. In such cases, testing for a unit root of the differenced series is equivalent to testing the adequacy of the trend plus noise model. The unit root problem also arises naturally in a signal plus noise model in which the signal is modeled as a random walk. The differenced series follows a MA(1) model and has a unit root if and only if the random walk signal is in fact a constant. The asymptotic theory of various estimators based on Gaussian likelihood has been developed for the unit root case and nearly unit root case ($θ=1+β/n,β\le0$). Unlike standard $1/\sqrt{n}$-asymptotics, these estimation procedures have $1/n$-asymptotics and a so-called pile-up effect, in which P$(\hatθ=1)$ converges to a positive value. One explanation for this pile-up phenomenon is the lack of identifiability of $θ$ in the Gaussian case. That is, the Gaussian likelihood has the same value for the two sets of parameter values $(θ,σ^2)$ and $(1/θ,θ^2σ^2$). It follows that $θ=1$ is always a critical point of the likelihood function. In contrast, for non-Gaussian noise, $θ$ is identifiable for all real values. Hence it is no longer clear whether or not the same pile-up phenomenon will persist in the non-Gaussian case. In this paper, we focus on limiting pile-up probabilities for estimates of $θ_0$ based on a Laplace likelihood. In some cases, these estimates can be viewed as Least Absolute Deviation (LAD) estimates. Simulation results illustrate the limit theory.

math.ST