SearcharxivSearch

arXiv subjects

Valentin Patilea

Publications and source records attributed to Valentin Patilea.

At least 19 recordsLinked to original sources

An information criterion for detecting periodicities in functional time series

We propose an information criterion for determining an unknown number of periodic components in functional time series. Identifying the number of frequencies in large-scale time series has been a central focus. To achieve this goal, we suggest an iterative procedure, utilizing the residual process obtained through least squares fitting. This iterative approach demonstrates broad applicability. We establish the consistency of the estimated number of periodic components by minimizing the information criterion. The efficacy of the procedure is illustrated through numerical simulations. In real data analysis, we apply this information criterion to temperature data and sunspot data.

stat.ME

Discrete-time Markov chains with random observation times

We propose a new approach for estimating the finite dimensional transition matrix of a Markov chain using a large number of independent sample paths observed at random times. The sample paths may be observed as few as two times, and the transitions are allowed to depend on covariates. Simple and easy to update kernel estimates are proposed, and their uniform convergence rates are derived. Simulation experiments show that our estimation approach performs well.

stat.ME

Optimal inference for the mean of random functions

We study estimation and inference for the mean of real-valued random functions defined on a hypercube. The independent random functions are observed on a discrete, random subset of design points, possibly with heteroscedastic noise. We propose a novel optimal-rate estimator based on Fourier series expansions and establish a sharp non-asymptotic error bound in $L^2-$norm. Additionally, we derive a non-asymptotic Gaussian approximation bound for our estimated Fourier coefficients. Pointwise and uniform confidence sets are constructed. Our approach is made adaptive by a plug-in estimator for the H\"older regularity of the mean function, for which we derive non-asymptotic concentration bounds.

math.ST

Continuously updated estimation of conditional hazard functions

Motivated by the need to analyze continuously updated data sets in the context of time-to-event modeling, we propose a novel nonparametric approach to estimate the conditional hazard function given a set of continuous and discrete predictors. The method is based on a representation of the conditional hazard as a ratio between a joint density and a conditional expectation determined by the distribution of the observed variables. It is shown that such ratio representations are available for uni- and bivariate time-to-events, in the presence of common types of random censoring, truncation, and with possibly cured individuals, as well as for competing risks. This opens the door to nonparametric approaches in many time-to-event predictive models. To estimate joint densities and conditional expectations we propose the recursive kernel smoothing, which is well suited for online estimation. Asymptotic results for such estimators are derived and it is shown that they achieve optimal convergence rates. Simulation experiments show the good finite sample performance of our recursive estimator with right censoring. The method is applied to a real dataset of primary breast cancer.

stat.ME

Rate accelerated inference for integrals of multivariate random functions

The computation of integrals is a fundamental task in the analysis of functional data, which are typically considered as random elements in a space of squared integrable functions. Borrowing ideas from recent advances in the Monte Carlo integration literature, we propose effective unbiased estimation and inference procedures for integrals of uni- and multivariate random functions. Several applications to key problems in functional data analysis involving random design points are studied and illustrated. In the absence of noise, the proposed estimates converge faster than the sample mean and the usual algorithms for numerical integration. Moreover, the proposed estimator facilitates effective inference by generally providing better coverage with shorter confidence and prediction intervals, in both noisy and noiseless setups.

stat.ME

Adaptive estimation for Weakly Dependent Functional Times Series

The local regularity of functional time series is studied under $L^p-m-$appro\-ximability assumptions. The sample paths are observed with error at possibly random design points. Non-asymptotic concentration bounds of the regularity estimators are derived. As an application, we build nonparametric mean and autocovariance functions estimators that adapt to the regularity and the design, which can be sparse or dense. We also derive the asymptotic normality of the mean estimator, which allows honest inference for irregular mean functions. Simulations and a real data application illustrate the performance of the new estimators.

math.ST

Learning the regularity of multivariate functional data

Combining information both within and between sample realizations, we propose a simple estimator for the local regularity of surfaces in the functional data framework. The independently generated surfaces are measured with errors at possibly random discrete times. Non-asymptotic exponential bounds for the concentration of the regularity estimators are derived. An indicator for anisotropy is proposed and an exponential bound of its risk is derived. Two applications are proposed. We first consider the class of multi-fractional, bi-dimensional, Brownian sheets with domain deformation, and study the nonparametric estimation of the deformation. As a second application, we build minimax optimal, bivariate kernel estimators for the reconstruction of the surfaces.

math.ST

Adaptive functional principal components analysis

Functional data analysis almost always involves smoothing discrete observations into curves, because they are never observed in continuous time and rarely without error. Although smoothing parameters affect the subsequent inference, data-driven methods for selecting these parameters are not well-developed, frustrated by the difficulty of using all the information shared by curves while being computationally efficient. On the one hand, smoothing individual curves in an isolated, albeit sophisticated way, ignores useful signals present in other curves. On the other hand, bandwidth selection by automatic procedures such as cross-validation after pooling all the curves together quickly become computationally unfeasible due to the large number of data points. In this paper we propose a new data-driven, adaptive kernel smoothing, specifically tailored for functional principal components analysis through the derivation of sharp, explicit risk bounds for the eigen-elements. The minimization of these quadratic risk bounds provide refined, yet computationally efficient bandwidth rules for each eigen-element separately. Both common and independent design cases are allowed. Rates of convergence for the estimators are derived. An extensive simulation study, designed in a versatile manner to closely mimic the characteristics of real data sets supports our methodological contribution. An illustration on a real data application is provided.

stat.ME

A 2-step estimation procedure for semiparametric mixture cure models

Cure models have been developed as an alternative modelling approach to conventional survival analysis in order to account for the presence of cured subjects that will never experience the event of interest. Mixture cure models, which model separately the cure probability and the survival of uncured subjects depending on a set of covariates, are particularly useful for distinguishing curative from life-prolonging effects. In practice, it is common to assume a parametric model for the cure probability and a semiparametric model for the survival of the susceptibles. Because of the latent cure status, maximum likelihood estimation is performed by means of the iterative EM algorithm. Here, we focus on the cure probabilities and propose a two-step procedure to improve upon the performance of the maximum likelihood estimator when the sample size is not large. The new method is based on the idea of presmoothing by first constructing a nonparametric estimator and then projecting it into the desired parametric class. We investigate the theoretical properties of the resulting estimator and show through an extensive simulation study for the logistic-Cox model that it outperforms the existing method. Practical use of the method is illustrated through two melanoma datasets.

stat.ME

Adaptive estimation of irregular mean and covariance functions

Nonparametric estimators for the mean and the covariance functions of functional data are proposed. The setup covers a wide range of practical situations. The random trajectories are, not necessarily differentiable, have unknown regularity, and are measured with error at discrete design points. The measurement error could be heteroscedastic. The design points could be either randomly drawn or common for all curves. The estimators depend on the local regularity of the stochastic process generating the functional data. We consider a simple estimator of this local regularity which exploits the replication and regularization features of functional data. Next, we use the ``smoothing first, then estimate'' approach for the mean and the covariance functions. They can be applied with both sparsely or densely sampled curves, are easy to calculate and to update, and perform well in simulations. Simulations built upon an example of real data set, illustrate the effectiveness of the new approach.

math.ST

Semiparametric inference for partially linear regressions with Box-Cox transformation

In this paper, a semiparametric partially linear model in the spirit of Robinson (1988) with Box- Cox transformed dependent variable is studied. Transformation regression models are widely used in applied econometrics to avoid misspecification. In addition, a partially linear semiparametric model is an intermediate strategy that tries to balance advantages and disadvantages of a fully parametric model and nonparametric models. A combination of transformation and partially linear semiparametric model is, thus, a natural strategy. The model parameters are estimated by a semiparametric extension of the so called smooth minimum distance (SmoothMD) approach proposed by Lavergne and Patilea (2013). SmoothMD is suitable for models defined by conditional moment conditions and allows the variance of the error terms to depend on the covariates. In addition, here we allow for infinite-dimension nuisance parameters. The asymptotic behavior of the new SmoothMD estimator is studied under general conditions and new inference methods are proposed. A simulation experiment illustrates the performance of the methods for finite samples.

econ.EM

Powers correlation analysis of non-stationary illiquid assets

In this paper, the higher order dynamics of individual illiquid stocks are investigated. We show that considering the classical powers correlation could lead to a spurious assessment of the volatility persistency or long memory volatility effects, if the zero returns probability is non-constant over time. In other words, the classical tools are not able to distinguish between long-run volatility effects, such as IGARCH, and the case where the zero returns are not evenly distributed over time. As a consequence, tools that are robust to changes in the degree of illiquidity are proposed. Since a time-varying zero returns probability could potentially be accompanied by a non-constant unconditional variance, we then develop powers correlations that are also robust in such a case. In addition, note that the tools proposed in the paper offer a rigorous analysis of the short-run volatility effects, while the use of the classical power correlations lead to doubtful conclusions. The Monte Carlo experiments, and the study of the absolute value correlation of some stocks taken from the Chilean financial market, suggest that the volatility effects are only short-run in many cases.

math.ST

Clustering multivariate functional data using unsupervised binary trees

We propose a model-based clustering algorithm for a general class of functional data for which the components could be curves or images. The random functional data realizations could be measured with error at discrete, and possibly random, points in the definition domain. The idea is to build a set of binary trees by recursive splitting of the observations. The number of groups are determined in a data-driven way. The new algorithm provides easily interpretable results and fast predictions for online data sets. Results on simulated datasets reveal good performance in various complex settings. The methodology is applied to the analysis of vehicle trajectories on a German roundabout.

stat.ML

Learning the smoothness of noisy curves with application to online curve estimation

Combining information both within and across trajectories, we propose a simple estimator for the local regularity of the trajectories of a stochastic process. Independent trajectories are measured with errors at randomly sampled time points. Non-asymptotic bounds for the concentration of the estimator are derived. Given the estimate of the local regularity, we build a nearly optimal local polynomial smoother from the curves from a new, possibly very large sample of noisy trajectories. We derive non-asymptotic pointwise risk bounds uniformly over the new set of curves. Our estimates perform well in simulations. Real data sets illustrate the effectiveness of the new approaches.

math.ST

A presmoothing approach for estimation in semiparametric mixture cure models

A challenge when dealing with survival analysis data is accounting for a cure fraction, meaning that some subjects will never experience the event of interest. Mixture cure models have been frequently used to estimate both the probability of being cured and the time to event for the susceptible subjects, by usually assuming a parametric (logistic) form of the incidence. We propose a new estimation procedure for a parametric cure rate that relies on a preliminary smooth estimator and is independent of the model assumed for the latency. We investigate the theoretical properties of the estimators and show through simulations that, in the logistic/Cox model, presmoothing leads to more accurate results compared to the maximum likelihood estimator. To illustrate the practical use, we apply the new estimation procedure to two studies of melanoma survival data.

stat.ME

Wilks' theorem for semiparametric regressions with weakly dependent data

The empirical likelihood inference is extended to a class of semiparametric models for stationary, weakly dependent series. A partially linear single-index regression is used for the conditional mean of the series given its past, and the present and past values of a vector of covariates. A parametric model for the conditional variance of the series is added to capture further nonlinear effects. We propose a fixed number of suitable moment equations which characterize the mean and variance model. We derive an empirical log-likelihood ratio which includes nonparametric estimators of several functions, and we show that this ratio has the same limit as in the case where these functions are known.

stat.ME

Orthogonal Impulse Response Analysis in Presence of Time-Varying Covariance

In this paper the orthogonal impulse response functions (OIRF) are studied in the non-standard, though quite common, case where the covariance of the error vector is not constant in time. The usual approach for taking into account such behavior of the covariance consists in applying the standard tools to sub-periods of the whole sample. We underline that such a practice may lead to severe upward bias. We propose a new approach intended to give what we argue to be a more accurate resume of the time-varying OIRF. This consists in averaging the Cholesky decomposition of nonparametric covariance estimators. In addition an index is developed to evaluate the heteroscedasticity effect on the OIRF analysis. The asymptotic behavior of the different estimators considered in the paper is investigated. The theoretical results are illustrated by Monte Carlo experiments. The analysis of the orthogonal response functions of the U.S. inflation to an oil price shock, shows the relevance of the tools proposed herein for an appropriate analysis of economic variables.

stat.ME

Modified Cox regression with current status data

In survival analysis, the lifetime under study is not always observed. In certain applications, for some individuals, the value of the lifetime is only known to be smaller or larger than some random duration. This framework represent an extension of standard situations where the lifetime is only left or only right randomly censored. We consider the case where the independent observation units include also some covariates, and we propose two semiparametric regression models. The new models extend the standard Cox proportional hazard model to the situation of a more complex censoring mechanism. However, like in Cox's model, in both models the nonparametric baseline hazard function still could be expressed as an explicit functional of the distribution of the observations. This allows to define the estimator of the finite-dimensional parameters as the maximum of a likelihood-type criterion which is an explicit function of the data. Given an estimate of the finite-dimensional parameter, the estimation of the baseline cumulative hazard function is straightforward.

math.ST