Searcharxiv⌕ Search

arXiv · 2610.04677

Estimation and inference in penalized function-on-function regression with dependent residual curves

Abstract

Penalized function-on-function regression, as implemented in the R package refund, treats all grid values of all curves as independent. Residual curves are smooth, so values within a curve are dependent. We study what this does to point estimates and pointwise confidence intervals, in a simulation with Gaussian, count and binary responses and in a plasmode study built on nine real datasets. Under within-curve dependence, model-based 95% intervals cover 0.63--0.82. A curve-clustered sandwich covariance with a leverage adjustment for the penalized hat matrix (CL2), computed after the usual REML fit, brings coverage close to nominal in the simulation; in the plasmode study it is a few points short on eight datasets and 0.84--0.91 for the coefficient surface on the ninth. Satterthwaite degrees of freedom recover most of the remaining shortfall except on that dataset, also with only 20 curves, at the cost of wider intervals. REML, however, undersmooths the coefficient surface: its error is two to four times that of a fit whose smoothing parameters are chosen by curve-blocked neighbourhood cross-validation (NCV), and its CL2 intervals for the surface are wide. Intervals centred at the smoother NCV estimate are narrow but undercover, the familiar tension between estimation and pointwise inference in nonparametric regression. We recommend the NCV estimate to describe the shape of the surface and REML + CL2 with Satterthwaite degrees of freedom for inference. For a surface that is zero everywhere, under dependent errors, these intervals exclude zero at about 5% of the grid points, and the NCV estimate is five to seven times closer to zero than REML's. For fitted means, intercepts and scalar effects, REML + CL2 intervals are close to nominal, except for binary intercepts and scalar effects and for misregistered counts. We state which defaults of refund the evidence supports.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fabian Scheipl. 2026-10-03. Estimation and inference in penalized function-on-function regression with dependent residual curves. https://arxiv.org/abs/2610.04677

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

GARCH copulas, v-transforms and D-vines for stochastic volatility

The bivariate copulas that describe the dependencies and partial dependencies of lagged variables in strictly stationary, first-order GARCH-type processes are investigated. It is shown that the copulas of symmetric GARCH processes are jointly symmetric but non-exchangeable, while the copulas of processes with symmetric innovation distributions and asymmetric leverage effects have weaker h-symmetry; copulas with asymmetric innovation distributions have neither form of symmetry. Since the true bivariate copulas are typically inaccessible, due to the unknown functional forms of the marginal distributions of GARCH processes, a new class of approximating copulas is proposed. These rely on copula density constructions that combine standard bivariate copula densities for positive dependence with two uniformity-preserving transformations known as v-transforms. The construction is shown to be particularly effective when applied to the density of the copula of the absolute values of a spherical t distribution. Tractable simplified D-vines incorporating the new pair copulas are developed for applications to time series showing stochastic volatility. The resulting models are shown to provide better fits to simulated data from GARCH processes, and to a dataset of financial exchange-rate returns, than have previously been obtained using vine copulas.

stat.ME↗

Selecting Informative Conformal Prediction Sets with an Optimized FCR-Controlled Approach

Conformal methods provide prediction sets for outcomes with confidence guarantees. We study their use in a selective inference setting, where inference is performed only when the prediction set is informative. The analyst may consider as informative, for example, cases with prediction sets that are sufficiently small, exclude null values, or satisfy other appropriate monotone constraints. Because inference is typically restricted to informative cases in practical applications, accounting for the resulting selection bias is crucial to maintaining false coverage rate (FCR) control. A general framework for constructing such informative conformal prediction sets while controlling the FCR on the selected sample was suggested in Gazin et al. (2025). In this work we focus on oracle-guided procedures. We derive the optimal decision policy under a suitable power objective in the oracle setting where the probability of belonging to each prediction set can be computed. In practice, of course, only estimated probabilities are available. We therefore introduce calibration procedures that adjust the oracle policy to maintain finite-sample FCR control. We show that this approach can achieve substantially higher power than available alternatives. We demonstrate the effectiveness of our new methods for classification outcomes on both real and simulated data.

stat.ME↗

Model--based clustering for spherical and hyper--spherical data using elliptically symmetric distributions

Model--based clustering for directional data data has attracted a lot of interest, but most methods utilize rotationally symmetric distributions. This paper suggests the use of elliptically symmetric distributions, namely the elliptically symmetric angular Gaussian and the spherical elliptically symmetric projected Cauchy distributions that were recently proposed in the literature for modelling spherical data. The expectation--maximization algorithm is employed and the inclusion of covariates is also examined. Simulation studies compare the two distributions in terms of choosing the optimal number of clusters and computational cost. We use the mixtures of these two distributions to cluster two datasets on the sphere (earthquake locations) and two hyper--spherical datasets.

stat.ME↗