Searcharxiv⌕ Search

arXiv · 2609.40242

Riemannian Regression

Abstract

Classical linear regression assumes that the relevant geometry of the predictor space is Euclidean and that all centered observations contribute to the least-squares fit in the same geometric scale. This paper proposes \emph{Riemannian Regression}, a regression framework in which the usual vector differences are replaced by locally weighted differences induced by a data-dependent similarity structure. We introduce a generalized framework, termed {\em Riemannian Regression}, extending classic regression to any data endowed with a local distance structure. By equipping data tables with local metrics, we adapt regression model to incorporate manifold geometry. Given a similarity matrix $S=(S_{ij})$, obtained from UMAP, ISOMAP, or DBSCAN \cite{mcinnes,isomap,dbscan}, we define the dissimilarity coefficient $ρ_{ij}=1-S_{ij}$ and the induced subtraction $ x_i\ominus x_j=ρ_{ij}(x_i-x_j). $ A Riemannian center $g=x_λ$ is selected as a discrete Fréchet mean, and regression is performed on the Riemannian-centered variables $X_R=W X_{c,λ}$ and $y_R=W y_{c,λ}$, where $W=\operatorname{diag}(ρ_{1λ},\ldots,ρ_{nλ})$. The resulting estimator has the weighted least-squares form $ \widehatβ_R=(X_{c,λ}^{t}W^2X_{c,λ})^{-1}X_{c,λ}^{t}W^2y_{c,λ}. $ The proposed approach preserves the linear form of the regression model while changing the geometry of the fit. The paper develops three ways to construct the local metric: UMAP-based fuzzy similarities, ISOMAP-based normalized geodesic distances, and DBSCAN-based density similarities. Simulated examples and the Abalone data set illustrate how Riemannian Regression can reduce the influence of locally anomalous observations and adapt to regions with different local densities.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Oldemar Rodríguez. 2026-09-30. Riemannian Regression. https://arxiv.org/abs/2609.40242

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Optimal Estimation of Large-Dimensional Nonlinear Factor Models

This paper studies optimal estimation of large-dimensional nonlinear factor models. The key challenge is that the observed variables are possibly nonlinear functions of some latent variables, with the functional forms left unspecified. A local principal component analysis method combining K-nearest neighbors matching and principal component analysis is proposed to estimate the factor structure and recover information on latent variables and latent functions. Large-sample properties are established, including a sharp bound on the matching discrepancy of nearest neighbors, sup-norm error bounds for estimated local factors and factor loadings, and the uniform convergence rate of the factor structure estimator. Under mild conditions our estimator of the latent factor structure can achieve the optimal rate of uniform convergence for nonparametric regression. The method is illustrated with a Monte Carlo experiment and an empirical application studying the effect of tax cuts on economic growth.

math.ST↗

Adaptive Density Estimation Using Projection Kernels and Penalized Comparison to Overfitting

In this work, we study wavelet projection estimators for density estimation, based on compactly supported $\mathcal S$-regular scaling functions. The main issue is the choice of the resolution level, which determines the bias--variance trade-off. We select this level by a Penalized Comparison to Overfitting (PCO) criterion: each candidate estimator is compared in $Ł^2(\R)$ with a single overfitting reference, and the additional stochastic fluctuation is corrected by an explicit penalty. For the selected estimator, we prove a high-probability oracle inequality and a risk oracle inequality in expectation. Over Besov balls of densities satisfying a common $Ł^\infty$ bound, the risk inequality yields the rate $n^{-2r/(2r+1)}$ uniformly over the class, with a sufficient penalty threshold that can be chosen uniformly over the class. This rate has the classical minimax order, and the procedure adapts to the unknown Besov regularity. Numerical experiments on several density shapes confirm that the data-driven level is close to the oracle one and adapts to the structure of the target.

math.ST↗

Minimax Optimal Estimation of Mean and Covariance Functions with Spectral Regularization

Estimation of the mean and covariance functions is a fundamental problem in functional data analysis, particularly for discretely observed functional data. In this work, we study a regularization-based framework for estimating the mean and the covariance functions within a reproducing kernel Hilbert space (RKHS) setting. Our approach utilizes a spectral regularization technique under Hölder-type source conditions, allowing for a broad class of regularization schemes and accommodating a wide range of smoothness assumptions on the target functions. In contrast to RKHS formulations that assume the target belongs to the underlying RKHS, our source-condition framework also accommodates misspecified targets. Convergence rates for the proposed estimators are derived, and we derive corresponding minimax lower bounds and identify regimes in which the upper bounds are optimal, or optimal up to logarithmic factors.

math.ST↗