SearcharxivSearch

arXiv subjects

Drew Lazar

Publications and source records attributed to Drew Lazar.

4 recordsLinked to original sources

Separating Spatial and Clinical Risk with Node-Splitting SVM Survival Trees

Recovering geographic variation in survival requires separating spatial risk from patients' clinical characteristics, a problem complicated by prognostic covariates that are themselves spatially structured. We develop a nonparametric two-stage method for this separation. A clinical survival tree fit to the covariates alone supplies leaf Nelson-Aalen cumulative hazard residuals, transferring the censored survival structure to a clinically adjusted scale without imposing a functional form on the clinical hazard, and a second tree fit to these residuals on the coordinates recovers the spatial structure. Both stages are kernel dipole-splitting survival trees, so the resulting spatial risk map is piecewise constant, with sharp, possibly curved boundaries. We establish when the residuals recover the spatial signal: under a multiplicative frailty and an exogeneity condition, they are free of the clinical covariates given location and stochastically ordered by the frailty. An expansion of the frailty Laplace exponent quantifies their approximate exponentiality and identifies the leading remainder. These results assume consistency of the first stage rather than a model class, so the construction extends to other survival estimators. On the LeukSurv leukemia data the method agrees with a Bayesian Gaussian random field frailty about where risk is elevated while resolving sharp adjacencies the smooth surface averages away, and an unadjusted spatial analysis misattributes clinical variation to location. Simulations with known zones, including a sweep through graded violations of exogeneity, locate the point at which the two contributions cease to be separately identifiable, with the smooth benchmark degrading in parallel as that point is approached.

stat.ME

Node Splitting SVMs for Survival Trees Based on an L2-Regularized Dipole Splitting Criteria

This paper proposes a novel, node-splitting support vector machine (SVM) for creating survival trees. This approach is capable of non-linearly partitioning survival data which includes continuous, right-censored outcomes. Our method improves on an existing non-parametric method, which uses at most oblique splits to induce survival regression trees. In the prior work, these oblique splits were created via a non-SVM approach, by minimizing a piece-wise linear objective, called a dipole splitting criterion, constructed from pairs of covariates and their associated survival information. We extend this method by enabling splits from a general class of non-linear surfaces. We achieve this by ridge regularizing the dipole-splitting criterion to enable application of kernel methods in a manner analogous to classical SVMs. The ridge regularization provides robustness and can be tuned. Using various kernels, we induce both linear and non-linear survival trees to compare their sizes and predictive powers on real and simulated data sets. We compare traditional univariate log-rank splits, oblique splits using the original dipole-splitting criterion and a variety of non-linear splits enabled by our method. In these tests, trees created by non-linear splits, using polynomial and Gaussian kernels show similar predictive power while often being of smaller sizes compared to trees created by univariate and oblique splits. This approach provides a novel and flexible array of survival trees that can be applied to diverse survival data sets.

stat.ME

Robust Optimization and Inference on Manifolds

We propose a robust and scalable procedure for general optimization and inference problems on manifolds leveraging the classical idea of `median-of-means' estimation. This is motivated by ubiquitous examples and applications in modern data science in which a statistical learning problem can be cast as an optimization problem over manifolds. Being able to incorporate the underlying geometry for inference while addressing the need for robustness and scalability presents great challenges. We address these challenges by first proving a key lemma that characterizes some crucial properties of geometric medians on manifolds. In turn, this allows us to prove robustness and tighter concentration of our proposed final estimator in a subsequent theorem. This estimator aggregates a collection of subset estimators by taking their geometric median over the manifold. We illustrate bounds on this estimator via calculations in explicit examples. The robustness and scalability of the procedure is illustrated in numerical examples on both simulated and real data sets.

stat.ME

Scale and curvature effects in principal geodesic analysis

There is growing interest in using the close connection between differential geometry and statistics to model smooth manifold-valued data. In particular, much work has been done recently to generalize principal component analysis (PCA), the method of dimension reduction in linear spaces, to Riemannian manifolds. One such generalization is known as principal geodesic analysis (PGA). This paper, in a novel fashion, obtains Taylor expansions in scaling parameters introduced in the domain of objective functions in PGA. It is shown this technique not only leads to better closed-form approximations of PGA but also reveals the effects that scale, curvature and the distribution of data have on solutions to PGA and on their differences to first-order tangent space approximations. This approach should be able to be applied not only to PGA but also to other generalizations of PCA and more generally to other intrinsic statistics on Riemannian manifolds.

stat.OT