Searcharxiv⌕ Search

arXiv subjects

Mary Frances Dorn

Publications and source records attributed to Mary Frances Dorn.

3 recordsLinked to original sources

Robust Wrapped Gaussian Process Inference for Noisy Angular Data

Angular data are commonly encountered in settings with a directional or orientational component. Regressing an angular response on real-valued features requires intrinsically capturing the circular or spherical manifold the data lie on, or using an appropriate extrinsic transformation. A popular example of the latter is the technique of distributional wrapping, in which functions are "wrapped" around the unit circle via a modulo-$2π$ transformation. This approach enables flexible, non-linear models like Gaussian processes (GPs) to properly account for circular structure. While straightforward in concept, the need to infer the latent unwrapped distribution along with its wrapping behavior makes inference difficult in noisy response settings, as misspecification of one can severely hinder estimation of the other. However, applications such as radiowave analysis (Shangguan et al., 2015) and biomedical engineering (Kurz and Hanebeck, 2015) encounter radial data where wrapping occurs in only one direction. We therefore propose a novel wrapped GP (WGP) model formulation that recognizes monotonic wrapping behavior for more accurate inference in these situations. This is achieved by estimating the locations where wrapping occurs and partitioning the input space accordingly. We also specify a more robust Student's t response likelihood, and take advantage of an elliptical slice sampling (ESS) algorithm for rejection-free sampling from the latent GP space. We showcase our model's preferable performance on simulated examples compared to existing WGP methodologies. We then apply our method to the problem of localizing radiofrequency identification (RFID) tags, in which we model the relationship between frequency and phase angle to infer how far away an RFID tag is from an antenna.

stat.AP↗

Beyond Trees: Classification with Sparse Pairwise Dependencies

Several classification methods assume that the underlying distributions follow tree-structured graphical models. Indeed, trees capture statistical dependencies between pairs of variables, which may be crucial to attain low classification errors. The resulting classifier is linear in the log-transformed univariate and bivariate densities that correspond to the tree edges. In practice, however, observed data may not be well approximated by trees. Yet, motivated by the importance of pairwise dependencies for accurate classification, here we propose to approximate the optimal decision boundary by a sparse linear combination of the univariate and bivariate log-transformed densities. Our proposed approach is semi-parametric in nature: we non-parametrically estimate the univariate and bivariate densities, remove pairs of variables that are nearly independent using the Hilbert-Schmidt independence criteria, and finally construct a linear SVM on the retained log-transformed densities. We demonstrate using both synthetic and real data that our resulting classifier, denoted SLB (Sparse Log-Bivariate density), is competitive with popular classification methods.

stat.ML↗

Locally Optimized Random Forests

Standard supervised learning procedures are validated against a test set that is assumed to have come from the same distribution as the training data. However, in many problems, the test data may have come from a different distribution. We consider the case of having many labeled observations from one distribution, $P_1$, and making predictions at unlabeled points that come from $P_2$. We combine the high predictive accuracy of random forests (Breiman, 2001) with an importance sampling scheme, where the splits and predictions of the base-trees are done in a weighted manner, which we call Locally Optimized Random Forests. These weights correspond to a non-parametric estimate of the likelihood ratio between the training and test distributions. To estimate these ratios with an unlabeled test set, we make the covariate shift assumption, where the differences in distribution are only a function of the training distributions (Shimodaira, 2000.) This methodology is motivated by the problem of forecasting power outages during hurricanes. The extreme nature of the most devastating hurricanes means that typical validation set ups will overly favor less extreme storms. Our method provides a data-driven means of adapting a machine learning method to deal with extreme events.

stat.ML↗