SearcharxivSearch

arXiv subjects

John A. Rice

Publications and source records attributed to John A. Rice.

4 recordsLinked to original sources

A Conversation with Richard A. Olshen

Richard Olshen was born in Portland, Oregon, on May 17, 1942. Richard spent his early years in Chevy Chase, Maryland, but has lived most of his life in California. He received an A.B. in Statistics at the University of California, Berkeley, in 1963, and a Ph.D. in Statistics from Yale University in 1966, writing his dissertation under the direction of Jimmie Savage and Frank Anscombe. He served as Research Staff Statistician and Lecturer at Yale in 1966-1967. Richard accepted a faculty appointment at Stanford University in 1967, and has held tenured faculty positions at the University of Michigan (1972-1975), the University of California, San Diego (1975-1989), and Stanford University (since 1989). At Stanford, he is Professor of Health Research and Policy (Biostatistics), Chief of the Division of Biostatistics (since 1998) and Professor (by courtesy) of Electrical Engineering and of Statistics. At various times, he has had visiting faculty positions at Columbia, Harvard, MIT, Stanford and the Hebrew University. Richard's research interests are in statistics and mathematics and their applications to medicine and biology. Much of his work has concerned binary tree-structured algorithms for classification, regression, survival analysis and clustering. Those for classification and survival analysis have been used with success in computer-aided diagnosis and prognosis, especially in cardiology, oncology and toxicology. He coauthored the 1984 book Classification and Regression Trees (with Leo Brieman, Jerome Friedman and Charles Stone) which gives motivation, algorithms, various examples and mathematical theory for what have come to be known as CART algorithms. The approaches to tree-structured clustering have been applied to problems in digital radiography (with Stanford EE Professor Robert Gray) and to HIV genetics, the latter work including studies on single nucleotide polymorphisms, which has helped to shed light on the presence of hypertension in certain subpopulations of women.

stat.OT

Kernel Density Estimation with Berkson Error

Given a sample $\{X_i\}_{i=1}^n$ from $f_X$, we construct kernel density estimators for $f_Y$, the convolution of $f_X$ with a known error density $f_ε$. This problem is known as density estimation with Berkson error and has applications in epidemiology and astronomy. Little is understood about bandwidth selection for Berkson density estimation. We compare three approaches to selecting the bandwidth both asymptotically, using large sample approximations to the MISE, and at finite samples, using simulations. Our results highlight the relationship between the structure of the error $f_ε$ and the optimal bandwidth. In particular, the results demonstrate the importance of smoothing when the error term $f_ε$ is concentrated near 0. We propose a data--driven bandwidth estimator and test its performance on NO$_2$ exposure data.

stat.ME

Optimizing Automated Classification of Periodic Variable Stars in New Synoptic Surveys

Efficient and automated classification of periodic variable stars is becoming increasingly important as the scale of astronomical surveys grows. Several recent papers have used methods from machine learning and statistics to construct classifiers on databases of labeled, multi--epoch sources with the intention of using these classifiers to automatically infer the classes of unlabeled sources from new surveys. However, the same source observed with two different synoptic surveys will generally yield different derived metrics (features) from the light curve. Since such features are used in classifiers, this survey-dependent mismatch in feature space will typically lead to degraded classifier performance. In this paper we show how and why feature distributions change using OGLE and \textit{Hipparcos} light curves. To overcome survey systematics, we apply a method, \textit{noisification}, which attempts to empirically match distributions of features between the labeled sources used to construct the classifier and the unlabeled sources we wish to classify. Results from simulated and real--world light curves show that noisification can significantly improve classifier performance. In a three--class problem using light curves from \textit{Hipparcos} and OGLE, noisification reduces the classifier error rate from 27.0% to 7.0%. We recommend that noisification be used for upcoming surveys such as Gaia and LSST and describe some of the promises and challenges of applying noisification to these surveys.

astro-ph.IM

Statistical Methods for Detecting Stellar Occultations by Kuiper Belt Objects: the Taiwanese-American Occultation Survey

The Taiwanese-American Occultation Survey (TAOS) will detect objects in the Kuiper Belt, by measuring the rate of occultations of stars by these objects, using an array of three to four 50cm wide-field robotic telescopes. Thousands of stars will be monitored, resulting in hundreds of millions of photometric measurements per night. To optimize the success of TAOS, we have investigated various methods of gathering and processing the data and developed statistical methods for detecting occultations. In this paper we discuss these methods. The resulting estimated detection efficiencies will be used to guide the choice of various operational parameters determining the mode of actual observation when the telescopes come on line and begin routine observations. In particular we show how real-time detection algorithms may be constructed, taking advantage of having multiple telescopes. We also discuss a retrospective method for estimating the rate at which occultations occur.

astro-ph