Searcharxiv⌕ Search

arXiv subjects

Kevin He

Publications and source records attributed to Kevin He.

58 records · Page 4Linked to original sources

The Outer Halo of the Milky Way as Probed by RR Lyr Variables from the Palomar Transient Facility

RR Lyr stars are ideal massless tracers that can be used to study the total mass and dark matter content of the outer halo of the Milky Way. This is because they are easy to find in the light curve databases of large stellar surveys and their distances can be determined with only knowledge of the light curve. We present here a sample of 112 RR Lyr beyond 50 kpc in the outer halo of the Milky Way, excluding the Sgr streams, for which we have obtained moderate resolution spectra with Deimos on the Keck 2 Telescope. Four of these have distances exceeding 100 kpc. These were selected from a much larger set of 447 candidate RR Lyr which were datamined using machine learning techniques applied to the light curves of variable stars in the Palomar Transient Facility database. The observed radial velocities taken at the phase of the variable corresponding to the time of observation were converted to systemic radial velocities in the Galactic standard of rest. From our sample of 112 RR Lyr we determine the radial velocity dispersion in the outer halo of the Milky Way to be ~90 km/s at 50 kpc falling to about 65 km/s near 100 kpc once a small number of major outliers are removed. With reasonable estimates of the completeness of our sample of 447 candidates and assuming a spherical halo, we find that the stellar density in the outer halo declines as the -4 power of r.

astro-ph.GA↗

Bayesian Posteriors For Arbitrarily Rare Events

We study how much data a Bayesian observer needs to correctly infer the relative likelihoods of two events when both events are arbitrarily rare. Each period, either a blue die or a red die is tossed. The two dice land on side $1$ with unknown probabilities $p_1$ and $q_1$, which can be arbitrarily low. Given a data-generating process where $p_1\ge c q_1$, we are interested in how much data is required to guarantee that with high probability the observer's Bayesian posterior mean for $p_1$ exceeds $(1-δ)c$ times that for $q_1$. If the prior densities for the two dice are positive on the interior of the parameter space and behave like power functions at the boundary, then for every $ε>0,$ there exists a finite $N$ so that the observer obtains such an inference after $n$ periods with probability at least $1-ε$ whenever $np_1\ge N$. The condition on $n$ and $p_1$ is the best possible. The result can fail if one of the prior densities converges to zero exponentially fast at the boundary.

math.ST↗

Latent Gaussian Mixture Models for Nationwide Kidney Transplant Center Evaluation

Five year post-transplant survival rate is an important indicator on quality of care delivered by kidney transplant centers in the United States. To provide a fair assessment of each transplant center, an effect that represents the center-specific care quality, along with patient level risk factors, is often included in the risk adjustment model. In the past, the center effects have been modeled as either fixed effects or Gaussian random effects, with various pros and cons. Our numerical analyses reveal that the distributional assumptions do impact the prediction of center effects especially when the effect is extreme. To bridge the gap between these two approaches, we propose to model the transplant center effect as a latent random variable with a finite Gaussian mixture distribution. Such latent Gaussian mixture models provide a convenient framework to study the heterogeneity among the transplant centers. To overcome the weak identifiability issues, we propose to estimate the latent Gaussian mixture model using a penalized likelihood approach, and develop sequential locally restricted likelihood ratio tests to determine the number of components in the Gaussian mixture distribution. The fitted mixture model provides a convenient means of controlling the false discovery rate when screening for underperforming or outperforming transplant centers. The performance of the methods is verified by simulations and by the analysis of the motivating data example.

stat.AP↗

Classification with Ultrahigh-Dimensional Features

Although much progress has been made in classification with high-dimensional features \citep{Fan_Fan:2008, JGuo:2010, CaiSun:2014, PRXu:2014}, classification with ultrahigh-dimensional features, wherein the features much outnumber the sample size, defies most existing work. This paper introduces a novel and computationally feasible multivariate screening and classification method for ultrahigh-dimensional data. Leveraging inter-feature correlations, the proposed method enables detection of marginally weak and sparse signals and recovery of the true informative feature set, and achieves asymptotic optimal misclassification rates. We also show that the proposed procedure provides more powerful discovery boundaries compared to those in \citet{CaiSun:2014} and \citet{JJin:2009}. The performance of the proposed procedure is evaluated using simulation studies and demonstrated via classification of patients with different post-transplantation renal functional types.

stat.ML↗