Searcharxiv⌕ Search

arXiv subjects

Ryan Martin

Publications and source records attributed to Ryan Martin.

At least 91 records · Page 5Linked to original sources

Applications of an algorithm for solving Fredholm equations of the first kind

In this paper we use an iterative algorithm for solving Fredholm equations of the first kind. The basic algorithm is known and is based on an EM algorithm when involved functions are non-negative and integrable. With this algorithm we demonstrate two examples involving the estimation of a mixing density and a first passage time density function involving Brownian motion. We also develop the basic algorithm to include functions which are not necessarily non-negative and again present illustrations under this scenario. A self contained proof of convergence of all the algorithms employed is presented.

math.ST↗

A mathematical characterization of confidence as valid belief

Confidence is a fundamental concept in statistics, but there is a tendency to misinterpret it as probability. In this paper, I argue that an intuitively and mathematically more appropriate interpretation of confidence is through belief/plausibility functions, in particular, those that satisfy a certain validity property. Given their close connection with confidence, it is natural to ask how a valid belief/plausibility function can be constructed directly. The inferential model (IM) framework provides such a construction, and here I prove a complete-class theorem stating that, for every nominal confidence region, there exists a valid IM whose plausibility regions are contained by the given confidence region. This characterization has implications for statistics understanding and communication, and highlights the importance of belief functions and the IM framework.

math.ST↗

Gibbs posterior inference on the minimum clinically important difference

IIt is known that a statistically significant treatment may not be clinically significant. A quantity that can be used to assess clinical significance is called the minimum clinically important difference (MCID), and inference on the MCID is an important and challenging problem. Modeling for the purpose of inference on the MCID is non-trivial, and concerns about bias from a misspecified parametric model or inefficiency from a nonparametric model motivate an alternative approach to balance robustness and efficiency. In particular, a recently proposed representation of the MCID as the minimizer of a suitable risk function makes it possible to construct a Gibbs posterior distribution for the MCID without specifying a model. We establish the posterior convergence rate and show, numerically, that an appropriately scaled version of this Gibbs posterior yields interval estimates for the MCID which are both valid and efficient even for relatively small sample sizes.

stat.ME↗

On recursive Bayesian predictive distributions

A Bayesian framework is attractive in the context of prediction, but a fast recursive update of the predictive distribution has apparently been out of reach, in part because Monte Carlo methods are generally used to compute the predictive. This paper shows that online Bayesian prediction is possible by characterizing the Bayesian predictive update in terms of a bivariate copula, making it unnecessary to pass through the posterior to update the predictive. In standard models, the Bayesian predictive update corresponds to familiar choices of copula but, in nonparametric problems, the appropriate copula may not have a closed-form expression. In such cases, our new perspective suggests a fast recursive approximation to the predictive density, in the spirit of Newton's predictive recursion algorithm, but without requiring evaluation of normalizing constants. Consistency of the new algorithm is shown, and numerical examples demonstrate its quality performance in finite-samples compared to fully Bayesian and kernel methods.

stat.ME↗

Rethinking probabilistic prediction in the wake of the 2016 U.S. presidential election

To many statisticians and citizens, the outcome of the most recent U.S. presidential election represents a failure of data-driven methods on the grandest scale. This impression has led to much debate and discussion about how the election predictions went awry -- Were the polls inaccurate? Were the models wrong? Did we misinterpret the probabilities? -- and how they went right -- Perhaps the analyses were correct even though the predictions were wrong, that's just the nature of probabilistic forecasting. With this in mind, we analyze the election outcome with respect to a core set of effectiveness principles. Regardless of whether and how the election predictions were right or wrong, we argue that they were ineffective in conveying the extent to which the data was informative of the outcome and the level of uncertainty in making these assessments. Among other things, our analysis sheds light on the shortcomings of the classical interpretations of probability and its communication to consumers in the form of predictions. We present here an alternative approach, based on a notion of validity, which offers two immediate insights for predictive inference. First, the predictions are more conservative, arguably more realistic, and come with certain guarantees on the probability of an erroneous prediction. Second, our approach easily and naturally reflects the (possibly substantial) uncertainty about the model by outputting plausibilities instead of probabilities. Had these simple steps been taken by the popular prediction outlets, the election outcome may not have been so shocking.

stat.OT↗

Efficient posterior inference on the volatility of a jump diffusion process

Jump diffusion processes are widely used to model asset prices over time, mainly for their ability to capture complex discontinuous behavior, but inference on the model parameters remains a challenge. Here our goal is posterior inference on the volatility coefficient of the diffusion part of the process based on discrete samples. A Bayesian approach requires specification of a model for the jump part of the process, prior distributions for the corresponding parameters, and computation of the joint posterior. Since the volatility coefficient is our only interest, it would be desirable to avoid the modeling and computational costs associated with the jump part of the process. Towards this, we consider a {\em purposely misspecified model} that ignores the jump part entirely. We work out precisely the asymptotic behavior of the Bayesian posterior under the misspecified model, propose some simple modifications to correct for the effects of misspecification, and demonstrate that our modified posterior inference on the volatility is efficient in the sense that its asymptotic variance equals the no-jumps model Cramér--Rao bound.

stat.ME↗

Exact prior-free probabilistic inference in a class of non-regular models

The use of standard statistical methods, such as maximum likelihood, is often justified based on their asymptotic properties. For suitably regular models, this theory is standard but, when the model is non-regular, e.g., the support depends on the parameter, these asymptotic properties may be difficult to assess. Recently, an inferential model (IM) framework has been developed that provides valid prior-free probabilistic inference without the need for asymptotic justification. In this paper, we construct an IM for a class of highly non-regular models with parameter-dependent support. This construction requires conditioning, which is facilitated through the solution of a particular differential equation. We prove that the plausibility intervals derived from this IM are exact confidence intervals, and we demonstrate their efficiency in a simulation study.

stat.ME↗

On an inferential model construction using generalized associations

The inferential model (IM) approach, like fiducial and its generalizations, depends on a representation of the data-generating process. Here, a particular variation on the IM construction is considered, one based on generalized associations. The resulting generalized IM is more flexible than the basic IM in that it does not require a complete specification of the data-generating process and is provably valid under mild conditions. Computation and marginalization strategies are discussed, and two applications of this generalized IM approach are presented.

stat.ME↗

A statistical inference course based on p-values

Introductory statistical inference texts and courses treat the point estimation, hypothesis testing, and interval estimation problems separately, with primary emphasis on large-sample approximations. Here I present an alternative approach to teaching this course, built around p-values, emphasizing provably valid inference for all sample sizes. Details about computation and marginalization are also provided, with several illustrative examples, along with a course outline.

stat.OT↗

Valid uncertainty quantification about the model in a linear regression setting

In scientific applications, there often are several competing models that could be fit to the observed data, so quantification of the model uncertainty is of fundamental importance. In this paper, we develop an inferential model (IM) approach for simultaneously valid probabilistic inference over a collection of assertions of interest without requiring any prior input. Our construction guarantees that the approach is optimal in the sense that it is the most efficient among those which are valid. Connections between the IM's simultaneous validity and post-selection inference are also made. We apply the general results to obtain valid uncertainty quantification about the set of predictor variables to be included in a linear regression model.

math.ST↗

On weighted Ramsey numbers

The weighted Ramsey number, ${\rm wR}(n,k)$, is the minimum $q$ such that there is an assignment of nonnegative real numbers (weights) to the edges of $K_n$ with the total sum of the weights equal to ${n\choose 2}$ and there is a Red/Blue coloring of edges of the same $K_n$, such that in any complete $k$-vertex subgraph $H$, of $K_n$, the sum of the weights on Red edges in $H$ is at most $q$ and the sum of the weights on Blue edges in $H$ is at most $q$. This concept was introduced recently by Fujisawa and Ota. We provide new bounds on ${\rm wR}(n,k)$, for $k\geq 4$ and $n$ large enough and show that determining ${\rm wR}(n,3)$ is asymptotically equivalent to the problem of finding the fractional packing number of monochromatic triangles in colorings of edges of complete graphs with two colors.

math.CO↗

Avoiding rainbow induced subgraphs in vertex-colorings

For a fixed graph $H$ on $k$ vertices, and a graph $G$ on at least $k$ vertices, we write $G\rightarrow H$ if in any vertex-coloring of $G$ with $k$ colors, there is an induced subgraph isomorphic to $H$ whose vertices have distinct colors. In other words, if $G\rightarrow H$ then a totally multicolored induced copy of $H$ is unavoidable in any vertex-coloring of $G$ with $k$ colors. In this paper, we show that, with a few notable exceptions, for any graph $H$ on $k$ vertices and for any graph $G$ which is not isomorphic to $H$, $G\not\!\rightarrow H$. We explicitly describe all exceptional cases. This determines the induced vertex-anti-Ramsey number for all graphs and shows that totally multicolored induced subgraphs are, in most cases, easily avoidable.

math.CO↗

On Posterior Concentration in Misspecified Models

We investigate the asymptotic behavior of Bayesian posterior distributions under independent and identically distributed ($i.i.d.$) misspecified models. More specifically, we study the concentration of the posterior distribution on neighborhoods of $f^{\star}$, the density that is closest in the Kullback--Leibler sense to the true model $f_0$. We note, through examples, the need for assumptions beyond the usual Kullback--Leibler support assumption. We then investigate consistency with respect to a general metric under three assumptions, each based on a notion of divergence measure, and then apply these to a weighted $L_1$-metric in convex models and non-convex models. Although a few results on this topic are available, we believe that these are somewhat inaccessible due, in part, to the technicalities and the subtle differences compared to the more familiar well-specified model case. One of our goals is to make some of the available results, especially that of , more accessible. Unlike their paper, our approach does not require construction of test sequences. We also discuss a preliminary extension of the $i.i.d.$ results to the independent but not identically distributed ($i.n.i.d.$) case.

math.ST↗

Ultra-Low Noise Mechanically Cooled Germanium Detector

Low capacitance, large volume, high purity germanium (HPGe) radiation detectors have been successfully employed in low-background physics experiments. However, some physical processes may not be detectable with existing detectors whose energy thresholds are limited by electronic noise. In this paper, methods are presented which can lower the electronic noise of these detectors. Through ultra-low vibration mechanical cooling and wire bonding of a CMOS charge sensitive preamplifier to a sub-pF p-type point contact HPGe detector, we demonstrate electronic noise levels below 40 eV-FWHM.

physics.ins-det↗

Status Update of the MAJORANA DEMONSTRATOR Neutrinoless Double Beta Decay Experiment

Neutrinoless double beta decay searches play a major role in determining neutrino properties, in particular the Majorana or Dirac nature of the neutrino and the absolute scale of the neutrino mass. The consequences of these searches go beyond neutrino physics, with implications for Grand Unification and leptogenesis. The \textsc{Majorana} Collaboration is assembling a low-background array of high purity Germanium (HPGe) detectors to search for neutrinoless double-beta decay in $^{76}$Ge. The \textsc{Majorana Demonstrator}, which is currently being constructed and commissioned at the Sanford Underground Research Facility in Lead, South Dakota, will contain 44 kg (30 kg enriched in $^{76}$Ge) of HPGe detectors. Its primary goal is to demonstrate the scalability and background required for a tonne-scale Ge experiment. This is accomplished via a modular design and projected background of less than 3 cnts/tonne-yr in the region of interest. The experiment is currently taking data with the first of its enriched detectors.

physics.ins-det↗

A semiparametric scale-mixture regression model and predictive recursion maximum likelihood

To avoid specification of the error distribution in a regression model, we propose a general nonparametric scale mixture model for the error distribution. For fitting such mixtures, the predictive recursion method is a simple and computationally efficient alternative to existing methods. We define a predictive recursion-based marginal likelihood function, and estimation of the regression parameters proceeds by maximizing this function. A hybrid predictive recursion--EM algorithm is proposed for this purpose. The method's performance is compared with that of existing methods in simulations and real data analyses.

stat.ME↗

Simulating from a gamma distribution with small shape parameter

Simulating from a gamma distribution with small shape parameter is a challenging problem. Towards an efficient method, we obtain a limiting distribution for a suitably normalized gamma distribution when the shape parameter tends to zero. Then this limiting distribution provides insight to the construction of a new, simple, and highly efficient acceptance--rejection algorithm. Comparisons based on acceptance rates show that the proposed procedure is more efficient than existing acceptance--rejection methods.

stat.CO↗

Prior-free probabilistic prediction of future observations

Prediction of future observations is a fundamental problem in statistics. Here we present a general approach based on the recently developed inferential model (IM) framework. We employ an IM-based technique to marginalize out the unknown parameters, yielding prior-free probabilistic prediction of future observables. Verifiable sufficient conditions are given for validity of our IM for prediction, and a variety of examples demonstrate the proposed method's performance. Thanks to its generality and ease of implementation, we expect that our IM-based method for prediction will be a useful tool for practitioners.

stat.ME↗