Searcharxiv⌕ Search

arXiv subjects

Yaakov Malinovsky

Publications and source records attributed to Yaakov Malinovsky.

At least 37 records · Page 2Linked to original sources

A note on the closed-form solution for the longest head run problem of Abraham de Moivre

The problem of the longest head run was introduced and solved by Abraham de Moivre in the second edition of his book Doctrine of Chances (de Moivre, 1738). The closed-form solution as a finite sum involving binomial coefficients was provided in Uspensky (1937). Since then, the problem and its variations and extensions have found broad interest and diverse applications. Surprisingly, a very simple closed form can be obtained, which we present in this note.

math.HO↗

Nested Group Testing Procedures for Screening

This article reviews a class of adaptive group testing procedures that operate under a probabilistic model assumption as follows. Consider a set of $N$ items, where item $i$ has the probability $p$ ($p_i$ in the generalized group testing) to be defective, and the probability $1-p$ to be non-defective independent from the other items. A group test applied to any subset of size $n$ is a binary test with two possible outcomes, positive or negative. The outcome is negative if all $n$ items are non-defective, whereas the outcome is positive if at least one item among the $n$ items is defective. The goal is complete identification of all $N$ items with the minimum expected number of tests.

stat.ME↗

Conjectures on Optimal Nested Generalized Group Testing Algorithm

Consider a finite population of $N$ items, where item $i$ has a probability $p_i$ to be defective. The goal is to identify all items by means of group testing. This is the generalized group testing problem (hereafter GGTP). In the case of $\displaystyle p_1=\cdots=p_{N}=p$ \cite{YH1990} proved that the pairwise testing algorithm is the optimal nested algorithm, with respect to the expected number of tests, for all $N$ if and only if $\displaystyle p \in [1-1/\sqrt{2},\,(3-\sqrt{5})/2]$ (R-range hereafter) (an optimal at the boundary values). In this note, we present a result that helps to define the generalized pairwise testing algorithm (hereafter GPTA) for the GGTP. We present two conjectures: (1) when all $p_i, i=1,\ldots,N$ belong to the R-range, GPTA is the optimal procedure among nested procedures applied to $p_i$ of nondecreasing order; (2) if all $p_i, i=1,\ldots,N$ belong to the R-range, GPTA the optimal nested procedure, i.e., minimises the expected total number of tests with respect to all possible testing orders in the class of nested procedures. Although these conjectures are logically reasonable, we were only able to empirically verify the first one up to a particular level of $N$. We also provide a short survey of GGTP.

stat.OT↗

An optimal design for hierarchical generalized group testing

Choosing an optimal strategy for hierarchical group testing is an important problem for practitioners who are interested in disease screening with limited resources. For example, when screening for infectious diseases in large populations, it is important to use algorithms that minimize the cost of potentially expensive assays. Black et al. (2015) described this as an intractable problem unless the number of individuals to screen is small. They proposed an approximation to an optimal strategy that is difficult to implement for large population sizes. In this article, we develop an optimal design with respect to the expected total number of tests that can be obtained using a novel dynamic programming algorithm. We show that this algorithm is substantially more efficient than the approach proposed by Black et al. (2015). In addition, we compare the two designs for imperfect tests. R code is provided for the practitioner.

stat.ME↗

A Unified Approach for Solving Sequential Selection Problems

In this paper we develop a unified approach for solving a wide class of sequential selection problems. This class includes, but is not limited to, selection problems with no-information, rank-dependent rewards, and considers both fixed as well as random problem horizons. The proposed framework is based on a reduction of the original selection problem to one of optimal stopping for a sequence of judiciously constructed independent random variables. We demonstrate that our approach allows exact and efficient computation of optimal policies and various performance metrics thereof for a variety of sequential selection problems, several of which have not been solved to date.

math.PR↗

Stochastic precedence and minima among dependent variables

The notion of stochastic precedence between two random variables emerges as a relevant concept in several fields of applied probability. When one consider a vector of random variables $X_1,...,X_n$, this notion has a preeminent role in the analysis of minima of the type $\min_{j \in A} X_j$ for $A \subset \{1, \ldots n\}$. In such an analysis, however, several apparently controversial aspects can arise (among which phenomena of "non-transitivity"). Here we concentrate attention on vectors of non-negative random variables with absolutely continuous joint distributions, in which a case the set of the multivariate conditional hazard rate (m.c.h.r.) functions can be employed as a convenient method to describe different aspects of stochastic dependence. In terms of the m.c.h.r. functions, we first obtain convenient formulas for the probability distributions of the variables $\min_{j \in A} X_j$ and for the probability of events $\{X_i=\min_{j \in A} X_j\}$. Then we detail several aspects of the notion of stochastic precedence. On these bases, we explain some controversial behavior of such variables and give sufficient conditions under which paradoxical aspects can be excluded. On the purpose of stimulating active interest of readers, we present several comments and pertinent examples.

math.PR↗

Efficient methods for the estimation of the multinomial parameter for the two-trait group testing model

Estimation of a single Bernoulli parameter using pooled sampling is among the oldest problems in the group testing literature. To carry out such estimation, an array of efficient estimators have been introduced covering a wide range of situations routinely encountered in applications. More recently, there has been growing interest in using group testing to simultaneously estimate the joint probabilities of two correlated traits using a multinomial model. Unfortunately, basic estimation results, such as the maximum likelihood estimator (MLE), have not been adequately addressed in the literature for such cases. In this paper, we show that finding the MLE for this problem is equivalent to maximizing a multinomial likelihood with a restricted parameter space. A solution using the EM algorithm is presented which is guaranteed to converge to the global maximizer, even on the boundary of the parameter space. Two additional closed form estimators are presented with the goal of minimizing the bias and/or mean square error. The methods are illustrated by considering an application to the joint estimation of transmission prevalence for two strains of the Potato virus Y by the aphid myzus persicae.

stat.ME↗

On the construction of unbiased estimators for the group testing problem

Debiased estimation has long been an area of research in the group testing literature. This has led to the development of several estimators with the goal of bias minimization and, recently, an unbiased estimator based on sequential binomial sampling. Previous research, however, has focused heavily on the simple case where no misclassification is assumed and only one trait is to be tested. In this paper, we consider the problem of unbiased estimation in these broader areas, giving constructions of such estimators for several cases. We show that, outside of the standard case addressed previously in the literature, it is impossible to find any proper unbiased estimator, that is, an estimator giving only values in the parameter space. This is shown to hold generally under any binomial or multinomial sampling plans

stat.ME↗

On optimal policy in the group testing with incomplete identification

Consider a very large (infinite) population of items, where each item independent from the others is defective with probability p, or good with probability q=1-p. The goal is to identify N good items as quickly as possible. The following group testing policy (policy A) is considered: test items together in the groups, if the test outcome of group i of size n_i is negative, then accept all items in this group as good, otherwise discard the group. Then, move to the next group and continue until exact N good items are found. The goal is to find an optimal testing configuration, i.e., group sizes, under policy A, such that the expected waiting time to obtain N good items is minimal. Recently, Gusev (2012) found an optimal group testing configuration under the assumptions of constant group size and N=\infty. In this note, an optimal solution under policy A for finite N is provided. Keywords: Dynamic programming; Optimal design; Partition problem; Shur-convexity

stat.OT↗

Follow Up on Detecting Deficiencies: An Optimal Group Testing Algorithm

In a recent volume of Mathematics Magazine (Vol. 90, No. 3, June 2017) there is an interesting article by Seth Zimmerman, titled Detecting Deficiencies: An Optimal Group Testing Algorithm. The claim in the summary is contradictory to well-known facts reported in the group- testing literature, which is easily verified, beginning with the work by Sobel and Groll (1959), which was cited by S. Zimmerman himself. Therefore, I feel compelled to offer a number of comments and clarifications. In addition, I have made some correction of mistaken claim made by Zimmerman (2017).

stat.AP↗

Proportional Closeness Estimation of Probability of Contamination Under Group Testing

The paper is focused on the problem of estimating the probability $p$ of individual contaminated sample, under group testing. The precision of the estimator is given by the probability of proportional closeness, a concept defined in the Introduction. Two-stage and sequential sampling procedures are characterized. An adaptive procedure is examined.

math.ST↗

Revisiting nested group testing procedures: new results, comparisons, and robustness

Group testing has its origin in the identification of syphilis in the US army during World War II. Much of the theoretical framework of group testing was developed starting in the late 1950s, with continued work into the 1990s. Recently, with the advent of new laboratory and genetic technologies, there has been an increasing interest in group testing designs for cost saving purposes. In this paper, we compare different nested designs, including Dorfman, Sterrett and an optimal nested procedure obtained through dynamic programming. To elucidate these comparisons, we develop closed-form expressions for the optimal Sterrett procedure and provide a concise review of the prior literature for other commonly used procedures. We consider designs where the prevalence of disease is known as well as investigate the robustness of these procedures when it is incorrectly assumed. This article provides a technical presentation that will be of interest to researchers as well as from a pedagogical perspective. Supplementary material for this article is available online.

stat.OT↗

Sterrett Procedure for the Generalized Group Testing Problem

Group testing is a useful method that has broad applications in medicine, engineering, and even in airport security control. Consider a finite population of $N$ items, where item $i$ has a probability $p_i$ to be defective. The goal is to identify all items by means of group testing. This is the generalized group testing problem. The optimum procedure, with respect to the expected total number of tests, is unknown even in case when all $p_i$ are equal. \cite{H1975} proved that an ordered partition (with respect to $p_i$) is the optimal for the Dorfman procedure (procedure $D$), and obtained an optimum solution (i.e., found an optimal partition) by dynamic programming. In this paper, we investigate the Sterrett procedure (procedure $S$). We provide close form expression for the expected total number of tests, which allows us to find the optimum arrangement of the items in the particular group. We also show that an ordered partition is not optimal for the procedure $S$ or even for a slightly modified Dorfman procedure (procedure $D^{\prime}$). This discovery implies that finding an optimal procedure $S$ appears to be a hard computational problem. However, by using an optimal ordered partition for all procedures, we show that procedure $D^{\prime}$ is uniformly better than procedure $D$, and based on numerical comparisons, procedure $S$ is uniformly and significantly better than procedures $D$ and $D^{\prime}$.

stat.OT↗

Sequential estimation in the group testing problem

Estimation using pooled sampling has long been an area of interest in the group testing literature. Such research has focused primarily on the assumed use of fixed sampling plans (i), although some recent papers have suggested alternative sequential designs that sample until a predetermined number of positive tests (ii). One major consideration, including in the new work on sequential plans, is the construction of debiased estimators which either reduce or keep the mean square error from inflating. Whether, however, under the above or other sampling designs unbiased estimation is in fact possible has yet to be established in the literature. In this paper, we introduce a design which samples until a fixed number of negatives (iii), and show that an unbiased estimator exists under this model, while unbiased estimation is not possible for either of the preceding designs (i) and (ii). We present new estimators under the different sampling plans that are either unbiased or that have reduced bias relative to those already in use as well as generally improve on the mean square error. Numerical studies are done in order to compare designs in terms of bias and mean square error under practical situations with small and medium sample sizes.

math.ST↗

On the structure of UMVUEs

In all setups when the structure of UMVUEs is known, there exists a subalgebra $\cal U$ (MVE-algebra) of the basic $σ$-algebra such that all $\cal U$-measurable statistics with finite second moments are UMVUEs. It is shown that MVE-algebras are, in a sense, similar to the subalgebras generated by complete sufficient statistics. Examples are given when these subalgebras differ, in these cases a new statistical structure arises.

math.ST↗

A note on the minimax solution for the two-stage group testing problem

Group testing is an active area of current research and has important applications in medicine, biotechnology, genetics, and product testing. There have been recent advances in design and estimation, but the simple Dorfman procedure introduced by R. Dorfman in 1943 is widely used in practice. In many practical situations the exact value of the probability p of being affected is unknown. We present both minimax and Bayesian solutions for the group size problem when p is unknown. For unbounded p we show that the minimax solution for group size is 8, while using a Bayesian strategy with Jeffreys prior results in a group size of 13. We also present solutions when p is bounded from above. For the practitioner we propose strong justification for using a group size of between eight to thirteen when a constraint on p is not incorporated and provide useable code for computing the minimax group size under a constrained p.

math.ST↗

Partially complete sufficient statistics are jointly complete

The theory of the basic statistical concept of (Lehmann-Scheffé-)completeness is perfected by providing the theorem indicated in the title and previously overlooked for several decades. Relations to earlier results are discussed and illustrating examples are presented. Of the two proofs offered for the main result, the first is direct and short, following the prototypical example of Landers and Rogge (1976), and the second is very short and purely statistical, utilizing the basic theory of optimal unbiased estimation in the little known version completed by Schmetterer and Strasser (1974).

math.ST↗

On the Nile Problem by Sir Ronald Fisher

The Nile problem by Ronald Fisher may be interpreted as the problem of making statistical inference for a special curved exponential family when the minimal sufficient statistic is incomplete. The problem itself and its versions for general curved exponential families pose a mathematical-statistical challenge: studying the subalgebras of ancillary statistics within the $σ$-algebra of the (incomplete) minimal sufficient statistics and closely related questions of the structure of UMVUEs. In this paper a new method is developed that, in particular, proves that in the classical Nile problem no statistic subject to mild natural conditions is a UMVUE. The method almost solves an old problem of the existence of UMVUEs. The method is purely statistical (vs. analytical) and works for any family possessing an ancillary statistic. It complements an analytical method that uses only the first order ancillarity (and thus works when the existence of ancillary subalgebras is an open problem) and works for curved exponential families with polynomial constraints on the canonical parameters of which the Nile problem is a special case.

math.ST↗