SearcharxivSearch

SEARCH · Searcharxiv

Results for “stat.OT”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Markovian growth dynamics on rooted binary trees evolving according to the Gompertz curve

Inspired by biological dynamics, we consider a growth Markov process taking values on the space of rooted binary trees, similar to the Aldous-Shields model. Fix $n\ge 1$ and $β>0$. We start at time 0 with the tree composed of a root only. At any time, each node with no descendants, independently from the other nodes, produces two successors at rate $β(n-k)/n$, where $k$ is the distance from the node to the root. Denote by $Z_n(t)$ the number of nodes with no descendants at time $t$ and let $T_n = β^{-1} n \ln(n /\ln 4) + (\ln 2)/(2 β)$. We prove that $2^{-n} Z_n(T_n + n τ)$, $τ\in\bb R$, converges to the Gompertz curve $\exp (- (\ln 2) e^{-βτ})$. We also prove a central limit theorem for the martingale associated to $Z_n(t)$.

q-bio.CB

"Not only defended but also applied": The perceived absurdity of Bayesian inference

The missionary zeal of many Bayesians of old has been matched, in the other direction, by a view among some theoreticians that Bayesian methods are absurd-not merely misguided but obviously wrong in principle. We consider several examples, beginning with Feller's classic text on probability theory and continuing with more recent cases such as the perceived Bayesian nature of the so-called doomsday argument. We analyze in this note the intellectual background behind various misconceptions about Bayesian statistics, without aiming at a complete historical coverage of the reasons for this dismissal.

math.ST

Cognitive Constructivism and the Epistemic Significance of Sharp Statistical Hypotheses in Natural Sciences

This book presents our case in defense of a constructivist epistemological framework and the use of compatible statistical theory and inference tools. The basic metaphor of decision theory is the maximization of a gambler's expected fortune, according to his own subjective utility, prior beliefs an learned experiences. This metaphor has proven to be very useful, leading the development of Bayesian statistics since its XX-th century revival, rooted on the work of de Finetti, Savage and others. The basic metaphor presented in this text, as a foundation for cognitive constructivism, is that of an eigen-solution, and the verification of its objective epistemic status. The FBST - Full Bayesian Significance Test - is the cornerstone of a set of statistical tolls conceived to assess the epistemic value of such eigen-solutions, according to their four essential attributes, namely, sharpness, stability, separability and composability. We believe that this alternative perspective, complementary to the one ofered by decision theory, can provide powerful insights and make pertinent contributions in the context of scientific research.

stat.OT

Statistical Inference in Dynamic Treatment Regimes

Dynamic treatment regimes are of growing interest across the clinical sciences as these regimes provide one way to operationalize and thus inform sequential personalized clinical decision making. A dynamic treatment regime is a sequence of decision rules, with a decision rule per stage of clinical intervention; each decision rule maps up-to-date patient information to a recommended treatment. We briefly review a variety of approaches for using data to construct the decision rules. We then review an interesting challenge, that of nonregularity that often arises in this area. By nonregularity, we mean the parameters indexing the optimal dynamic treatment regime are nonsmooth functionals of the underlying generative distribution. A consequence is that no regular or asymptotically unbiased estimator of these parameters exists. Nonregularity arises in inference for parameters in the optimal dynamic treatment regime; we illustrate the effect of nonregularity on asymptotic bias and via sensitivity of asymptotic, limiting, distributions to local perturbations. We propose and evaluate a locally consistent Adaptive Confidence Interval (ACI) for the parameters of the optimal dynamic treatment regime. We use data from the Adaptive Interventions for Children with ADHD study as an illustrative example. We conclude by highlighting and discussing emerging theoretical problems in this area.

stat.ME

Hotelling's test for highly correlated data

This paper is motivated by the analysis of gene expression sets, especially by finding differentially expressed gene sets between two phenotypes. Gene $\log_2$ expression levels are highly correlated and, very likely, have approximately normal distribution. Therefore, it seems reasonable to use two-sample Hotelling's test for such data. We discover some unexpected properties of the test making it different from the majority of tests previously used for such data. It appears that the Hotelling's test does not always reach maximal power when all marginal distributions are differentially expressed. For highly correlated data its maximal power is attained when about a half of marginal distributions are essentially different. For the case when the correlation coefficient is greater than 0.5 this test is more powerful if only one marginal distribution is shifted, omparing to the case when all marginal distributions are equally shifted. Moreover, when the correlation coefficient increases the power of Hotelling's test increases as well.

stat.OT

Development and Initial Validation of a Scale to Measure Instructors' Attitudes toward Concept-Based Teaching of Introductory Statistics in the Health and Behavioral Sciences

Despite more than a decade of reform efforts, students continue to experience difficulty understanding and applying statistical concepts. The predominant focus of reform has been on content, pedagogy, technology and assessment, with little attention to instructor characteristics. However, there is strong theoretical and empirical evidence that instructors' attitudes impact the quality of teaching and learning. The objective of this study was to develop and initially validate a scale to measure instructors' attitudes toward reform-oriented (or concept-based) teaching of introductory statistics in the health and behavioral sciences, at the tertiary level. This scale will be referred to as FATS (Faculty Attitudes Toward Statistics). Data were obtained from 227 instructors (USA and international), and analyzed using factor analysis, multidimensional scaling and hierarchical cluster analysis. The overall scale consists of five sub-scales with a total of 25 items, and an overall alpha of 0.89. Construct validity was established. Specifically, the overall scale, and subscales (except perceived difficulty) plausibly differentiated between low-reform and high-reform practice instructors. Statistically significant differences in attitude were observed with respect to age, but not gender, employment status, membership status in professional organizations, ethnicity, highest academic qualification, and degree concentration. This scale can be considered a reliable and valid measure of instructors' attitudes toward reform-oriented (concept-based or constructivist) teaching of introductory statistics in the health and behavioral sciences at the tertiary level. These five dimensions influence instructors' attitudes. Additional studies are required to confirm these structural and psychometric properties.

stat.OT

$L_p$-nested symmetric distributions

Tractable generalizations of the Gaussian distribution play an important role for the analysis of high-dimensional data. One very general super-class of Normal distributions is the class of $ν$-spherical distributions whose random variables can be represented as the product $\x = r\cdot \u$ of a uniformly distribution random variable $\u$ on the $1$-level set of a positively homogeneous function $ν$ and arbitrary positive radial random variable $r$. Prominent subclasses of $ν$-spherical distributions are spherically symmetric distributions ($ν(\x)=\|\x\|_2$) which have been further generalized to the class of $L_p$-spherically symmetric distributions ($ν(\x)=\|\x\|_p$). Both of these classes contain the Gaussian as a special case. In general, however, $ν$-spherical distributions are computationally intractable since, for instance, the normalization constant or fast sampling algorithms are unknown for an arbitrary $ν$. In this paper we introduce a new subclass of $ν$-spherical distributions by choosing $ν$ to be a nested cascade of $L_p$-norms. This class is still computationally tractable, but includes all the aforementioned subclasses as a special case. We derive a general expression for $L_p$-nested symmetric distributions as well as the uniform distribution on the $L_p$-nested unit sphere, including an explicit expression for the normalization constant. We state several general properties of $L_p$-nested symmetric distributions, investigate its marginals, maximum likelihood fitting and discuss its tight links to well known machine learning methods such as Independent Component Analysis (ICA), Independent Subspace Analysis (ISA) and mixed norm regularizers. Finally, we derive a fast and exact sampling algorithm for arbitrary $L_p$-nested symmetric distributions, and introduce the Nested Radial Factorization algorithm (NRF), which is a form of non-linear ICA.

stat.OT

Tests of Non-Equivalence among Absolutely Nonsingular Tensors through Geometric Invariants

4x4x3 absolutely nonsingular tensors are characterized by their determinant polynomial. Non-quivalence among absolutely nonsingular tensors with respect to a class of linear transformations, which do not chage the tensor rank,is studied. It is shown theoretically that affine geometric invariants of the constant surface of a determinant polynomial is useful to discriminate non-equivalence among absolutely nonsingular tensors. Also numerical caluculations are presented and these invariants are shown to be useful indeed. For the caluculation of invarinats by 20-spherical design is also commented. We showed that an algebraic problem in tensor data analysis can be attacked by an affine geometric method.

stat.OT

Measurement error and deconvolution in spaces of generalized functions

This paper considers convolution equations that arise from problems such as measurement error and non-parametric regression with errors in variables with independence conditions. The equations are examined in spaces of generalized functions to account for possible singularities; this makes it possible to consider densities for arbitrary and not only absolutely continuous distributions, and to operate with Fourier transforms for polynomially growing regression functions. Results are derived for identification and well-posedness in the topology of generalized functions for the deconvolution problem and for some regression models. Conditions for consistency of plug-in estimation for these models are derived.

math.ST

A brief history of the Fail Safe Number in Applied Research

Rosenthal's (1979) Fail-Safe-Number (FSN) is probably one of the best known statistics in the context of meta-analysis aimed to estimate the number of unpublished studies in meta-analyses required to bring the meta-analytic mean effect size down to a statistically insignificant level. Already before Scargle's (2000) and Schonemann & Scargle's (2008) fundamental critique on the claimed stability of the basic rationale of the FSN approach, objections focusing on the basic assumption of the FSN which treats the number of studies as unbiased with averaging null were expressed throughout the history of the FSN by different authors (Elashoff, 1978; Iyengar & Greenhouse, 1988a; 1988b; see also Scargle, 2000). In particular, Elashoff's objection appears to be important because it was the very first critique pointing directly to the central problem of the FSN: "R & R claim that the number of studies hidden in the drawers would have to be 65,000 to achieve a mean effect size of zero when combined with the 345 studies reviewed here. But surely, if we allowed the hidden studies to be negative, on the average no more than 345 hidden studies would be necessary to obtain a zero mean effect size" (p. 392). Thus, users of meta-analysis could have been aware right from the beginning that something was wrong with the statistical reasoning of the FSN. In particular, from an applied research perspective, it is therefore of interest whether any of the fundamental objections on the FSN are reflected in standard handbooks on meta-analysis as well as -and of course even more importantly- in meta-analytic studies itself.

stat.OT

Degrees of Equivalence in a Key Comparison

In an interlaboratory key comparison, a data analysis procedure for this comparison was proposed and recommended by CIPM [1, 2, 3], therein the degrees of equivalence of measurement standards of the laboratories participated in the comparison and the ones between each two laboratories were introduced but a corresponding clear and plausible measurement model was not given. Authors in [4] offered possible measurement models for a given comparison and a suitable model was selected out after rigorous analyzing steps for expectation values of these degrees of equivalence. The systematic laboratory-effects model was then selected as a right one in this report. Those models were all based on the one true value existence assumption. However in the year 2008, a new version of the Vocabulary for International Metrology (VIM) [7] was issued where the true value of a given measurement standard should be now perceived as multi true values which following a given statistics distribution. Applying this perception of true values of a measurement standard with combination of the steps in [4], measurement models have been developed and degrees of equivalence have been analyzed. The results show that although with new definition, the systematic laboratory-effects model is still the reasonable one in a given key comparison.

stat.OT

Inherent Difficulties of Non-Bayesian Likelihood-based Inference, as Revealed by an Examination of a Recent Book by Aitkin

For many decades, statisticians have made attempts to prepare the Bayesian omelette without breaking the Bayesian eggs; that is, to obtain probabilistic likelihood-based inferences without relying on informative prior distributions. A recent example is Murray Aitkin's recent book, {\em Statistical Inference}, which presents an approach to statistical hypothesis testing based on comparisons of posterior distributions of likelihoods under competing models. Aitkin develops and illustrates his method using some simple examples of inference from iid data and two-way tests of independence. We analyze in this note some consequences of the inferential paradigm adopted therein, discussing why the approach is incompatible with a Bayesian perspective and why we do not find it relevant for applied work.

stat.ME

A Unified MGF-Based Capacity Analysis of Diversity Combiners over Generalized Fading Channels

Unified exact average capacity results for L-branch coherent diversity receivers including equal-gain combining (EGC) and maximal-ratio combining (MRC) are not known. This paper develops a novel generic framework for the capacity analysis of $L$-branch EGC/MRC over generalized fading channels. The framework is used to derive new results for the Gamma shadowed generalized Nakagami-m fading model which can be a suitable model for the fading environments encountered by high frequency (60 GHz and above) communications. The mathematical formalism is illustrated with some selected numerical and simulation results confirming the correctness of our newly proposed framework.

cs.IT

Squaring the Circle and Cubing the Sphere: Circular and Spherical Copulas

Do there exist circular and spherical copulas in $R^d$? That is, do there exist circularly symmetric distributions on the unit disk in $R^2$ and spherically symmetric distributions on the unit ball in $R^d$, $d\ge3$, whose one-dimensional marginal distributions are uniform? The answer is yes for $d=2$ and 3, where the circular and spherical copulas are unique and can be determined explicitly, but no for $d\ge4$. A one-parameter family of elliptical bivariate copulas is obtained from the unique circular copula in $R^2$ by oblique coordinate transformations. Copulas obtained by a non-linear transformation of a uniform distribution on the unit ball in $R^d$ are also described, and determined explicitly for $d=2$.

stat.OT

Simultaneous concentration of order statistics

Let $μ$ be a probability measure on $\mathbb{R}$ with cumulative distribution function $F$, $(x_{i})_{1}^{n}$ a large i.i.d. sample from $μ$, and $F_{n}$ the associated empirical distribution function. The Glivenko-Cantelli theorem states that with probability 1, $F_{n}$ converges uniformly to $F$. In so doing it describes the macroscopic structure of $\{x_{i}\}_{1}^{n}$, however it is insensitive to the position of individual points. Indeed any subset of $o(n)$ points can be perturbed at will without disturbing the convergence. We provide several refinements of the Glivenko-Cantelli theorem which are sensitive not only to the global structure of the sample but also to individual points. Our main result provides conditions that guarantee simultaneous concentration of all order statistics. The example of main interest is the normal distribution.

math.PR