Comment: Classifier Technology and the Illusion of Progress
Comment on Classifier Technology and the Illusion of Progress [math.ST/0606441]
SEARCH · Searcharxiv
Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Comment on Classifier Technology and the Illusion of Progress [math.ST/0606441]
Comment on Classifier Technology and the Illusion of Progress--Credit Scoring [math.ST/0606441]
Rejoinder: Classifier Technology and the Illusion of Progress [math.ST/0606441]
Comment on The Place of Death in the Quality of Life [math.ST/0612783]
Comment on Causal Inference in the Medical Area [math.ST/0612783]
Comment on Complex Causal Questions Require Careful Model Formulation: Discussion of Rubin on Experiments with ``Censoring'' Due to Death [math.ST/0612783]
Rejoinder on Causal Inference Through Potential Outcomes and Principal Stratification: Application to Studies with ``Censoring'' Due to Death by D. B. Rubin [math.ST/0612783]
Comment on "Support Vector Machines with Applications" [math.ST/0612817]
Comment on "Support Vector Machines with Applications" [math.ST/0612817]
Comment on ``Support Vector Machines with Applications'' [math.ST/0612817]
Rejoinder to ``Support Vector Machines with Applications'' [math.ST/0612817]
In ref [math.ST/0411462] the notion of statistically dual distributions is introduced. The reconstruction of confidence density [AIP Conference Proceedings 803 (2005) 398] for the location parameter for several pairs of statistically dual distributions (Poisson and Gamma, normal and normal, Cauchy and Cauchy, Laplace and Laplace) in the case of single observation of the random variable is a unique. It allows to introduce the Transform between the space of observed values and the space of possible values of the parameter.
These notes provide a pedagogical introduction to the role of transversality theory in the analysis of statistical degeneracies within the framework of distributional statistical models. The classical question of when a statistical model is well-behaved - in the sense of being identifiable, having non-singular Fisher information, and admitting robust estimation - is reformulated as a question about the geometry of a kernel-induced feature map. Statistical pathologies correspond to geometric degeneracies of this map, and transversality theory provides a precise language for understanding when and why such degeneracies are non-generic. The exposition is organised in three parts. Part I surveys the statistical phenomena that motivate the geometric treatment: representation failure, non-identifiability, moment indeterminacy, singular information, nuisance parameters, and the Behrens-Fisher problem. Part II develops the necessary geometric toolkit - smooth maps, Sard's theorem, transversality, jets, stratifications, and the parametric transversality theorem - at a level accessible to students with a background in analysis and linear algebra but no prior exposure to differential topology. Part~III returns to the statistical problems of Part~I and shows how each one admits a unified geometric interpretation as a transversality condition on the feature map. These notes are a pedagogical companion to the research paper Labouriau (2026) "Transversality and Geometric Regularisation in Distributional Statistical Models" (arXiv:2605.04536 [math.ST]), expanding its arguments with motivating examples, geometric intuition, and exercises aimed at advanced Master's and PhD students with a background in mathematical statistics and measure theory. They are designed to support seminars or reading groups.
In this paper, we consider simultaneous estimation of Poisson parameters in situations where we can use side information in aggregated data. We use standardized squared error and entropy loss functions. Bayesian shrinkage estimators are derived based on conjugate priors. We compare the risk functions of direct estimators and Bayesian estimators with respect to different priors that are constructed based on different subsets of observations. We obtain conditions for domination and also prove minimaxity and admissibility in a simple setting.
We propose novel necessary and sufficient conditions for a sensing matrix to be "$s$-good" - to allow for exact $\ell_1$-recovery of sparse signals with $s$ nonzero entries when no measurement noise is present. Then we express the error bounds for imperfect $\ell_1$-recovery (nonzero measurement noise, nearly $s$-sparse signal, near-optimal solution of the optimization problem yielding the $\ell_1$-recovery) in terms of the characteristics underlying these conditions. Further, we demonstrate (and this is the principal result of the paper) that these characteristics, although difficult to evaluate, lead to verifiable sufficient conditions for exact sparse $\ell_1$-recovery and to efficiently computable upper bounds on those $s$ for which a given sensing matrix is $s$-good. We establish also instructive links between our approach and the basic concepts of the Compressed Sensing theory, like Restricted Isometry or Restricted Eigenvalue properties.
Network detection is an important capability in many areas of applied research in which data can be represented as a graph of entities and relationships. Oftentimes the object of interest is a relatively small subgraph in an enormous, potentially uninteresting background. This aspect characterizes network detection as a "big data" problem. Graph partitioning and network discovery have been major research areas over the last ten years, driven by interest in internet search, cyber security, social networks, and criminal or terrorist activities. The specific problem of network discovery is addressed as a special case of graph partitioning in which membership in a small subgraph of interest must be determined. Algebraic graph theory is used as the basis to analyze and compare different network detection methods. A new Bayesian network detection framework is introduced that partitions the graph based on prior information and direct observations. The new approach, called space-time threat propagation, is proved to maximize the probability of detection and is therefore optimum in the Neyman-Pearson sense. This optimality criterion is compared to spectral community detection approaches which divide the global graph into subsets or communities with optimal connectivity properties. We also explore a new generative stochastic model for covert networks and analyze using receiver operating characteristics the detection performance of both classes of optimal detection techniques.
We show that the class of conditional distributions satisfying the coarsening at random (CAR) property for discrete data has a simple and robust algorithmic description based on randomized uniform multicovers: combinatorial objects generalizing the notion of partition of a set. However, the complexity of a given CAR mechanism can be large: the maximal "height" of the needed multicovers can be exponential in the number of points in the sample space. The results stem from a geometric interpretation of the set of CAR distributions as a convex polytope and a characterization of its extreme points. The hierarchy of CAR models defined in this way could be useful in parsimonious statistical modeling of CAR mechanisms, though the results also raise doubts in applied work as to the meaningfulness of the CAR assumption in its full generality.