Searcharxiv⌕ Search

arXiv subjects

Nils Lid Hjort

Publications and source records attributed to Nils Lid Hjort.

At least 19 recordsLinked to original sources

How many cards, until the first ace: variations, extensions, lachrymae, confidence, Dirichlets

From a deck of cards, how many cards do I need to draw, until the first ace? I identify the distribution for this waiting time $T$, and its satisfyingly nice expected value ${\rm E}\,T=(N+1)/(n+1)$, with $N$ the number of cards and $n$ the number of aces; hence $53/5=10.6$ for the standard setup. After having solved this Question One I go on to certain alternative solutions and extensions, involving e.g. Beta approximations. I also consider the distributions and means for the 2nd, the 3rd, the 4th occurrences of aces, with generalisations, where there is a Dirichlet distribution in wait for us, with further links to order statistics for the uniform. Furthermore, an apparatus is developed for obtaining estimators and full confidence distributions for applications where one knows the number $n$ of aces, but not the deck size $N$; and correspondingly for inference about the unknown population size $N$ when $n$ is known. If you have 1000 people in a room, and need to interview 11 of them until you've found the first left-handed person, how may left-handed are there in the room -- here we need both an estimate and a clear measure of uncertainty.

stat.OT↗

Nonparametric quantile inference using Dirichlet processes

This chapter deals with nonparametric inference for quantiles from a Bayesian perspective, using the Dirichlet process. The posterior distribution for quantiles is characterised, enabling also explicit formulae for posterior mean and variance. Unlike the Bayes estimator for the distribution function, our Bayes estimator for the quantile function is a smooth curve. A Bernstein--von Mises type theorem is given, exhibiting the limiting posterior distribution of the quantile process. Links to kernel-smoothed quantile estimators are provided. As a side product we develop an automatic nonparametric density estimator, free of smoothing parameters, with support exactly matching that of the data range. Nonparametric Bayes estimators are also provided for other quantile-related quantities, including the Lorenz curve and the Gini index, for Doksum's shift curve and for Parzen's comparison distribution in two-sample situations, and finally for the quantile regression function in situations with covariates.

math.ST↗

Focused Information Criteria

The focused information criterion is used to make a choice among several statistical models, or among several variables to include in a model. Different from other such information criteria, the focused information criterion is constructed to select the best model for a given interest quantity, the focus of the research question. Different such focus parameters may lead to different selected models, each one best for the corresponding focus. What is `best' is defined by a risk function, often the mean squared error. Other risks can be considered too for focused selection. Selections by the focused information criterion include using parametric (generalized) linear models, non- and semiparametric models, quantile regression models, graphical models, models for survival data, for longitudinal data, time series models, and many more. Extensions of the basic version include versions for high-dimensional data, regularized estimation, and Bayesian methods.

stat.ME↗

Notes on the Theory of Statistical Symbol Recognition

This document is a pdf generated from old plain-TeX files of 1986, of Nils Lid Hjort's `Notes on the Theory of Statistical Symbol Recogntion', a limited circulation 207-pages monograph published at the Norwegian Computing Centre, as Report no. 778/1986. It gives the basics of the statistical pattern recognition theory developed to suit the needs of several industrial projects, related to contracts with the Norwegian-German firm SysScan, the Royal Norwegian Research Council, and yet others, involving symbol recognition and classification analysis from noisy images, related to maps, documents, satellite imaging, etc. The methods and algorithms developed also needed to fit the technology of that time, anno c. 1986, with machines scanning documents, converting these to vector representation, within computational and machine system boundaries. There is an accompanying and also limited circulation booklet, `Statistical Symbol Recognition: Development of a System', by Knut Bråten, Erik Holbæk-Hanssen, and Torfinn Taxt (Report No.777/1986, Norwegian Computing Centre, Oslo), detailing the system developments. Thus developments took form and shape on two frontiers, in close collaboration, Hjort's statistical methods and getting the technology to work, with its multiple components, hardware and software.

math.ST↗

Counting the uncounted: How many were killed in Guatemala, 1978-1995?

In various application domains, there is a certain `null cell', inside a multinomial setup, where observations are recorded for the other cells, but where one cannot count the number of occurrences for the null cell. I develop inference theory for assessing such unknown numbers, counting the uncounted, in situations where counts are available for the other cells, via parametric modelling. The methods are used to estimate the number of persons killed in Guatemala during the Genocidio guatemalteco years 1978--1995. There are three carefully curated lists of killed people, where the information can be mapped to a Venn diagram with $2^3=8$ cells. Summing over the seven observed cells, $R=\hbox{47,803}$ killed individuals can be identified, but how big is $N_{0,0,0}$, and hence $N=N_{0,0,0}+R$?

stat.AP↗

On multiplicative bias correction in kernel density estimation

Hjort and Glad (1995) present a method for semiparametric density estimation. Relative to the ordinary kernel density estimator, this technique performs much better when a parametric vehicle distribution fits the data, and otherwise performs at broadly the same level. Jones, Linton, and Nielsen (1995) present a somewhat similar method for density estimation which has higher order bias for all sufficiently smooth densities. In this paper, we combine the two methods. We show that, theoretically, the desired properties of general higher order bias allied with even better performance for an appropriate vehicle model are achieved. Simulations suggest that the new estimator realises only a little of its theoretical potential in practice for small to moderately large sample sizes.

math.ST↗

Estimating the logistic regression equation when the model is incorrect

Protesting mildly against the notion of an exactly correct parametric model the view is adopted that the logistic regression equation is merely an approximation to the underlying, true function. The behaviour of likelihood based estimators is investigated in such a general framework. The maximum likelihood estimator is shown to be consistent for a certain least false parameter value minimising a weighted average of quantities that measure the distance from the true to the parametric model. Asymptotic normality is also demonstrated. Finally a number of additional remarks are offered, some pointing to natural generalisations and some to new questions for research, like weighted and local likelihood estimation methods.

math.ST↗

Post-Processing Posterior Predictive P-values

This article addresses issues of model criticism and model comparison in Bayesian contexts, and focusses on the use of the so-called posterior predictive p-values (ppp values). These involve a general discrepancy or conflict measure and depend on the prior, the model, and the data. They are used in statistical practice to quantify the degree of surprise or conflict in data, and for purposes of comparing different combinations of prior and model. The distribution of such ppp values is however far from uniform, as we demonstrate for different models, making their interpretation and comparison a difficult matter. We propose a natural calibration of the ppp values, where the resulting cppp values are uniform on the unit interval under model conditions. The cppp values, which in general rely on a double simulation scheme for their computation, may then be used to assess and compare different priors and models. Our methods also make it possible to compare parametric with nonparametric model specifications, in that genuine `measures of surprise' are put on the same canonical uniform scale. Our techniques are illustrated for some applications to real data. We also present supplementing theoretical results on various properties of the ppp and cppp.

stat.ME↗

Bayesian Nonparametrics: Principles and Practice

This extended preface [to the Book `Bayesian Nonparametrics', Cambridge University Press, 2010, by NL Hjort, CC Holmes, P Mueller, SG Walker] is meant to explain why you are right to be curious about Bayesian nonparametrics -- why you may actually need it and how you can manage to understand it and use it. The preface also serves as an introductory chapter, giving an overview of the aims and contents of the book. We also explain the background for how the book came into existence, delve briefly on the history of the still relatively young field of Bayesian nonparametrics, and offer some concluding remarks, pertaining to various challenges and likely future developments of the area.

stat.ME↗

Topics in Nonparametric Bayesian Statistics

The intersection set of Bayesian and nonparametric statistics was almost empty until about 1973, but now is growing at a healthy rate. This chapter, for the {\it Highly Structured Stochastic Systems} book (Oxford University Press, 2003) gives an overview of various theoretical and applied research themes inside this field, partly complementing and extending recent reviews of Dey, M{ü}ller and Sinha (1998) and Walker, Damien, Laud and Smith (1999). The intention is not to be complete or exhaustive, but rather to touch on research areas of interest, partly by example.

stat.ME↗

Modelling pairs of Poissons and binomials with negative correlation

Suppose $f_1(x)$ and $f_2(y)$ are given marginals for pairs $(x,y)$. I consider the construction $f_1(x)f_2(y)\{ 1+αh_1(x)h_2(y) \}$, where $h_1$ and $h_2$ are seen as bounded adjustment functions, normalised to have means zero under $f_1$ and $f_2$. This defines a bivariate distribution for $(X,Y)$ with the specified marginal densities $f_1$ and $f_2$, with an interval of permissible values of $α$, both positive and negative; in particular, independence corresponds to an innter point in the adjustments parameter region. Applications to bivariate Poisson distributions, allowing both positive and negative correlation, are discussed. As illustration I provide a more accurate and extended analysis of a Poisson pairs dataset, pertaining to competing seeds and plants, for $n=958$ plots of soil, earlier analysed in the well-cited paper Lakshminarayana, Pandit, Rao, Srinivasa (1999). The general apparatus is also shown to work for negatively correlated binomials. Those methods are illustrated in a meta-analysis framework for two-by-two tables across different studies, pertaining to the Audit-C screening questionnaire for alcohol use disorders, where again negative correlation is demonstrated, between $X$, the number of correct `yes', and $Y$, the number of correct `no'.

stat.ME↗

Recent advances in statistical methodology applied to the Hjort liver index time series (1859-2012) and associated influential factors

Certain recent advances in statistical methodology have promising potential for fruitful use in general biology and the fisheries sciences. This paper reviews and discusses some of the relevant themes, including accurate modelling via focused model selection techniques, dynamic goodness-of-fit testing of processes evolving over time, finding break points for phenomena experiencing changes, prediction uncertainty, and optimal combination of information across diverse sources via confidence distributions. The methods are illustrated for the Hjort liver quality index time series. Its roots lie in the classic Hjort (`Fluctuations in the Great Fisheries of Northern Europe, Viewed in the Light of Biological Research', 1914), where liver quality of the Atlantic cod {\it (Gadus morhua)} for 1880--1912 is reported on and studied, along with related factors, making it one of the first teleost time series ever published. Diligent work by Kjesbu et al. (`Making use of Johan Hjort's `unknown' legacy: reconstruction of a 150-year coastal time-series on northeast Arctic cod (Gadus morhua) liver data reveals long-term trends in energy allocation patterns', 2014), involving both archival and calibration efforts, have extended the series both backwards and forwards in time, to 1859--2012, yielding one of the longest time series of marine science. Our study offers a detailed examination of this series and how it relates to and interacts with associated factors, including Kola winter temperatures, length distribution parameters, cod mortality, and a certain index related to availability of food.

stat.AP↗

Bayesian and Empirical Bayesian Bootstrapping

Let $X_1,\ldots,X_n$ be a random sample from an unknown probability distribution $P$ on the sample space ${\cal X}$, and let $θ=θ(P)$ be a parameter of interest. The present paper proposes a nonparametric `Bayesian bootstrap' method of obtaining Bayes estimates and Bayesian confidence limits for $θ$. It uses a simple simulation technique to numerically approximate the exact posterior distribution of $θ$ using a (non-degenerate) Dirichlet process prior for $P$. Asymptotic arguments are given which justify the use of the Bayesian bootstrap for any smooth functional $θ(P)$. When the prior is fixed and the sample size grows five approaches become first-order equivalent: the exact Bayesian, the Bayesian bootstrap, Rubin's degenerate-prior bootstrap, Efron's bootstrap, and the classical one using delta methods. The Bayesian bootstrap method is also extended to the semiparametric regression case. A separate section treats similar ideas for censored data and for more general hazard rate models, where a connection is made to a `weird bootstrap' proposed by Gill. Finally empirical Bayesian versions of the procedure are discussed, where suitable parameters of the Dirichlet process prior are inferred from data. Our results lend Bayesian support to the classic Efron bootstrap. It is the Bayesian bootstrap under a noninformative reference prior; it is a limit of natural approximations to good Bayes solutions; it is an approximation to a natural empirical Bayesian strategy; and the formally incorrect reading of a bootstrap histogram as a posterior distribution for the parameter isn't so incorrect after all.

math.ST↗

Density Estimation Using the Sinc Kernel

This paper deals with the kernel density estimator based on the so-called sinc (or Fourier integral) kernel $K(x)=(πx)^{-1}\sin x$. We study in detail both asymptotic and finite sample properties of this estimator. It is shown that, contrary to widespread opinion, the sinc estimator is superior to other estimators in many respects: it is more accurate for quite moderate values of the sample size, has better asymptotics in non-smooth case (the density to be estimated has only first derivative), is more convenient for the bandwidth selection, etc.

math.ST↗

Tests for constancy of model parameters Over time

Suppose that a sequence of data points follows a distribution of a certain parametric form, but that one or more of the underlying parameters may change over time. This paper addresses various natural questions in such a framework. We construct canonical monitoring processes which under the hypothesis of no change converge in distribution to independent Brownian bridges, and use these to construct natural goodness-of-fit statistics. Weighted versions of these are also studied, and optimal weight functions are derived to give maximum local power against alternatives of interest. We also discuss how our results can be used to pinpoint where and what type of changes have occurred, in the event that initial screening tests indicate that such exist. Our unified large-sample methodology is quite general and applies to all regular parametric models, including regression, Markov chains, and time series situations.

stat.ME↗

Nonparametric density estimation with a parametric start

The traditional kernel density estimator of an unknown density is by construction completely nonparametric, in the sense that it has no preferences and will work reasonably well for all shapes. The present paper develops a class of semiparametric methods that are designed to work better than the kernel estimator in a broad nonparametric neighbourhood of a given parametric class of densities, for example the normal, while not losing much in precision when the true density is far from the parametric class. The idea is to multiply an initial parametric density estimate with a kernel type estimate of the necessary correction factor. This works well in cases where the correction factor function is less rough than the original density itself. Extensive comparisons with the kernel estimator are carried out, including exact analysis for the class of all normal mixtures. The new method, with a normal start, wins quite often, even in many cases where the true density is far from normal. Procedures for choosing the smoothing parameter of the estimator are also discussed. The new estimator should be particularly useful in higher dimensions, where the usual nonparametric methods have problems. The idea is also spelled out for nonparametric regression.

stat.ME↗

Sudoku Solving and Finding Magic Squares by Probability Models and Markov Chains

The sudoku puzzles have a long history, with variations going back more than a hundred years, but its current and perhaps surprising world-wide prominence goes back to certain initiatives and then puzzle-generating computer programmes from just after 2000. To solve a sudoko puzzle, a statistician can put up a probabilitymodel on the enormous space of $9\times9$ matrix possibilities, constructed to favour `good attempts', and then engineer a Markov chain to sample a long enough chain of sudoku table realisations from that model, until the solution is found. The methods work also for other types of puzzles, like constructing `magic squares' with wished-for properties (sums of rows, columns, diagonals equal, etc.), as is also illustrated in this article; via magic models and equally magic Markov chains I find impressively magic $8\times8$ and $10\times10$ squares.

stat.OT↗

Focused Information Criteria for the Linear Hazard Regression Model

The linear hazard regression model developed by Aalen is becoming an increasingly popular alternative to the Cox multiplicative hazard regression model. There are no methods in the literature for selecting among different candidate models of this nonparametric type, however. In the present paper a focused information criterion is developed for this task. The criterion works for each specified covariate vector, by estimating the mean squared error for each candidate model's estimate of the associated cumulative hazard rate; the finally selected model is the one with lowest estimated mean squared error. Averaged versions of the criterion are also developed.

stat.ME↗