SearcharxivSearch

arXiv subjects

Nicholas G. Polson

Publications and source records attributed to Nicholas G. Polson.

At least 19 recordsLinked to original sources

Riemann, Thorin, van Dantzig Pairs, Wald Couples and Hadamard Factorisation

The Hadamard-Weierstrass factorisation of an entire function in the Laguerre-Pólya class is dual to a pair of probabilistic objects: a van Dantzig pair of characteristic functions and a Wald couple of infinitely divisible random variables. The reciprocal of such a function is the Laplace transform of a generalised gamma convolution (GGC) whose Thorin measure encodes the zeros, and Thorin's condition, a real Laplace identity on $(0,\infty)$, is equivalent to the absence of zeros off the critical axis. We prove this duality, together with a closed-form formula for the Thorin density in terms of any Lévy representation. The framework is then applied to the Gamma, hyperbolic, Bessel and Macdonald functions, to the Riemann $ξ$-function, to Dirichlet and modular $L$-functions, and to Dedekind's $η$. For Ramanujan's $τ$ it is set up but not carried through, since reality of the zeros of the associated $L$-function is open. For $ξ$ we prove, unconditionally, that the heat trace $W(t)=\tfrac12\sum_ρ\exp\{(ρ-\tfrac12)^{2}t\}$ is positive and strictly decreasing on $(0,\infty)$, so that $ξ(\tfrac12)/ξ(\tfrac12+\sqrt{u})$ is the Laplace transform of a self-decomposable law, and that $ξ(α)/ξ(α+\sqrt{s})$ is the Laplace transform of a GGC for every $α$ exceeding the real part of every non-trivial zero, hence unconditionally for every $α\ge1$. Complete monotonicity of $W$, equivalently the GGC property at the centre, is shown to be equivalent to the Riemann hypothesis. Self-decomposability is therefore the unconditional ceiling.

math.PR

On Hilbert's 8th Problem

Let $W(t)=\tfrac12\sum_ρ\exp\{(ρ-\tfrac12)^2t\}$ be the Xi heat trace, with zeros counted with multiplicity. We prove that verification of the zeros through height $T_N=20\,000N\sqrt{1+\log N}$ suffices for positive definiteness of both leading $N\times N$ heat trace Hankel blocks at every positive time. For the three by three blocks, height $10\,000$ suffices. Published rigorous zero verification makes the conclusion unconditional for $N\le30\,000\,000$. A separate Cauchy estimate for the prime kernel proves $(-1)^kW^{(k)}>0$ through order $10^{17}$ and yields positive finite order integral representations of $W$. A quantitative perturbation estimate constructs functions with off line zeros preserving both these finite blocks and all the stated derivative signs. The reciprocal Xi clock is infinitely divisible and self decomposable, and the sine squared lemma gives a regularised arithmetic formula for its Thorin functional. Positivity of the full functional and the generalised gamma convolution property are reformulations of RH; for the latter it suffices to establish the transform identity for real $0<s<1$. The positive boundary measure on the critical line records only line zeros. A separate interval certificate proves a nonreal pole of the Pólya Mellin ratio, excluding the sufficient Bernstein Pick criterion of Konstantopoulos, Patie and Sarkar. A beta multiplier obstruction excludes every positive KPS parameter for scale mixtures of arcsine laws, including symmetric unimodal densities with all moments. The sufficient KPS route to RH considered here is therefore closed; a Hausdorff moment criterion characterises the remaining membership question at $θ=0$. The finite clock and matrix properties proved here do not imply RH. The complete monotonicity condition remains open.

math.GM

Bayes with No Shame: Admissibility Geometries of Predictive Inference

Modern predictive systems combine predictors, sequential monitors, prediction sets, and online strategies, each with a different certificate of optimality. We study four criterion-relative geometries: Blackwell risk dominance, anytime-valid admissibility, fixed-level marginal coverage with expected-length efficiency within a declared rank-indexed family, and choice-based approachability (CApp) boundary-feasibility. We embed the four procedure types in a common product space and prove witness-based pairwise non-nesting: for every ordered pair of criterion classes, an explicit predictive system is active in both relevant coordinates, belongs to one class, and fails the other. The result records non-nesting across different object spaces and partial orders; it does not assert practical incompatibility. We separate three measure-relative coherence notions. In conditionally i.i.d. models, posterior predictive means under a single prior are martingales under the prior predictive law. For a point null, anytime-valid admissibility within e-processes is equivalent to the nonnegative martingale property. Self-consistency under a predictor's own predictive law does not imply Blackwell admissibility, as shown by a Bernoulli log-loss counterexample. Coverage admissibility is certified by exchangeability ranks within the declared rank-indexed family, while CApp boundary-feasibility uses Cesaro steering. A constrained-Bayes design schema organizes the four paradigms without collapsing their distinct decision spaces, partial orders, or risk functionals. Admissibility is criterion-relative.

stat.ML

De Finetti + Sanov = Bayes: Exchangeable Prediction under Moment Constraints

We study exchangeable prediction when empirical-moment constraints define each active finite horizon N. The relevant law is the de Finetti mixture conditioned on E_N = {Phat_N in E_{eps_N}}. By permutation invariance, the target may be any fixed block of m coordinates within the active horizon, regardless of whether those coordinates are labeled past, held out, or future relative to any finite cut. Conditionally on the directing measure mu, the Gibbs-conditioning principle sends the law of such a block to the m-fold product of the I-projection P*_mu = argmin_{Q in E} D(Q || mu). On a finite alphabet we give an elementary master inequality for general polyhedral moment windows. After mixing over the constraint posterior Pi_{N,E}, and under weak convergence plus posterior-averaged component control, the finite-dimensional marginals converge to a consistent exchangeable law whose random directing measure is the I-projection P*_mu, with mu drawn from the weak limit Pi_E. Sequential prediction under this limiting law is therefore Bayesian prediction from a mixture of componentwise I-projections. The limiting behavior depends on reachability. For a reachable constraint, the projection is asymptotically the identity on the selected subfamily. Under additional regularity, an unreachable constraint makes the constraint posterior concentrate on the rate-minimizing subfamily, while the projections remain nontrivial. In our examples, at least one operation is asymptotically inactive, though both enter the finite-horizon construction. The master bound also reads as an equivalence of ensembles. We reserve "maximum entropy" for a uniform or flat baseline and use "minimum relative entropy" or "I-projection" in general.

math.ST

Bayesian Prediction under Moment Conditioning

Moment restrictions specify a class of laws, not a predictive model. We obtain one by conditioning an independent sample from a reference law on its empirical moments, and define prediction as the law of a fixed block selected from that conditioned ensemble. On a finite partition this law is an exact mixture over empirical types. Under exact feasibility and lattice regularity, the mixing law has a Gaussian limit on the feasible tangent space, governed by the reduced Hessian, and the selected block approaches independent sampling from the Kullback-Leibler projection. A separate finite-sample bound gives the same product limit for general real-valued restrictions without lattice assumptions. Refinement recovers the projection on the original sample space. Parameterizing the projected family produces a predictive product criterion with a local inverse-covariance expansion, connecting the construction to generalized method of moments.

math.ST

Martingale Posterior Predictive Coherence: Hausdorff Moment Hierarchy

For an exchangeable Bernoulli sequence with de Finetti mixing measure Pi, the k-step predictive probability P(X_{n+1}=...=X_{n+k}=0 | F_n) equals the posterior expectation E[(1-theta)^k | F_n]. By binomial expansion, this depends on all posterior moments up to order k. We show that the first moment alone is not sufficient to uniquely identify these quantities: for k >= 2, the mapping from posterior mean to k-step predictive is set-valued. The martingale posterior framework of Fong, Holmes, and Walker (which constrains only the first conditional moment of the terminal value) does not, in general, uniquely identify multi-step predictive distributions. Under any strictly proper scoring rule, the plug-in predictive is strictly dominated by the Bayes predictive whenever the posterior is non-degenerate. A closure theorem establishes that a martingale posterior determines all k-step predictives if and only if the conditional law of the terminal value is uniquely specified. Hill's A_{(n)} rule under the Jeffreys Beta(1/2,1/2) prior is a positive example. The discrepancy is O(Var(theta | F_n)) and vanishes as the posterior concentrates. These results clarify the structural requirements for predictive completeness under exchangeability.

math.ST

An Old Look at Empirical Bayes

Dennis Lindley once said that there is only one thing worse than a frequentist, and that is an empirical Bayesian. The quip has the air of caricature, but its technical content is serious: empirical Bayes uses the same data twice, conflates levels of a hierarchy, and produces posterior-shaped summaries whose uncertainty quantification differs from what a fully hierarchical model delivers. David Blei's 2026 IMS Medallion Lecture, "A Fresh Look at Empirical Bayes," revives the program under three new banners: empirical Bayes via probabilistic symmetries (rebranded "Bayesian empirical Bayes"), empirical Bayes with implicit likelihoods through simulation-based inference, and empirical Bayes for combining experimental and observational data through calibration studies. This is a continuation of Blei and Kucukelbir's earlier "population empirical Bayes" (PopEB, 2015). We argue, in the spirit of Lindley, I. J. Good, William DuMouchel, Thomas Louis, and our own recent work with Datta, that Blei's machinery targets inferential objects distinct from the posterior conditional on the realized data, and that the cost of maintaining the full hierarchical discipline has fallen low enough that the computational trade-off no longer favors the shortcut. The case study is the Tweedie formula. Efron's f-modeling empirical Bayes plugs an estimated score function into a posterior-mean identity, but a smoothed score need not arise from any prior. The horseshoe Tweedie formula does. We conclude by recommending that the impressive computational machinery of modern empirical Bayes (variational inference, neural amortization, simulation-based inference) be redeployed in service of properly hierarchical Bayes.

stat.ME

E-Values, Bayes Risk, Dual Role of Markov's Inequality

Two approaches to hypothesis testing, e-value testing and Bayes risk minimisation, both invoke Markov's inequality to control error probabilities. They differ in which distribution certifies the unit-moment condition: the null for Type I error, the alternative for Type II error. The likelihood ratio is not intrinsically an e-value; it acquires that status only relative to the experiment under which its expectation is certified. This note makes the resulting role-reversal symmetry explicit, traces its asymptotic sharpening through the information-theoretic arguments of Barron and Clarke (1994), and situates the duality within the typed evidence calculus of Polson, Sokolov, and Zantedeschi (2026).

math.ST

Kakeya Conjecture and Conditional Kolmogorov Complexity

This paper develops an information-theoretic framework for algorithmic complexity under regular identifiable fibering. The central question is: when a decoder is given information about the fiber label in a fibered geometric set, how much can the residual description length be reduced, and when does this reduction fail to bring dimension below the ambient rate? We formulate a directional compression principle, proposing that sets admitting regular, identifiable fiber decompositions should remain informationally incompressible at ambient dimension, unless the fiber structure is degenerate or adaptively chosen. The principle is phrased in the language of algorithmic dimension and the point-to-set principle of Lutz and Lutz, which translates pointwise Kolmogorov complexity into Hausdorff dimension. We prove an exact analytical result: under effectively bi-Lipschitz, identifiable, and computable fibering, the complexity of a point splits additively as the sum of fiber-label complexity and along-fiber residual complexity, up to logarithmic overhead, via the chain rule for Kolmogorov complexity. The Kakeya conjecture (asserting that sets containing a unit segment in every direction have Hausdorff dimension n) motivates the framework. The conjecture was recently resolved in R^3 by Wang and Zahl; it remains open in dimension n >= 4, precisely because adaptive fiber selection undermines the naive conditional split in the general case. We isolate this adaptive-fibering obstruction as the key difficulty and propose a formal research program connecting geometric measure theory, algorithmic complexity, and information-theoretic compression.

cs.IT

Prediction-Powered Inference with Inverse Probability Weighting

Prediction-powered inference (PPI) is a recent framework for valid statistical inference with partially labeled data, combining model-based predictions on a large unlabeled set with bias correction from a smaller labeled subset. Building on existing PPI results under covariate shift, we show that PPI rectification admits a direct design-based interpretation, and that informative labeling can be handled naturally by Horvitz--Thompson and Hájek-style corrections. This connection unites design-based survey sampling ideas with modern prediction-assisted inference, yielding estimators that remain valid when labeling probabilities vary across units. We consider the common setting where the inclusion probabilities are not known but estimated from a correctly specified model. In simulations, the performance of IPW-adjusted PPI with estimated propensities closely matches the known-probability case, retaining both nominal coverage and the variance-reduction benefits of PPI.

stat.ML

A New Look at Bayesian Testing

We identify the critical deviation scale governing Bayesian evidence accumulation in regular parametric testing. Under integrated Bayes risk with zero-one loss, the risk-optimal rejection boundary lies in a moderate deviation regime, with a square-root logarithmic inflation relative to the usual local asymptotic normal scale. Under Cramer regularity, local prior smoothness at the null, and symmetric loss, we derive the sharp threshold and show that its leading logarithmic term is universal across regular priors, while lower-order constants depend on the local prior density, Fisher information, and prior model odds. The result extends to one-parameter exponential families through local asymptotic normality and places Jeffreys' testing threshold, the Bayesian information criterion penalty, and Chernoff-Stein type error-exponent arguments within a common asymptotic moderate deviation framework.

math.ST

Bayes, E-values and Testing

E-values and E-processes (nonnegative supermartingales) provide anytime-valid evidence for sequential testing via Ville's inequality, yet their connection to Bayesian reasoning, representational structure, and computational feasibility are often conflated in the literature. We develop a typed framework that separates sequential evidence into three layers: (i) representation (Radon-Nikodym / likelihood-ratio geometry), (ii) validity (supermartingale certificates under optional stopping), and (iii) decision (boundary design and efficiency calibration). Our main results are: (a) under log-loss and Bayes-risk minimization, the likelihood ratio is the unique evidence representation within the coherent predictive subclass; (b) the likelihood-ratio stopping time satisfies E_1[tau_b] = (log b)/mu + O(sqrt(log b)) under Cramer conditions, while validity-only thresholds admit no such growth-rate guarantee; and (c) regret-optimal codes (e.g., NML/MDL) do not in general yield valid E-processes, while prequential codes do. Monte Carlo experiments confirm the theoretical predictions. The framework applies to online model validation, adaptive experimentation, conformal prediction, and sequential changepoint detection.

math.ST

Bayes Risk for Goodness of Fit Tests

We develop a unified framework for goodness-of-fit (GOF) testing through the lens of Bayes risk. Classical GOF procedures are commonly calibrated either at fixed significance level (CLT scale) or through exponential error exponents (LDP scale). We establish that Bayes-risk optimal calibration operates on the moderate-deviation (MDP) scale, producing canonical $\sqrt{\log n}$ inflation of rejection thresholds and polynomially decaying Type I error. Our main contributions are: (i) we formalise the Rubin--Sethuraman program for KS-type statistics as a risk-calibration theorem with explicit regularity conditions on priors and empirical-process functionals; (ii) we develop the precise connection between Bayes-risk expansions and Sanov information asymptotics, showing how $\log n$-order truncations arise naturally when risk, rather than pure exponents, is the evaluation criterion; (iii) we provide detailed applications to location testing under Laplace families, shape testing via Bayes factors, and connections to Fisher information geometry. The organizing principle throughout is that sample size enters Bayes-optimal GOF cutoffs through the MDP scale, unifying KS-based and Sanov-based perspectives under a single risk criterion.

math.ST

Entropy-Regularized Inference: A Predictive Approach

Predictive inference requires balancing statistical accuracy against informational complexity, yet the choice of complexity measure is usually imposed rather than derived. We treat econometric objects as predictive rules, mappings from information to reported predictive distributions, and impose three structural requirements on evaluation: locality, strict propriety, and coherence under aggregation (coarsening/refinement) of outcome categories. These axioms characterize (uniquely, up to affine transformations) the logarithmic score and induce Shannon mutual information (Kullback-Leibler divergence) as the corresponding measure of predictive complexity. The resulting entropy-regularized prediction problem admits Gibbs-form optimal rules, and we establish an essentially complete-class result for the admissible rules we study under joint risk-complexity dominance. Rational inattention emerges as the constrained dual, corresponding to frontier points with binding information capacity. The entropy penalty contributes additive curvature to the predictive criterion; in weakly identified settings, such as weak instruments in IV regression, where the unregularized objective is flat, this curvature stabilizes the predictive criterion. We derive a local quadratic (LAQ) expansion connecting entropy regularization to classical weak-identification diagnostics.

math.ST

Polynomial Log-Marginals and Tweedie's Formula : When Is Bayes Possible?

Motivated by Tweedie's formula for the Compound Decision problem, we examine the theoretical foundations of empirical Bayes estimators that directly model the marginal density $m(y)$. Our main result shows that polynomial log-marginals of degree $k \ge 3 $ cannot arise from any valid prior distribution in exponential family models, while quadratic forms correspond exactly to Gaussian priors. This provides theoretical justification for why certain empirical Bayes decision rules, while practically useful, do not correspond to any formal Bayes procedures. We also strengthen the diagnostic by showing that a marginal is a Gaussian convolution only if it extends to a bounded solution of the heat equation in a neighborhood of the smoothing parameter, beyond the convexity of $c(y)=\tfrac12 y^2+\log m(y)$.

math.ST

Conformal Prediction = Bayes?

Conformal prediction (CP) is widely presented as distribution-free predictive inference with finite-sample marginal coverage under exchangeability. We argue that CP is best understood as a rank-calibrated descendant of the Fisher-Dempster-Hill fiducial/direct-probability tradition rather than as Bayesian conditioning in disguise. We establish four separations from coherent countably additive predictive semantics. First, canonical conformal constructions violate conditional extensionality: prediction sets can depend on the marginal design P(X) even when P(Y|X) is fixed. Second, any finitely additive sequential extension preserving rank calibration is nonconglomerable, implying countable Dutch-book vulnerabilities. Third, rank-calibrated updates cannot be realized as regular conditionals of any countably additive exchangeable law on Y^infty. Fourth, formalizing both paradigms as families of one-step predictive kernels, conformal and Bayesian kernels coincide only on a Baire-meagre subset of the space of predictive laws. We further show that rank- and proxy-based reductions are generically Blackwell-deficient relative to full-data experiments, yielding positive Le Cam deficiency for suitable losses. Extending the analysis to prediction-powered inference (PPI) yields an analogous message: bias-corrected, proxy-rectified estimators can be valid as confidence devices while failing to define transportable belief states across stages, shifts, or adaptive selection. Together, the results sharpen a general limitation of wrappers: finite-sample calibration guarantees do not by themselves supply composable semantics for sequential updating or downstream decision-making.

math.ST

Bayesian ICA with super-Gaussian Source Priors

Independent Component Analysis (ICA) plays a central role in modern machine learning as a flexible framework for feature extraction. We introduce a horseshoe-type prior with a latent Polya-Gamma scale mixture representation, yielding scalable algorithms for both point estimation via expectation-maximization (EM) and full posterior inference via Markov chain Monte Carlo (MCMC). This hierarchical formulation unifies several previously disparate estimation strategies within a single Bayesian framework. We also establish the first theoretical guarantees for hierarchical Bayesian ICA, including posterior contraction and local asymptotic normality results for the unmixing matrix. Comprehensive simulation studies demonstrate that our methods perform competitively with widely used ICA tools. We further discuss implementation of conditional posteriors, envelope-based optimization, and possible extensions to flow-based architectures for nonlinear feature extraction and deep learning. Finally, we outline several promising directions for future work.

stat.ME

Quantile Importance Sampling

In Bayesian inference, the approximation of integrals of the form $ψ= \mathbb{E}_{F}{l(X)} = \int_χ l(\mathbf{x}) d F(\mathbf{x})$ is a fundamental challenge. Such integrals are crucial for evidence estimation, which is important for various purposes, including model selection and numerical analysis. The existing strategies for evidence estimation are classified into four categories: deterministic approximation, density estimation, importance sampling, and vertical representation (Llorente et al., 2020). In this paper, we show that the Riemann sum estimator due to Yakowitz (1978) can be used in the context of nested sampling (Skilling, 2006) to achieve a $O(n^{-4})$ rate of convergence, faster than the usual Ergodic Central Limit Theorem. We provide a brief overview of the literature on the Riemann sum estimators and the nested sampling algorithm and its connections to vertical likelihood Monte Carlo. We provide theoretical and numerical arguments to show how merging these two ideas may result in improved and more robust estimators for evidence estimation, especially in higher dimensional spaces. We also briefly discuss the idea of simulating the Lorenz curve that avoids the problem of intractable $Λ$ functions, essential for the vertical representation and nested sampling.

stat.CO