SearcharxivSearch

arXiv subjects

Daniel Zantedeschi

Publications and source records attributed to Daniel Zantedeschi.

16 recordsLinked to original sources

An Old Look at Empirical Bayes

Dennis Lindley once said that there is only one thing worse than a frequentist, and that is an empirical Bayesian. The quip has the air of caricature, but its technical content is serious: empirical Bayes uses the same data twice, conflates levels of a hierarchy, and produces posterior-shaped summaries whose uncertainty quantification differs from what a fully hierarchical model delivers. David Blei's 2026 IMS Medallion Lecture, "A Fresh Look at Empirical Bayes," revives the program under three new banners: empirical Bayes via probabilistic symmetries (rebranded "Bayesian empirical Bayes"), empirical Bayes with implicit likelihoods through simulation-based inference, and empirical Bayes for combining experimental and observational data through calibration studies. This is a continuation of Blei and Kucukelbir's earlier "population empirical Bayes" (PopEB, 2015). We argue, in the spirit of Lindley, I. J. Good, William DuMouchel, Thomas Louis, and our own recent work with Datta, that Blei's machinery targets inferential objects distinct from the posterior conditional on the realized data, and that the cost of maintaining the full hierarchical discipline has fallen low enough that the computational trade-off no longer favors the shortcut. The case study is the Tweedie formula. Efron's f-modeling empirical Bayes plugs an estimated score function into a posterior-mean identity, but a smoothed score need not arise from any prior. The horseshoe Tweedie formula does. We conclude by recommending that the impressive computational machinery of modern empirical Bayes (variational inference, neural amortization, simulation-based inference) be redeployed in service of properly hierarchical Bayes.

stat.ME

E-Values, Bayes Risk, Dual Role of Markov's Inequality

Two approaches to hypothesis testing, e-value testing and Bayes risk minimisation, both invoke Markov's inequality to control error probabilities. They differ in which distribution certifies the unit-moment condition: the null for Type I error, the alternative for Type II error. The likelihood ratio is not intrinsically an e-value; it acquires that status only relative to the experiment under which its expectation is certified. This note makes the resulting role-reversal symmetry explicit, traces its asymptotic sharpening through the information-theoretic arguments of Barron and Clarke (1994), and situates the duality within the typed evidence calculus of Polson, Sokolov, and Zantedeschi (2026).

math.ST

Horseshoe Priors and MDP

Carvalho (2010) established two foundational theorems for the horseshoe prior: tight two-sided logarithmic bounds on the marginal density near the origin (Theorem~1.1), and a super-efficient rate of convergence of the Bayes predictive density to the true sampling density in sparse situations (Theorem~2). The ``Shrink Globally, Act Locally'' paper \citep{polson2010shrink} formalised necessary and sufficient conditions on the prior's behaviour at the origin for sparsity adaptation as $p \to \infty$. We show that these results are not merely descriptive properties of the horseshoe -- they are the finite-sample precursors to the asymptotic moderate deviation principle (MDP) of \citet{datta2026newlook}. The log-pole singularity $\piH(\theta) \asymp -\log\abs{\theta}$ is precisely the origin integrability boundary that selects the MDP threshold $\tcrit = \sqrt{\log(\pi n/2)}$; super-efficiency below the threshold and tail robustness above it together produce the ABOS Bayes risk $p_0 \log(p/p_0)/n$; and the Clarke--Barron information-theoretic asymptotics of Bayes methods provide the unifying framework in which all three results are faces of a single logarithmic budget principle.

math.ST

Bell's Inequality, Causal Bounds, and Quantum Bayesian Computation: A Unified Framework

Bell inequalities characterize the boundary of the local-realist correlation polytope -- the set of joint probability distributions achievable by classical hidden-variable models. Quantum mechanics exceeds this boundary through non-commutativity, reaching the Tsirelson bound $2\sqrt{2}$ for CHSH. We show that this polytope structure is not specific to quantum foundations: it appears identically in the causal inference literature, where the instrumental inequality, the Balke--Pearl linear programming bounds, and the Tian--Pearl probabilities of causation all arise as facets of the same marginal compatibility polytope. Fine's theorem -- that CHSH inequalities hold if and only if a joint distribution exists -- is precisely the pivot: the instrumental variable model in causal inference is structurally equivalent to the Bell local hidden-variable model, with the instrument playing the role of the measurement setting and the latent confounder playing the role of the hidden variable $\lambda$. We develop this correspondence in detail, extending it to algorithmic (Kolmogorov complexity) and entropic formulations of Bell inequalities, the NPA semidefinite programming hierarchy, and the MIP$^*$=RE undecidability result. We further show that the Born-rule / Bayes-rule duality underlying quantum Bayesian computation exploits the same non-commutativity that enables Bell violation, providing polynomial speedups for posterior inference. The framework yields a concrete dictionary between quantum information theory, causal econometrics, and Bayesian computation, and suggests new directions including NPA-based quantum causal inference algorithms and quantum architectures for function approximation.

quant-ph

Modal Exchangeability: Centered Symmetry and the Credal Architecture of Kripke Frames

We ask what happens when the index set carries modal structure, with possibilities organized into a Kripke frame. We define modal exchangeability as invariance under accessibility-preserving automorphisms that fix a designated base world, and derive a representation theorem for countable frames. The orbit decomposition of the centered symmetry group governs the within-orbit structure: worlds in the same orbit are conditionally identically distributed, and on orbits satisfying a richness condition and countable infinitude they are conditionally i.i.d. given a rigid orbit-specific directing measure. Point-homogeneous S5 frames yield a single de Finetti parameter; S4 frames may admit multiple orbits, with the richer orbits carrying rigid directing measures and the remainder carrying only weaker invariant structure. Two applications follow. First, the orbit decomposition determines whether learning pools globally or remains orbit-local. Second, it supplies a mechanism for structural credal fine-graining indexed to orbit regions, distinct from hyperintensionality in the strict sense of distinguishing coextensive propositions.

math.LO

Kakeya Conjecture and Conditional Kolmogorov Complexity

This paper develops an information-theoretic framework for algorithmic complexity under regular identifiable fibering. The central question is: when a decoder is given information about the fiber label in a fibered geometric set, how much can the residual description length be reduced, and when does this reduction fail to bring dimension below the ambient rate? We formulate a directional compression principle, proposing that sets admitting regular, identifiable fiber decompositions should remain informationally incompressible at ambient dimension, unless the fiber structure is degenerate or adaptively chosen. The principle is phrased in the language of algorithmic dimension and the point-to-set principle of Lutz and Lutz, which translates pointwise Kolmogorov complexity into Hausdorff dimension. We prove an exact analytical result: under effectively bi-Lipschitz, identifiable, and computable fibering, the complexity of a point splits additively as the sum of fiber-label complexity and along-fiber residual complexity, up to logarithmic overhead, via the chain rule for Kolmogorov complexity. The Kakeya conjecture (asserting that sets containing a unit segment in every direction have Hausdorff dimension n) motivates the framework. The conjecture was recently resolved in R^3 by Wang and Zahl; it remains open in dimension n >= 4, precisely because adaptive fiber selection undermines the naive conditional split in the general case. We isolate this adaptive-fibering obstruction as the key difficulty and propose a formal research program connecting geometric measure theory, algorithmic complexity, and information-theoretic compression.

cs.IT

Bayes with No Shame: Admissibility Geometries of Predictive Inference

Modern predictive systems combine predictors, sequential monitors, prediction sets, and online strategies, each with a different certificate of optimality. We study four criterion-relative geometries: Blackwell risk dominance, anytime-valid admissibility, fixed-level marginal coverage with expected-length efficiency within a declared rank-indexed family, and choice-based approachability (CApp) boundary-feasibility. We embed the four procedure types in a common product space and prove witness-based pairwise non-nesting: for every ordered pair of criterion classes, an explicit predictive system is active in both relevant coordinates, belongs to one class, and fails the other. The result records non-nesting across different object spaces and partial orders; it does not assert practical incompatibility. We separate three measure-relative coherence notions. In conditionally i.i.d. models, posterior predictive means under a single prior are martingales under the prior predictive law. For a point null, anytime-valid admissibility within e-processes is equivalent to the nonnegative martingale property. Self-consistency under a predictor's own predictive law does not imply Blackwell admissibility, as shown by a Bernoulli log-loss counterexample. Coverage admissibility is certified by exchangeability ranks within the declared rank-indexed family, while CApp boundary-feasibility uses Cesaro steering. A constrained-Bayes design schema organizes the four paradigms without collapsing their distinct decision spaces, partial orders, or risk functionals. Admissibility is criterion-relative.

stat.ML

Mini-Batch Covariance, Diffusion Limits, and Oracle Complexity in Stochastic Gradient Descent: A Sampling-Design Perspective

Stochastic gradient descent (SGD) is central to simulation optimization, stochastic programming, and online M-estimation, where sampling effort is a decision variable. We study the mini-batch gradient noise as a sampling-design object. Under exchangeable fresh-sampling mini-batches, the conditional covariance given the de Finetti directing measure mu is b^{-1} G_mu(theta), and under identifiability the projected population object is b^{-1} G*(theta) -- projected Fisher information for correctly specified likelihoods, the sandwich partner of the Hessian otherwise. This identification fixes the noise matrix entering the diffusion analysis of constant-step SGD: the raw iterate path has a deterministic fluid limit, and the sqrt(b/eta)-scaled fluctuations satisfy a functional CLT with noise covariance G*; near a nondegenerate optimum the limit is Ornstein-Uhlenbeck, and its Lyapunov covariance scaled by eta/b matches the linearized discrete recursion at leading order. Under a curvature-noise compatibility condition mu_F > 0, we prove 1/N mean-square upper bounds and an i.i.d. parametric Fisher van Trees lower bound of the same rate order, with oracle-complexity guarantees depending on an effective dimension d_eff and condition number kappa_F. Numerical experiments verify the identification and confirm the Lyapunov predictions in direct SGD.

stat.ML

Martingale Posterior Predictive Coherence: Hausdorff Moment Hierarchy

For an exchangeable Bernoulli sequence with de Finetti mixing measure Pi, the k-step predictive probability P(X_{n+1}=...=X_{n+k}=0 | F_n) equals the posterior expectation E[(1-theta)^k | F_n]. By binomial expansion, this depends on all posterior moments up to order k. We show that the first moment alone is not sufficient to uniquely identify these quantities: for k >= 2, the mapping from posterior mean to k-step predictive is set-valued. The martingale posterior framework of Fong, Holmes, and Walker (which constrains only the first conditional moment of the terminal value) does not, in general, uniquely identify multi-step predictive distributions. Under any strictly proper scoring rule, the plug-in predictive is strictly dominated by the Bayes predictive whenever the posterior is non-degenerate. A closure theorem establishes that a martingale posterior determines all k-step predictives if and only if the conditional law of the terminal value is uniquely specified. Hill's A_{(n)} rule under the Jeffreys Beta(1/2,1/2) prior is a positive example. The discrepancy is O(Var(theta | F_n)) and vanishes as the posterior concentrates. These results clarify the structural requirements for predictive completeness under exchangeability.

math.ST

Bayes Risk for Goodness of Fit Tests

We develop a unified framework for goodness-of-fit (GOF) testing through the lens of Bayes risk. Classical GOF procedures are commonly calibrated either at fixed significance level (CLT scale) or through exponential error exponents (LDP scale). We establish that Bayes-risk optimal calibration operates on the moderate-deviation (MDP) scale, producing canonical $\sqrt{\log n}$ inflation of rejection thresholds and polynomially decaying Type I error. Our main contributions are: (i) we formalise the Rubin--Sethuraman program for KS-type statistics as a risk-calibration theorem with explicit regularity conditions on priors and empirical-process functionals; (ii) we develop the precise connection between Bayes-risk expansions and Sanov information asymptotics, showing how $\log n$-order truncations arise naturally when risk, rather than pure exponents, is the evaluation criterion; (iii) we provide detailed applications to location testing under Laplace families, shape testing via Bayes factors, and connections to Fisher information geometry. The organizing principle throughout is that sample size enters Bayes-optimal GOF cutoffs through the MDP scale, unifying KS-based and Sanov-based perspectives under a single risk criterion.

math.ST

A New Look at Bayesian Testing

We identify the critical deviation scale governing Bayesian evidence accumulation in regular parametric testing. Under integrated Bayes risk with zero-one loss, the risk-optimal rejection boundary lies in a moderate deviation regime, with a square-root logarithmic inflation relative to the usual local asymptotic normal scale. Under Cramer regularity, local prior smoothness at the null, and symmetric loss, we derive the sharp threshold and show that its leading logarithmic term is universal across regular priors, while lower-order constants depend on the local prior density, Fisher information, and prior model odds. The result extends to one-parameter exponential families through local asymptotic normality and places Jeffreys' testing threshold, the Bayesian information criterion penalty, and Chernoff-Stein type error-exponent arguments within a common asymptotic moderate deviation framework.

math.ST

Bayes, E-values and Testing

E-values and E-processes (nonnegative supermartingales) provide anytime-valid evidence for sequential testing via Ville's inequality, yet their connection to Bayesian reasoning, representational structure, and computational feasibility are often conflated in the literature. We develop a typed framework that separates sequential evidence into three layers: (i) representation (Radon-Nikodym / likelihood-ratio geometry), (ii) validity (supermartingale certificates under optional stopping), and (iii) decision (boundary design and efficiency calibration). Our main results are: (a) under log-loss and Bayes-risk minimization, the likelihood ratio is the unique evidence representation within the coherent predictive subclass; (b) the likelihood-ratio stopping time satisfies E_1[tau_b] = (log b)/mu + O(sqrt(log b)) under Cramer conditions, while validity-only thresholds admit no such growth-rate guarantee; and (c) regret-optimal codes (e.g., NML/MDL) do not in general yield valid E-processes, while prequential codes do. Monte Carlo experiments confirm the theoretical predictions. The framework applies to online model validation, adaptive experimentation, conformal prediction, and sequential changepoint detection.

math.ST

Conformal Prediction = Bayes?

Conformal prediction (CP) is widely presented as distribution-free predictive inference with finite-sample marginal coverage under exchangeability. We argue that CP is best understood as a rank-calibrated descendant of the Fisher-Dempster-Hill fiducial/direct-probability tradition rather than as Bayesian conditioning in disguise. We establish four separations from coherent countably additive predictive semantics. First, canonical conformal constructions violate conditional extensionality: prediction sets can depend on the marginal design P(X) even when P(Y|X) is fixed. Second, any finitely additive sequential extension preserving rank calibration is nonconglomerable, implying countable Dutch-book vulnerabilities. Third, rank-calibrated updates cannot be realized as regular conditionals of any countably additive exchangeable law on Y^infty. Fourth, formalizing both paradigms as families of one-step predictive kernels, conformal and Bayesian kernels coincide only on a Baire-meagre subset of the space of predictive laws. We further show that rank- and proxy-based reductions are generically Blackwell-deficient relative to full-data experiments, yielding positive Le Cam deficiency for suitable losses. Extending the analysis to prediction-powered inference (PPI) yields an analogous message: bias-corrected, proxy-rectified estimators can be valid as confidence devices while failing to define transportable belief states across stages, shifts, or adaptive selection. Together, the results sharpen a general limitation of wrappers: finite-sample calibration guarantees do not by themselves supply composable semantics for sequential updating or downstream decision-making.

math.ST

Entropy-Regularized Inference: A Predictive Approach

Predictive inference requires balancing statistical accuracy against informational complexity, yet the choice of complexity measure is usually imposed rather than derived. We treat econometric objects as predictive rules, mappings from information to reported predictive distributions, and impose three structural requirements on evaluation: locality, strict propriety, and coherence under aggregation (coarsening/refinement) of outcome categories. These axioms characterize (uniquely, up to affine transformations) the logarithmic score and induce Shannon mutual information (Kullback-Leibler divergence) as the corresponding measure of predictive complexity. The resulting entropy-regularized prediction problem admits Gibbs-form optimal rules, and we establish an essentially complete-class result for the admissible rules we study under joint risk-complexity dominance. Rational inattention emerges as the constrained dual, corresponding to frontier points with binding information capacity. The entropy penalty contributes additive curvature to the predictive criterion; in weakly identified settings, such as weak instruments in IV regression, where the unregularized objective is flat, this curvature stabilizes the predictive criterion. We derive a local quadratic (LAQ) expansion connecting entropy regularization to classical weak-identification diagnostics.

math.ST

Bayesian Prediction under Moment Conditioning

Moment restrictions specify a class of laws, not a predictive model. We obtain one by conditioning an independent sample from a reference law on its empirical moments, and define prediction as the law of a fixed block selected from that conditioned ensemble. On a finite partition this law is an exact mixture over empirical types. Under exact feasibility and lattice regularity, the mixing law has a Gaussian limit on the feasible tangent space, governed by the reduced Hessian, and the selected block approaches independent sampling from the Kullback-Leibler projection. A separate finite-sample bound gives the same product limit for general real-valued restrictions without lattice assumptions. Refinement recovers the projection on the original sample space. Parameterizing the projected family produces a predictive product criterion with a local inverse-covariance expansion, connecting the construction to generalized method of moments.

math.ST

De Finetti + Sanov = Bayes: Exchangeable Prediction under Moment Constraints

We study exchangeable prediction when empirical-moment constraints define each active finite horizon N. The relevant law is the de Finetti mixture conditioned on E_N = {Phat_N in E_{eps_N}}. By permutation invariance, the target may be any fixed block of m coordinates within the active horizon, regardless of whether those coordinates are labeled past, held out, or future relative to any finite cut. Conditionally on the directing measure mu, the Gibbs-conditioning principle sends the law of such a block to the m-fold product of the I-projection P*_mu = argmin_{Q in E} D(Q || mu). On a finite alphabet we give an elementary master inequality for general polyhedral moment windows. After mixing over the constraint posterior Pi_{N,E}, and under weak convergence plus posterior-averaged component control, the finite-dimensional marginals converge to a consistent exchangeable law whose random directing measure is the I-projection P*_mu, with mu drawn from the weak limit Pi_E. Sequential prediction under this limiting law is therefore Bayesian prediction from a mixture of componentwise I-projections. The limiting behavior depends on reachability. For a reachable constraint, the projection is asymptotically the identity on the selected subfamily. Under additional regularity, an unreachable constraint makes the constraint posterior concentrate on the rate-minimizing subfamily, while the projections remain nontrivial. In our examples, at least one operation is asymptotically inactive, though both enter the finite-horizon construction. The master bound also reads as an equivalence of ensembles. We reserve "maximum entropy" for a uniform or flat baseline and use "minimum relative entropy" or "I-projection" in general.

math.ST