SearcharxivSearch

arXiv subjects

Nick Polson

Publications and source records attributed to Nick Polson.

At least 19 recordsLinked to original sources

Perpetual Futures for Stocks: The SpaceX Pre-IPO Market

Robert Shiller proposed perpetual futures in 1993 to create derivative markets for assets that are illiquid or whose price cannot be observed directly, such as single family homes, human capital, and the consumer price index. The crypto markets later built the instrument for a different reason and with a different funding rule. We give a single no arbitrage result that nests both designs: the perpetual price is the present value of a benchmark flow discounted at the funding rate, so the funding rule chooses both the benchmark and the discount. We give a random time change representation in which the price is the expected spot at the first event of a clock whose intensity is the funding rate, use it to show that stochastic volatility moves the basis only through the carry, so a volatility risk premium and not volatility itself can break the peg, read price discovery as the convergence of a Doob martingale driven by a stochastic approximation, and give a segmented market equilibrium that makes the pre listing premium structural rather than behavioural. The June 2026 SpaceX pre-IPO market is the first large scale realization of the idea for an equity claim, and we find that the perpetual consensus forecast the secondary clearing price more accurately than the bookbuilt offer. We close with other equity applications.

q-fin.PR

Horseshoe Priors and MDP

Carvalho (2010) established two foundational theorems for the horseshoe prior: tight two-sided logarithmic bounds on the marginal density near the origin (Theorem~1.1), and a super-efficient rate of convergence of the Bayes predictive density to the true sampling density in sparse situations (Theorem~2). The ``Shrink Globally, Act Locally'' paper \citep{polson2010shrink} formalised necessary and sufficient conditions on the prior's behaviour at the origin for sparsity adaptation as $p \to \infty$. We show that these results are not merely descriptive properties of the horseshoe -- they are the finite-sample precursors to the asymptotic moderate deviation principle (MDP) of \citet{datta2026newlook}. The log-pole singularity $\piH(\theta) \asymp -\log\abs{\theta}$ is precisely the origin integrability boundary that selects the MDP threshold $\tcrit = \sqrt{\log(\pi n/2)}$; super-efficiency below the threshold and tail robustness above it together produce the ABOS Bayes risk $p_0 \log(p/p_0)/n$; and the Clarke--Barron information-theoretic asymptotics of Bayes methods provide the unifying framework in which all three results are faces of a single logarithmic budget principle.

math.ST

Bell's Inequality, Causal Bounds, and Quantum Bayesian Computation: A Unified Framework

Bell inequalities characterize the boundary of the local-realist correlation polytope -- the set of joint probability distributions achievable by classical hidden-variable models. Quantum mechanics exceeds this boundary through non-commutativity, reaching the Tsirelson bound $2\sqrt{2}$ for CHSH. We show that this polytope structure is not specific to quantum foundations: it appears identically in the causal inference literature, where the instrumental inequality, the Balke--Pearl linear programming bounds, and the Tian--Pearl probabilities of causation all arise as facets of the same marginal compatibility polytope. Fine's theorem -- that CHSH inequalities hold if and only if a joint distribution exists -- is precisely the pivot: the instrumental variable model in causal inference is structurally equivalent to the Bell local hidden-variable model, with the instrument playing the role of the measurement setting and the latent confounder playing the role of the hidden variable $\lambda$. We develop this correspondence in detail, extending it to algorithmic (Kolmogorov complexity) and entropic formulations of Bell inequalities, the NPA semidefinite programming hierarchy, and the MIP$^*$=RE undecidability result. We further show that the Born-rule / Bayes-rule duality underlying quantum Bayesian computation exploits the same non-commutativity that enables Bell violation, providing polynomial speedups for posterior inference. The framework yields a concrete dictionary between quantum information theory, causal econometrics, and Bayesian computation, and suggests new directions including NPA-based quantum causal inference algorithms and quantum architectures for function approximation.

quant-ph

Synthetic Priors

Bayesian inference in generalized linear models requires a prior on the coefficient vector $\beta$. Practitioners naturally reason about response probabilities at specific covariate values, not about abstract log-odds parameters. We develop synthetic priors: informative Bayesian priors for GLMs grounded in Good's device of imaginary observations -- the principle that every conjugate prior is equivalent to a likelihood on pseudo-data from the same exponential family. The conditional means prior of Bedrick (1996) elicits independent Beta priors on the conditional mean response at $p$ expert-chosen design points; the induced prior on $\beta$ is a product of binomial likelihoods at synthetic data points. Combined with P\'{o}lya-Gamma data augmentation \citep{polson2013}, the posterior admits an exact conjugate Gibbs sampler -- no tuning, no Metropolis step -- by treating the augmented dataset as a standard logistic regression. We show that ridge regression and catalytic priors \citep{huang2020} are instances of Good's device, and identify prediction-powered inference \citep{angelopoulos2023ppi} as a structural analogue in the frequentist setting -- all three mediate a variance-bias tradeoff through a single informativeness parameter. We illustrate the approach on two benchmark problems: the Challenger O-ring data \citep{dalal1989}, where the BCJ prior provides a more moderate posterior predictive at the 31{\deg}F launch temperature; and a Phase~II atopic dermatitis dose-finding trial ($n = 300$), where the synthetic prior narrows 95\% credible intervals by 3-6\% and raises decision probabilities by up to 2 percentage points relative to a flat prior.

stat.ME

Generative Bayesian Computation as a Scalable Alternative to Gaussian Process Surrogates

Gaussian process (GP) surrogates are the default tool for emulating expensive computer experiments, but cubic cost, stationarity assumptions, and Gaussian predictive distributions limit their reach. We propose Generative Bayesian Computation (GBC) via Implicit Quantile Networks (IQNs) as a surrogate framework that targets all three limitations. GBC learns the full conditional quantile function from input--output pairs; at test time, a single forward pass per quantile level produces draws from the predictive distribution. Across fourteen benchmarks we compare GBC to four GP-based methods. GBC improves CRPS by 11--26\% on piecewise jump-process benchmarks, by 14\% on a ten-dimensional Friedman function, and scales linearly to 90,000 training points where dense-covariance GPs are infeasible. A boundary-augmented variant matches or outperforms Modular Jump GPs on two-dimensional jump datasets (up to 46\% CRPS improvement). In active learning, a randomized-prior IQN ensemble achieves nearly three times lower RMSE than deep GP active learning on Rocket LGBB. Overall, GBC records a favorable point estimate in 12 of 14 comparisons. GPs retain an edge on smooth surfaces where their smoothness prior provides effective regularization.

cs.LG

Photons = Tokens: The Physics of AI and the Economics of Knowledge

Debates about artificial intelligence capabilities and risks are often conducted without quantitative grounding. This paper applies the methodology of MacKay (2009) -- who reframed energy policy as arithmetic -- to the economy of AI computation. We define the token, the elementary unit of large language model input and output, as a physical quantity with measurable thermodynamic cost. Using Landauer's principle, Shannon's channel capacity, and current infrastructure data, we construct a supply-and-demand balance sheet for global token production. We then derive a finite question budget: the number of meaningful queries humanity can direct at AI systems under physical, information-theoretic, and economic constraints. We apply Coase's theory of the firm and the durable-goods monopoly problem to the AI value chain -- from photon to atom to chip to power to token to question -- to identify where economic value concentrates and where regulatory intervention is warranted. We argue that the expansion of the token budget does not resolve a deeper constraint: under structural uncertainty, the decisive variable is not how many questions can be answered but which questions are worth asking -- a problem of agency and direction that computation alone cannot solve. We connect limits of measurement in the token economy to a structural parallel between Goodhart's law and the Heisenberg uncertainty principle, and to Arrow's impossibility result for efficient information pricing. The framework yields order-of-magnitude estimates that discipline policy discussion: at current efficiency, the projected 2028 US AI energy allocation of 326~TWh could support roughly $6.5 \times 10^{17}$ tokens per year, or 225,000 tokens per person per day -- more than three orders of magnitude above estimated mid-2024 utilization.

physics.soc-ph

Fast Compute for ML Optimization

We study optimization for losses that admit a variance-mean scale-mixture representation. Under this representation, each EM iteration is a weighted least squares update in which latent variables determine observation and parameter weights; these play roles analogous to Adam's second-moment scaling and AdamW's weight decay, but are derived from the model. The resulting Scale Mixture EM (SM-EM) algorithm removes user-specified learning-rate and momentum schedules. On synthetic ill-conditioned logistic regression benchmarks with $p \in \{20, \ldots, 500\}$, SM-EM with Nesterov acceleration attains up to $13\times$ lower final loss than Adam tuned by learning-rate grid search. For a 40-point regularization path, sharing sufficient statistics across penalty values yields a $10\times$ runtime reduction relative to the same tuned-Adam protocol. For the base (non-accelerated) algorithm, EM monotonicity guarantees nonincreasing objective values; adding Nesterov extrapolation trades this guarantee for faster empirical convergence.

stat.CO

Some Bayesian Perspectives on Clinical Trials

We examine three landmark clinical trials -- ECMO, CALGB~49907, and I-SPY~2 -- through a unified Bayesian framework connecting prior specification, sequential adaptation, and decision-theoretic optimisation. For ECMO, the posterior probability of treatment superiority is robust across the range of priors examined. For CALGB, predictive probability monitoring stopped enrolment at 633 instead of 1800 patients. For I-SPY~2, adaptive enrichment graduated nine of 23 arms to Phase~III. These case studies motivate a methodological contribution: exact backward induction for two-arm binary trials, where Beta-Binomial conjugacy yields closed-form transitions on the integer lattice of success counts with no quadrature. A P\'olya-Gamma augmentation bridges this to covariate-adjusted logistic regression. Simulation reveals a fundamental tension: the optimal Bayesian design reduces expected sample sizes to 14--26 per arm (versus 42--100 for alternatives) but with substantially lower power. A calibrated variant embedding the declaration threshold in the terminal utility improves power while maintaining sample-size savings; varying the per-stage cost traces a power frontier for selecting the preferred operating point, with suitability highest in patient-sparing contexts such as rare diseases and paediatrics. The P\'olya-Gamma Laplace approximation is validated against exact calculations (mean absolute error below 0.01). We discuss implications for the 2026 FDA draft guidance on Bayesian methodology.

stat.ME

Horseshoe Mixtures-of-Experts (HS-MoE)

Horseshoe mixtures-of-experts (HS-MoE) models provide a Bayesian framework for sparse expert selection in mixture-of-experts architectures. We combine the horseshoe prior's adaptive global-local shrinkage with input-dependent gating, yielding data-adaptive sparsity in expert usage. Our primary methodological contribution is a particle learning algorithm for sequential inference, in which the filter is propagated forward in time while tracking only sufficient statistics. We also discuss how HS-MoE relates to modern mixture-of-experts layers in large language models, which are deployed under extreme sparsity constraints (e.g., activating a small number of experts per token out of a large pool).

stat.ML

Generative Bayesian Hyperparameter Tuning

\noindent Hyper-parameter selection is a central practical problem in modern machine learning, governing regularization strength, model capacity, and robustness choices. Cross-validation is often computationally prohibitive at scale, while fully Bayesian hyper-parameter learning can be difficult due to the cost of posterior sampling. We develop a generative perspective on hyper-parameter tuning that combines two ideas: (i) optimization-based approximations to Bayesian posteriors via randomized, weighted objectives (weighted Bayesian bootstrap), and (ii) amortization of repeated optimization across many hyper-parameter settings by learning a transport map from hyper-parameters (including random weights) to the corresponding optimizer. This yields a ``generator look-up table'' for estimators, enabling rapid evaluation over grids or continuous ranges of hyper-parameters and supporting both predictive tuning objectives and approximate Bayesian uncertainty quantification. We connect this viewpoint to weighted $M$-estimation, envelope/auxiliary-variable representations that reduce non-quadratic losses to weighted least squares, and recent generative samplers for weighted $M$-estimators.

stat.ML

Bayesian Global-Local Regularization

We propose a unified framework for global-local regularization that bridges the gap between classical techniques -- such as ridge regression and the nonnegative garotte -- and modern Bayesian hierarchical modeling. By estimating local regularization strengths via marginal likelihood under order constraints, our approach generalizes Stein's positive-part estimator and provides a principled mechanism for adaptive shrinkage in high-dimensional settings. We establish that this isotonic empirical Bayes estimator achieves near-minimax risk (up to logarithmic factors) over sparse ordered model classes, constituting a significant advance in high-dimensional statistical inference. Applications to orthogonal polynomial regression demonstrate the methodology's flexibility, while our theoretical results clarify the connections between empirical Bayes, shape-constrained estimation, and degrees-of-freedom adjustments.

stat.ME

$MC^2$ Mixed Integer and Linear Programming

In this paper, we design $MC^2$ algorithms for Mixed Integer and Linear Programming. By expressing a constrained optimisation as one of simulation from a Boltzmann distribution, we reformulate integer and linear programming as Monte Carlo optimisation problems. The key insight is that solving these optimisation problems requires the ability to simulate from truncated distributions, namely multivariate exponentials and Gaussians. Efficient simulation can be achieved using the algorithms of Kent and Davis. We demonstrate our methodology on portfolio optimisation and the classical farmer problem from stochastic programming. Finally, we conclude with directions for future research.

stat.CO

Generative Quantile Bayesian Prediction

Prediction is a central task of machine learning. Our goal is to solve large scale prediction problems using Generative Quantile Bayesian Prediction (GQBP).By directly learning predictive quantiles rather than densities we achieve a number of theoretical and practical advantages. We contrast our approach with state-of-the-art methods including conformal prediction, fiducial prediction and marginal likelihood. Our distinguishing feature of our method is the use of generative methods for predictive quantile maps. We illustrate our methodology for normal-normal learning and causal inference. Finally, we conclude with directions for future research.

stat.ME

Bayesian Double Descent

Double descent is a phenomenon of over-parameterized statistical models such as deep neural networks which have a re-descending property in their risk function. As the complexity of the model increases, risk exhibits a U-shaped region due to the traditional bias-variance trade-off, then as the number of parameters equals the number of observations and the model becomes one of interpolation where the risk can be unbounded and finally, in the over-parameterized region, it re-descends -- the double descent effect. Our goal is to show that this has a natural Bayesian interpretation. We also show that this is not in conflict with the traditional Occam's razor -- simpler models are preferred to complex ones, all else being equal. Our theoretical foundations use Bayesian model selection, the Dickey-Savage density ratio, and connect generalized ridge regression and global-local shrinkage methods with double descent. We illustrate our approach for high dimensional neural networks and provide detailed treatments of infinite Gaussian means models and non-parametric regression. Finally, we conclude with directions for future research.

stat.ML

Generative Modeling: A Review

We organize the generative-modeling literature around three classes of generators, corresponding to three distinct inferential tasks: estimating counterfactual outcome distributions in causal inference, recovering posteriors from simulated parameter--outcome pairs, and forming predictive outcome distributions. The unifying representation relies on the noise outsourcing theorem of Kallenberg, which expresses a conditional distribution as a deterministic function of its inputs and an independent noise variable. Within this organization we develop generative Bayesian computation, a method in the parameter--outcome class: a quantile neural network, trained on simulated pairs under the pinball loss, that targets the posterior of the parameter directly, without invertible architectures or density evaluation, and that serves equally as a predictive generator once the roles of parameter and outcome are exchanged. We illustrate the framework on an agent-based Ebola transmission application, where generative Bayesian computation recovers accurate posteriors at substantially lower cost than rejection-based simulation inference, while avoiding the density-evaluation and invertibility constraints of competing generators.

stat.CO

Kramnik vs Nakamura: A Chess Scandal

We provide a statistical analysis of the recent controversy between Vladimir Kramnik (ex chess world champion) and Hikaru Nakamura. Hikaru Nakamura is a chess prodigy and a five-time United States chess champion. Kramnik called into question Nakamura's 45.5 out of 46 win streak in an online blitz contest at chess.com. We assess the weight of evidence using a priori assessment of Viswanathan Anand and the streak evidence. Based on this evidence, we show that Nakamura has a 99.6 percent chance of not cheating. We study the statistical fallacies prevalent in both their analyses. On the one hand Kramnik bases his argument on the probability of such a streak is very small. This falls precisely into the Prosecutor's Fallacy. On the other hand, Nakamura tries to refute the argument using a cherry-picking argument. This violates the likelihood principle. We conclude with a discussion of the relevant statistical literature on the topic of fraud detection and the analysis of streaks in sports data.

stat.AP

Generative Bayesian Computation for Maximum Expected Utility

Generative Bayesian Computation (GBC) methods are developed to provide an efficient computational solution for maximum expected utility (MEU). We propose a density-free generative method based on quantiles that naturally calculates expected utility as a marginal of quantiles. Our approach uses a deep quantile neural estimator to directly estimate distributional utilities. Generative methods assume only the ability to simulate from the model and parameters and as such are likelihood-free. A large training dataset is generated from parameters and output together with a base distribution. Our method a number of computational advantages primarily being density-free with an efficient estimator of expected utility. A link with the dual theory of expected utility and risk taking is also discussed. To illustrate our methodology, we solve an optimal portfolio allocation problem with Bayesian learning and a power utility (a.k.a. fractional Kelly criterion). Finally, we conclude with directions for future research.

stat.CO

Counting $N$ Queens

Gauss proposed the problem of how to enumerate the number of solutions for placing $N$ queens on an $N\times N$ chess board, so no two queens attack each other. The N-queen problem is a classic problem in combinatorics. We describe a variety of Monte Carlo (MC) methods for counting the number of solutions. In particular, we propose a quantile re-ordering based on the Lorenz curve of a sum that is related to counting the number of solutions. We show his approach leads to an efficient polynomial-time solution. Other MC methods include vertical likelihood Monte Carlo, importance sampling, slice sampling, simulated annealing, energy-level sampling, and nested-sampling. Sampling binary matrices that identify the locations of the queens on the board can be done with a Swendsen-Wang style algorithm. Our Monte Carlo approach counts the number of solutions in polynomial time.

stat.CO