Searcharxiv⌕ Search

arXiv subjects

Laurence A. Jacobs

Publications and source records attributed to Laurence A. Jacobs.

3 recordsLinked to original sources

A prior-free blind detection of information leakage from model predictions

Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise. None operates on the artifact an auditor most often holds: the model's output. We ask what can be decided about leakage from predictions and outcomes alone. We give a decision-theoretic framework in which leakage diagnostics are functionals of the predicted-risk/outcome law, parameterized by a threshold-weighting linked to proper scoring rules and decision-curve analysis. We prove a sharp impossibility: a recalibrated leak matching an honest model's calibration and discrimination is indistinguishable from honest performance by \emph{any} function of the predictions, so the broad class is detectable only against an externally supplied ceiling on achievable discrimination. We then prove what leakage cannot hide: a near-deterministic subgroup -- the signature of a near-label leak -- produces a sustained unit-purity head that no legitimate predictor of a non-deterministic outcome can manufacture, yielding a prior-free test. These results organize leakage into a trichotomy -- miscalibrated, broad-calibrated, and deterministic -- each with a matched detector and failure mode. We validate on UK Biobank using time-windowed comorbidity leakage with known, graded severity, measuring a detection floor of $Δ\cstar \approx 0.007$ on this endpoint, below which residual leakage is undetectable from output and too small to alter conclusions. The numerical floor is cohort- and endpoint-specific; the structural lesson is general: output-only detection fails where residual leakage is indistinguishable from an honestly stronger predictor. The test returns a verdict on a prediction vector in under a second on commodity hardware.

cs.LG↗

Multicriticality and Scaling: Mellin Spectral Theory, and the Decoupling of Geometric and Spectral Exponents

We develop a spectral theory of scale-invariant operators on the multiplicative half-line $(\mathbb{R}_+, dx/x)$. A symmetric kernel $M(x, y)$ satisfying $M(kx, ky) = k^{-a}M(x, y)$ necessarily factorizes as $(xy)^{-a/2}F(x/y)$, where the shape function $F$ depends only on the ratio of its arguments. The Mellin transform diagonalizes such operators: the generalized eigenfunctions are $ψ_ω(x) = x^{-a/2+iω}$, and the eigenvalues are the Mellin multiplier $\tilde{F}(ω)$. This structure reveals a fundamental decoupling of two exponents. The geometric exponent $a$, carried by the power-law envelope $(xy)^{-a/2}$, governs the matrix scaling under dilation. The spectral exponent $b$, measured from the eigenvalue decay of the finite-dimensional truncation, is an effective quantity determined by the shape of $\tilde{F}(ω)$. For the explicit kernel $F(t) = c ρ^{|\ln t|}$, the Mellin multiplier is a Lorentzian of width $σ= -\ln ρ$, not a power law -- so $b$ is generically distinct from $a$. This decoupling provides a precise mathematical characterization of multicriticality: the equality $a = b$ corresponds to a simple critical fixed point of the Renormalization Group, while $a \neq b$ signals the presence of multiple independent scaling dimensions. We prove that the discrete self-similarity condition forces eigenvector collapse on the lattice, motivating the continuum formulation. Finite-size corrections from lattice sampling are quantified numerically.

math.GM↗

Temporal Matrix Scale Invariance and the Classification of Tipping Points

We introduce temporal matrix scale invariance (tMSI), a mathematical structure for the two-time correlation kernel of a multivariate observable. A kernel $C(t,t')$ satisfies tMSI of order $α$ if $C(kt, kt') = k^{-α}C(t,t')$ for all $k>0$; this condition holds near a tipping point, where the divergence of the coherence time produces temporal scale freedom. By a kernel factorization theorem, every tMSI kernel separates into a power-law envelope $(tt')^{-α/2}$ and a shape function $F(t/t')$ diagonalized by the Mellin transform. This reveals a decoupling of two independent exponents: the dynamical exponent $α$, carried by the envelope, and the spectral relaxation exponent $β$, determined by the eigenvalue decay of the finite-dimensional truncation. Their equality $α= β$ characterizes a simple critical point; their inequality $α\neq β$ is the signature of temporal multicriticality. We provide a classification of tipping points. The Landau quartic coefficient $a_4$ is given exactly by $a_4 = p^2 + q^2 - 2λpq - g^2_{ααβ}Γ(σ_α, σ_β)$, where $λ= 2\sqrt{σ_ασ_β}/(σ_α+σ_β) \in (0, 1]$, $g_{ααβ}$ is the three-point structure constant, and $Γ> 0$ is in explicit closed form. The transition is continuous for $a_4 > 0$, tricritical for $a_4 = 0$, and discontinuous for $a_4 < 0$. The simple critical point $α= β$ is maximally fragile: any nonzero operator mixing drives $a_4 < 0$, placing the synchronized state generically at the edge of catastrophe. The framework yields a matrix-valued early warning diagnostic, computable from a multivariate time series without knowledge of the underlying equations, that classifies an approaching tipping point as recoverable or catastrophic. Applications to epilepsy and acute myocardial infarction are discussed.

nlin.CD↗