SearcharxivSearch

arXiv subjects

Philip Kennerberg

Publications and source records attributed to Philip Kennerberg.

13 recordsLinked to original sources

Special Dirichlet Processes: Structure, Uniqueness and Stability

We introduce the class of \emph{Special--Dirichlet processes}, consisting of c\`adl\`ag adapted processes admitting a decomposition \[ X=M+\Gamma, \] where \(M\) is a local martingale and \(\Gamma\) is an adapted c\`adl\`ag process with vanishing continuous quadratic variation whose jumps are predictable and \(\mathcal F_{s-}\)-measurable. This class arises naturally from transformations of special semimartingales. Classical results imply that sufficiently regular functions of special semimartingales belong to the broad class of Dirichlet processes. We show that such transformed processes possess substantially more structure: they admit a canonical decomposition in which the predictable jump component is explicitly separated from the martingale component. This yields a refinement of the traditional classification, which previously identified these processes only as Dirichlet processes. We establish uniqueness of the decomposition and prove that the class is stable under a large family of nonsmooth transformations, including primitives of locally bounded functions with at most countably many discontinuities. An explicit It\^o-type decomposition is obtained in terms of the martingale jump measure and its compensator. Finally, we investigate stability properties of the canonical decomposition. Under convergence in quadratic variation and Skorokhod \(J_1\)-convergence, we prove stability of both the martingale and singular components after transformation. The proof relies on a threshold isolation principle for jump structures, allowing large jumps to be separated from small-jump contributions and yielding convergence of the transformed decompositions.

math.PR

Stability of Compensated Jump Integrals under Quadratic Variation Convergence

We study the stability of compensated jump integrals under convergence of quadratic variation alone. Let \(X\) and \(\{X^n\}_{n\ge1}\) be c\`adl\`ag processes with jump measures \(\mu,\mu_n\) and predictable compensators \(\nu,\nu_n\). Under the assumption \[ [X^n-X]_t \to 0 \qquad\text{in probability}, \] we establish ucp convergence of compensated jump integrals of the form \[ \int_0^. \int_{\mathbb R} f_n(s,x)(\mu_n-\nu_n)(ds,dx) \] under local linear growth and locally uniform convergence assumptions on the integrands. The proof is based on two structural mechanisms. The first is a forbidden bands principle, showing that quadratic variation convergence prevents jumps from crossing suitably chosen moving threshold regions. The second is a compensator mass control mechanism, which combines threshold-separated alignment of large predictable jumps with a counting argument for the associated compensator atoms. The results require neither semimartingale convergence, convergence of characteristics, uniform tightness, nor global structural assumptions such as independence, stationarity, or Markovianity. More broadly, they show that quadratic variation convergence imposes a substantially stronger rigidity on the jump organization of c\`adl\`ag processes than one might initially expect.

math.PR

An Extremal Reconstruction Principle under Covariance Domination

We identify a structural extremal principle governing residual \(L^2\)-norms over operator-ordered covariance envelopes. In contrast to the centered setting, where such quantities reduce to trace expressions involving covariance operators, the non-centered framework generates mixed terms that cannot be recovered from covariance ordering alone. We show that the worst-case squared residual \(L^2\)-norm over an operator-ordered covariance envelope is attained at a canonical envelope representative, possibly belonging only to the closure of the admissible class. The resulting extremal identity holds uniformly over all admissible reconstruction operators. The result is obtained without convexity, compactness, or a global Hilbert space structure governing all components of the system. As a consequence, the associated minimax reconstruction problem over covariance envelopes reduces to evaluation at a canonical representative under covariance domination.

math.FA

Functional structural equation models with out-of-sample guarantees

Statistical learning methods typically assume that the training and test data originate from the same distribution, enabling effective risk minimization. However, real-world applications frequently involve distributional shifts, leading to poor model generalization. To address this, recent advances in causal inference and robust learning have introduced strategies such as invariant causal prediction and anchor regression. While these approaches have been explored for traditional structural equation models (SEMs), their extension to functional systems remains limited. This paper develops a risk minimization framework for functional SEMs using linear, potentially unbounded operators. We introduce a functional worst-risk minimization approach, ensuring robust predictive performance across shifted environments. Our key contribution is a novel worst-risk decomposition theorem, which expresses the maximum out-of-sample risk in terms of observed environments. We establish conditions for the existence and uniqueness of the worst-risk minimizer and provide consistent estimation procedures. Empirical results on functional systems illustrate the advantages of our method in mitigating distributional shifts. These findings contribute to the growing literature on robust functional regression and causal learning, offering practical guarantees for out-of-sample generalization in dynamic environments.

math.ST

Functional worst risk minimization

The aim of this paper is to extend worst risk minimization, also called worst average loss minimization, to the functional realm. This means finding a functional regression representation that will be robust to future distribution shifts on the basis of data from two environments. In the classical non-functional realm, structural equations are based on a transfer matrix $B$. In section~\ref{sec:sfr}, we generalize this to consider a linear operator $\mathcal{T}$ on square integrable processes that plays the the part of $B$. By requiring that $(I-\mathcal{T})^{-1}$ is bounded -- as opposed to $\mathcal{T}$ -- this will allow for a large class of unbounded operators to be considered. Section~\ref{sec:worstrisk} considers two separate cases that both lead to the same worst-risk decomposition. Remarkably, this decomposition has the same structure as in the non-functional case. We consider any operator $\mathcal{T}$ that makes $(I-\mathcal{T})^{-1}$ bounded and define the future shift set in terms of the covariance functions of the shifts. In section~\ref{sec:minimizer}, we prove a necessary and sufficient condition for existence of a minimizer to this worst risk in the space of square integrable kernels. Previously, such minimizers were expressed in terms of the unknown eigenfunctions of the target and covariate integral operators (see for instance \cite{HeMullerWang} and \cite{YaoAOS}). This means that in order to estimate the minimizer, one must first estimate these unknown eigenfunctions. In contrast, the solution provided here will be expressed in any arbitrary ON-basis. This completely removes any necessity of estimating eigenfunctions. This pays dividends in section~\ref{sec:estimation}, where we provide a family of estimators, that are consistent with a large sample bound. Proofs of all the results are provided in the appendix.

math.ST

Worst-risk minimization in generalized structural equation models

We consider rather general structural equation models (SEMs) between a target and its covariates in several shifted environments. Given $k\in\mathbb{N}$ shifts we consider the set of shifts that are at most $γ$-times as strong as a given weighted linear combination of these $k$ shifts and the worst (quadratic) risk over this entire space. This worst risk has a nice decomposition which we refer to as the "worst risk decomposition". Then we find an explicit arg-min solution that minimizes the worst risk and consider its corresponding plug-in estimator which is the main object of this paper. This plug-in estimator is (almost surely) consistent and we first prove a concentration in measure result for it. The solution to the worst risk minimizer is rather reminiscent of the corresponding ordinary least squares solution in that it is product of a vector and an inverse of a Grammian matrix. Due to this, the central moments of the plug-in estimator is not well-defined in general, but we instead consider these moments conditioned on the Grammian inverse being bounded by some given constant. We also study conditional variance of the estimator with respect to a natural filtration for the incoming data. Similarly we consider the conditional covariance matrix with respect to this filtration and prove a bound for the determinant of this matrix. This SEM model generalizes the linear models that have been studied previously for instance in the setting of casual inference or anchor regression but the concentration in measure result and the moment bounds are new even in the linear setting.

math.ST

Optimal worst-risk minimization in structural equation models with random coefficients

The insight that causal parameters are particularly suitable for out-of-sample prediction has sparked a lot development of causal-like predictors. However, the connection with strict causal targets, has limited the development with good risk minimization properties, but without a direct causal interpretation. In this manuscript we derive the optimal out-of-sample risk minimizing predictor of a certain target $Y$ in a non-linear system $(X,Y)$ that has been trained in several within-sample environments. We consider data from an observation environment, and several shifted environments. Each environment corresponds to a structural equation model (SEM), with random coefficients and with its own shift and noise vector, both in $L^2$. Unlike previous approaches, we also allow shifts in the target value. We define a sieve of out-of-sample environments, consisting of all shifts $\tilde{A}$ that are at most $γ$ times as strong as any weighted average of the observed shift vectors. For each $β\in\mathbb{R}^p$ we show that the supremum of the risk functions $R_{\tilde{A}}(β)$ has a worst-risk decomposition into a (positive) non-linear combination of risk functions, depending on $γ$. We then define the set $\mathcal{B}_γ$, as minimizers of this risk. The main result of the paper is that there is a unique minimizer ($|\mathcal{B}_γ|=1$) that can be consistently estimated by an explicit estimator, outside a set of zero Lebesgue measure in the parameter space. A practical obstacle for the initial method of estimation is that it involves the solution of a general degree polynomials. Therefore, we prove that an approximate estimator using the bisection method is also consistent.

math.ST

Constructive and consistent estimation of quadratic minimax

We consider $k$ square integrable random variables $Y_1,...,Y_k$ and $k$ random (row) vectors of length $p$, $X_1,...,X_k$ such that $X_i(l)$ is square integrable for $1\le i\le k$ and $1\le l\le p$. No assumptions whatsoever are made of any relationship between the $X_i$:s and $Y_i$:s. We shall refer to each pairing of $X_i$ and $Y_i$ as an environment. We form the square risk functions $R_i(\beta)=\mathbb{E}\left[(Y_i-\beta X_i)^2\right]$ for every environment and consider $m$ affine combinations of these $k$ risk functions. Next, we define a parameter space $\Theta$ where we associate each point with a subset of the unique elements of the covariance matrix of $(X_i,Y_i)$ for an environment. Then we study estimation of the $\arg\min$-solution set of the maximum of a the $m$ affine combinations the of quadratic risk functions. We provide a constructive method for estimating the entire $\arg\min$-solution set which is consistent almost surely outside a zero set in $\Theta^k$. This method is computationally expensive, since it involves solving polynomials of general degree. To overcome this, we define another approximate estimator that also provides a consistent estimation of the solution set based on the bisection method, which is computationally much more efficient. We apply the method to worst risk minimization in the setting of structural equation models.

math.ST

Stability in quadratic variation

Consider a sequence of cadlag processes $\{X^n\}_n$, and some fixed function $f$. If $f$ is continuous then under several modes of convergence $X^n\to X$ implies corresponding convergence of $f(X^n)\to f(X)$, due to continuous mapping. We study conditions (on $f$, $\{X^n\}_n$ and $X$) under which convergence of $X^n\to X$ implies $\left[f(X^n)-f(X)\right]\to 0$. While interesting in its own right, this also directly relates (through integration by parts and the Kunita-Watanabe inequality) to convergence of integrators in the sense $\int_0^t Y_{s-}df(X^n_s)\to\int_0^t Y_{s-}df(X_s)$. We use two different types of quadratic variations, weak sense and strong sense which our two main results deal with. For weak sense quadratic variations we show stability when $f\in C^1$, $\{X^n\}_n,X$ are Dirichlet processes defined as in \cite{NonCont} $X^n\xrightarrow{a.s.}X$, $[X^n-X]\xrightarrow{a.s.}0$ and $\{(X^n)^*_t\}_n$ is bounded in probability. For strong sense quadratic variations we are able to relax the conditions on $f$ to being the primitive function of a cadlag function but with the additional assumption on $X$, that the continuous and discontinuous parts of $X$ are independent stochastic processes (this assumption is not imposed on $\{X^n\}_n$ however), and $\{X^n\}_n,X$ are Dirichlet processes with quadratic variations along any stopping time refining sequence. To prove the result regarding strong sense quadratic variation we prove a new Itô decomposition for this setting.

math.PR

Stability in quadratic variation, with applications

We show that non continuous Dirichlet processes, defined as in \cite{NonCont} are closed under a wide family of locally Lipschitz continuous maps (similar to the time-homogeneous variants of the maps considered in \cite{Low}) thus extending Theorem 2.1. from that paper. We provide an Itô formula for these transforms and apply it to study of how $[f(X^n)-f(X)]\to 0$ when $X^n\to X$ (in some appropriate sense) for certain Dirichlet processes $\{X^n\}_n$, $X$ and certain locally Lipschitz continuous maps. We also consider how $[f_n(X^n)-f(X)]\to 0$ for $C^1$ maps $\{f_n\}_n$, $f$ when $f_n'\to f'$ uniformly on compacts. For applications we give examples of jump removal and stability of integrators.

math.PR

A local barycentric version of the Bak-Sneppen model

We study the behaviour of the interacting particle system, arising from the Bak-Sneppen model and Jante's law process. Let $N$ vertices be placed on a circle, such that each vertex has exactly two neighbours. To each vertex assign a real number, called {\em fitness}. Now find the vertex which fitness deviates most from the average of the fitnesses of its two immediate neighbours (in case of a tie, draw uniformly among such vertices), and replace it by a random value drawn independently according to some distribution $ζ$. We show that in case where $ζ$ is a uniform or a discrete uniform distribution, all the fitnesses except one converge to the same value.

math.PR

Convergence in the $p$-contest

We study asymptotic properties of the following Markov system of $N \geq 3$ points in~$[0,1]$. At each time step, the point farthest from the current centre of mass, multiplied by a constant $p>0$, is removed and replaced by an independent $ζ$-distributed point; the problem, inspired by variants of the Bak--Sneppen model of evolution and called a $p$-contest, was posed in [Grinfeld, M, Knight, P.A., and Wade, A.R. Rank-driven Markov processes, J. Stat. Phys. 146 (2012)]. We obtain various criteria for the convergences of the system, both for $p<1$ and $p>1$. In particular, when $p<1$ and $ζ\sim U[0,1]$, we show that the limiting configuration converges to zero. When $p>1$, we show that the configuration must converge to either zero or one, and we present an example where both outcomes are possible. Finally, when $p>1$, $N=3$ and $ζ$ satisfies certain conditions (e.g.~$ζ\sim U[0,1]$), we prove that the configuration can only converge to one a.s. Our paper substantially extends the results of [Grinfeld, M., Volkov, S., and Wade, A.R. Convergence in a multidimensional randomized Keynesian beauty contest. Adv. in Appl. Probab. 47 (2015)] and [Kennerberg, P., and Volkov, S. Jante's law process. Adv. in Appl. Probab. 50 (2018)] where it was assumed that $p=1$. Unlike the previous models, one can no longer use the Lyapunov function based just on the radius of gyration; when $0<p<1$ one has to find a much finer tuned function which turns out to be a supermartingale; the proof of this fact constitutes an unwieldy, albeit necessary, part of the paper.

math.PR

Jante's law process

Consider the process which starts with $N\ge 3$ distinct points on ${\mathbb R}^d$, and fix a positive integer~$K<N$. Of the total $N$ points keep those $N-K$ which minimize the energy (defined as the sum of all pairwise distances squared) amongst all the possible subsets of size $N-K$, and then replace the removed points by $K$ i.i.d.\ points sampled according to some fixed distribution $ζ$. Repeat this process ad infinitum. We obtain various quite non-restrictive conditions under which the set of points converges to a certain limit. This is a very substantial generalization of the "Keynesian beauty contest process" studied by Grinfeld, Volkov and Wade, where $K=1$ and the distribution $ζ$ was uniform on the unit cube.

math.PR