SearcharxivSearch

arXiv subjects

Bert Zwart

Publications and source records attributed to Bert Zwart.

At least 19 recordsLinked to original sources

Decision-Centric Large Deviations for Data-Driven Capital Buffers in Ruin Models

We consider an insurance risk model with a random walk structure in which the underlying probability law is unknown. A decision maker observes a statistic $Q_n$ computed from $n$ historical observations and then needs to choose a capital buffer of the form $C_n=n f(c,Q_n)$, where $c=\log(1/\delta)/n$ balances the amount of data and the tolerated ruin probability $\delta$. The classical safe capital buffer is inversely proportional to the adjustment coefficient $\gamma$, the exponential rate at which the ruin probability decays; when the law is unknown, this coefficient must be inferred from the data. The profile $f$ couples the statistical cost of observing an atypical historical statistic with the future ruin exponent induced by the resulting decision. We study the joint ex ante probability (over both the historical sample and an independent future risk process) that the future maximum exceeds the data-dependent buffer. We show that the naive plug-in rule fails to achieve the prescribed logarithmic decay exponent $c$, illustrating the adverse impact of model uncertainty when making decisions under rare-event constraints. We then identify a profile $f^*=f^*(c,Q_n)$ with the following appealing properties: its lower semicontinuous majorants are safe, while regular continuous rules that fall strictly below $f^*$ are, under mild conditions, unsafe. We illustrate the potential applicability of our framework by developing three parametric examples. In nonparametric settings, we show that a single exponential envelope leads to degenerate capital buffers, whereas a two-level exponential envelope yields a nondegenerate capital buffer which we prove to be safe.

math.ST

Optimal Estimators for Heavy-Tailed Mean Estimation via Convex Analysis

We study optimal estimation of the location parameter of a distribution known only to lie in a symmetric moment class $\mathcal C_0$: the mean-zero distributions with bounded moment $\int\phi\, d\mathbb P\le B$ for a fixed even $\phi$. Our main result concerns the fixed-margin regime, where the error margin $\Delta$ is fixed as $n\to\infty$: we give an exact large-deviation characterization of the smallest worst-case probability $\beta_n(\Delta)$ of an error exceeding $\Delta$ that any measurable estimator can guarantee with $n$ observations. Its exponential rate is exactly a two-point Hellinger exponent over the class shifted to means $\pm\Delta$, $r(\Delta)=-\log\sup_{\mathbb P_{\pm\Delta}\in\mathcal C_{\pm\Delta}}\int\sqrt{d\mathbb P_{-\Delta}\, d\mathbb P_{\Delta}}$, achieved non-asymptotically, $\beta_n(\Delta)\le e^{-nr(\Delta)}$, by a monotone $M$-estimator synthesized from a two-parameter convex program. Lagrangian duality collapses the infinite-dimensional search over estimating functions to two multipliers, which determine a pair of envelopes characterizing the optimal estimating functions; the sandwich shape posited ad hoc in prior constructions emerges naturally. For bounded variance ($\phi(x)=x^2$, $B=\sigma^2$) the exponent is $r(\Delta)=\tfrac12\log(1+\Delta^2/\sigma^2)$. In the fixed-confidence regime, holding $\beta$ fixed and letting the optimal margin $\Delta_n(\beta)$ shrink with $n$, the same synthesis stays optimal to leading order for several concrete classes. As $\beta\downarrow0$ it attains the sharp constant $\sqrt2$ of Catoni for bounded variance and the constant $L(\alpha)$ of Lee and Bhatt et al. for bounded $\alpha$-moments, $\alpha\in(1,2)$, thereby shown tight; for slowly varying $\phi$ it is leading-order minimax at every fixed $\beta$. The least-favorable distributions are simple, supported on at most three atoms.

math.ST

Large deviations for subgraphs in inhomogeneous random graphs

Inhomogeneous random graphs are fundamental models for real-world networks, where prescribed degrees are imposed as soft constraints. A common assumption in such models is that the degree distribution follows a power-law, capturing the heavy-tailed nature observed in many contexts. While various graph functionals have been studied in this setting, inhomogeneity makes their analysis significantly more challenging. The goal of this paper is to investigate the large deviations of subgraph counts in inhomogeneous random graphs. Rare events concerning these functionals translate into quantifying the probability that extremely large hubs appear in the graph. This can be achieved by defining a specific optimization problem that captures the most likely way to generate numerous additional subgraphs. When the expected number of subgraphs is sublinear in the graph size, polynomially large deviations are possible, and in this case, we can derive sharp results on clique counts.

math.PR

Heavy tails in dynamic flow networks: Universal explanation of their emergence

Overload-induced cascading failures can cause extreme disruptions in a wide range of networked systems, such as power grids, transportation networks, or financial systems. Empirical studies across domains report that the size of such disruptions often follows a Pareto- or heavy-tailed distribution. While many models reproduce this scaling behavior, they are either tailored to specific domains or based on simplified mechanisms that overlook key aspects of overload cascading behavior. Hence, a general understanding of the mechanisms driving scale-free behavior in these settings remains incomplete. In this paper, we develop a universal and analytically tractable model of overload cascading failures on flow networks, offering a new perspective on how Pareto-tailed disruptions emerge across networks. Our framework shows, under mild assumptions, that heavy-tailed disruptions can arise naturally from Pareto-tailed external inputs, and it establishes a transformation law linking the input and output tail exponents. We further identify broad conditions under which the resulting cascade cost exhibits a heavy-tailed distribution and show that the mechanism is robust across several domains, including power transmission, traffic networks, and processing systems. Our results provide a unified explanation for the emergence of scale-free failures in overload-driven systems and connect previously disparate, application-specific models under a unified framework.

physics.soc-ph

Sample Path Large Deviations for Multivariate Heavy-Tailed Hawkes Processes and Related L\'evy Processes

In this paper, we develop sample path large deviations for multivariate Hawkes processes with heavy-tailed mutual excitation rates. Our results address a broad class of rare events in Hawkes processes at the sample path level and, via the cluster representation of Hawkes processes and a recent result on the tail asymptotics of the cluster sizes, unravel the most likely configurations of (multiple) large clusters that could trigger the target events. Our proof hinges on establishing the asymptotic equivalence, in terms of M-convergence, between a suitably scaled multivariate Hawkes process and a coupled L\'evy process with multivariate hidden regular variation. Hence, along the way, we derive a sample path large deviations principle for a class of L\'evy processes with multivariate hidden regular variation, which not only plays an auxiliary role in our analysis but is also of independent interest.

math.PR

Robust Mean Estimation for Optimization: The Impact of Heavy Tails

We consider the problem of constructing a least conservative estimator of the expected value $\mu$ of a non-negative heavy-tailed random variable. We require that the probability of overestimating the expected value $\mu$ is kept appropriately small; a natural requirement if its subsequent use in a decision process is anticipated. In this setting, we show it is optimal to estimate $\mu$ by solving a distributionally robust optimization (DRO) problem using the Kullback-Leibler (KL) divergence. We further show that the statistical properties of KL-DRO compare favorably with other estimators based on truncation, variance regularization, or Wasserstein DRO.

math.OC

Tail Asymptotics of Cluster Sizes in Multivariate Heavy-Tailed Hawkes Processes

We examine a distributional fixed-point equation related to a multi-type branching process that is key in the cluster sizes analysis of multivariate heavy-tailed Hawkes processes. Specifically, we explore the tail behavior of its solution and demonstrate the emergence of a form of multivariate hidden regular variation. Large values of the cluster size vector result from one or several significant jumps. A discrete optimization problem involving any given rare event set of interest determines the exact configuration of these large jumps and the degree of hidden regular variation. Our proofs rely on a detailed probabilistic analysis of the spatiotemporal structure of multiple large jumps in multi-type branching processes.

math.PR

Emergence of Scale-Free Traffic Jams in Highway Networks: A Probabilistic Approach

Traffic congestion continues to escalate with urbanization and socioeconomic development, necessitating advanced modeling to understand and mitigate its impacts. In large-scale networks, traffic congestion can be studied using cascade models, where congestion not only impacts isolated segments, but also propagates through the network in a domino-like fashion. One metric for understanding these impacts is congestion cost, which is typically defined as the additional travel time caused by traffic jams. Recent data suggests that congestion cost exhibits a universal scale-free-tailed behavior. However, the mechanism driving this phenomenon is not yet well understood. To address this gap, we propose a stochastic cascade model of traffic congestion. We show that traffic congestion cost is driven by the scale-free distribution of traffic intensities. This arises from the catastrophe principle, implying that severe congestion is likely caused by disproportionately large traffic originating from a single location. We also show that the scale-free nature of congestion cost is robust to various congestion propagation rules, explaining the universal scaling observed in empirical data. These findings provide a new perspective in understanding the fundamental drivers of traffic congestion and offer a unifying framework for studying congestion phenomena across diverse traffic networks.

physics.soc-ph

SIR on locally converging dynamic random graphs

In this paper, we study the trajectory of a classic SIR epidemic on a family of dynamic random graphs of fixed size, whose set of edges continuously evolves over time. We set general infection and recovery times, and start the epidemic from a positive, yet small, proportion of vertices. We show that in such a case, the spread of an infectious disease around a typical individual can be approximated by the spread of the disease in a local neighbourhood of a uniformly chosen vertex. We formalize this by studying general dynamic random graphs that converge dynamically locally in probability and demonstrate that the epidemic on these graphs converges to the epidemic on their dynamic local limit graphs. We provide a detailed treatment of the theory of dynamic local convergence, which remains a relatively new topic in the study of random graphs. One main conclusion of our paper is that a specific form of dynamic local convergence is required for our results to hold.

math.PR

Bidding in Ancillary Service Markets: An Analytical Approach Using Extreme Value Theory

To enable the participation of stochastic distributed energy resources in ancillary service markets, the Danish transmission system operator, Energinet, mandates that flexibility providers satisfy a minimum 90% reliability requirement for reserve bids. This paper examines the bidding strategy of an electric vehicle aggregator under this regulation and develops a chance-constrained optimization model. In contrast to conventional sample-based approaches that demand large datasets to capture uncertainty, we propose an analytical reformulation that leverages extreme value theory to characterize the tail behavior of flexibility distributions. A case study with real-world charging data from 1400 residential electric vehicles in Denmark demonstrates that the analytical solution improves out-of-sample reliability, reducing bid violation rates by up to 8% relative to a sample-based benchmark. The method is also computationally more efficient, solving optimization problems up to 4.8 times faster while requiring substantially fewer samples to ensure compliance. Moreover, the proposed approach enables the construction of feasible bids with reliability levels as high as 99.95%, which would otherwise require prohibitively large scenario sets under the sample-based method. Beyond its computational and reliability advantages, the framework also provides actionable insights into how reliability thresholds influence aggregator bidding behavior and market participation. This study establishes a regulation-compliant, tractable, and risk-aware bidding methodology for stochastic flexibility aggregators, enhancing both market efficiency and power system security.

eess.SY

Dynamic Dimensioning of Frequency Containment Reserves: The Case of the Nordic Grid

One of the main responsibilities of a Transmission System Operator (TSO) operating an electric grid is to maintain a designated frequency (e.g., 50 Hz in Europe). To achieve this, TSOs have created several products called frequency-supporting ancillary services. The Frequency Containment Reserve (FCR) is one of these ancillary service products. This article focuses on the TSO problem of determining the volume procured for FCR. Specifically, we investigate the potential benefits and impact on grid security when transitioning from a traditionally \textit{static} procurement method to a \textit{dynamic} strategy for FCR volume. We take the Nordic synchronous area in Europe as a case study and use a diffusion model to capture its frequency development. We introduce a controlled mean reversal parameter to assess changes in FCR obligations, in particular for the Nordic FCR-N ancillary service product. We establish closed-form expressions for exceedance probabilities and use historical frequency data as input to calibrate the model. We show that a dynamic dimensioning approach for FCR has the potential to significantly reduce the exceedance probabilities (up to $37\%$) while maintaining the total yearly procured FCR volume equal to that of the current static approach. Alternatively, a dynamic dimensioning approach could significantly increase security at limited extra cost.

eess.SY

Optimization under rare events: scaling laws for linear chance-constrained programs

We consider a class of chance-constrained programs in which profit needs to be maximized while enforcing that a given adverse event remains rare. Using techniques from large deviations and extreme value theory, we show how the optimal value scales as the prescribed bound on the violation probability becomes small and how convex programs emerge in the limit. We use our results to analyze the performance of existing popular approaches in the rare-event regime. We show that the popular CVaR and sample approximations have optimality properties under light-tailed assumptions on the randomness, while they behave sub-optimal in a heavy-tailed setting. Our results are derived using large deviations theory, extreme value theory, process techniques, and random set theory.

math.OC

Large deviations of the giant component in scale-free inhomogeneous random graphs

We study large deviations of the size of the largest connected component in a general class of inhomogeneous random graphs with iid weights, parametrized so that the degree distribution is regularly varying. We derive a large-deviation principle with logarithmic speed: the rare event that the largest component contains linearly more vertices than expected is caused by the presence of constantly many vertices with linear degree. Conditionally on this rare event, we prove distributional limits of the weight distribution and component-size distribution.

math.PR

Sample-path large deviations for a class of heavy-tailed Markov additive processes

For a class of additive processes driven by the affine recursion $X_{n+1} = A_n X_n + B_n$, we develop a sample-path large deviations principle in the $M_1'$ topology on $D [0,1]$. We allow $B_n$ to have both signs and focus on the case where Kesten's condition holds on $A_1$, leading to heavy-tailed distributions. The most likely paths in our large deviations results are step functions with both positive and negative jumps.

math.PR

Large deviations for triangles in scale-free random graphs

We provide large deviations estimates for the upper tail of the number of triangles in scale-free inhomogeneous random graphs where the degrees have power law tails with index $-α, α\in (1,2)$. We show that upper tail probabilities for triangles undergo a phase transition. For $α<4/3$, the upper tail is caused by many vertices of degree of order $n$, and this probability is semi-exponential. In this regime, additional triangles consist of two hubs. For $α>4/3$ on the other hand, the upper tail is caused by one hub of a specific degree, and this probability decays polynomially in $n$, leading to additional triangles with one hub. In the intermediate case $α=4/3$, we show polynomial decay of the tail probability caused by multiple but finitely many hubs. In this case, the additional triangles contain either a single hub or two hubs. Our proofs are partly based on various concentration inequalities. In particular, we tailor concentration bounds for empirical processes to make them well-suited for analyzing heavy-tailed phenomena in nonlinear settings.

math.PR

Optimization of inventory and capacity in large-scale assembly systems using extreme-value theory

High-tech systems are typically produced in two stages: 1) Production of components using specialized equipment and staff; 2) System assembly/integration. Component production capacity is subject to fluctuations, causing a high risk of shortages of at least one component, which results in costly delays. Companies hedge this risk by strategic investments in excess production capacity and in buffer inventories of components. To optimize these, it is crucial to characterize the relation between component shortage risk and capacity and inventory investments. We suppose that component production capacity and produce demand are normally distributed over finite time intervals, and we accordingly model the production system as a symmetric fork-join queueing network with $N$ statistically identical queues with a common arrival process and independent service processes. Assuming a symmetric cost structure, we subsequently apply extreme value theory to gain analytic insights into this optimization problem. We derive several new results for this queueing network, notably that the scaled maximum of $N$ steady-state queue lengths converges in distribution to a Gaussian random variable. These results translate into asymptotically optimal methods to dimension the system. Tests on a range of problems reveal that these methods typically work well for systems of moderate size.

math.PR

Scale-free graphs with many edges

We develop tail estimates for the number of edges in a Chung-Lu random graph with regularly varying weight distribution. Our results show that the most likely way to have an unusually large number of edges is through the presence of one or more hubs, i.e.\ vertices with degree of order $n$.

math.PR

Grid-Aware Real-Time Control and Balancing Between Microgrids

Due to the energy transition, lots of research has been conducted within the last decade on the topics of energy management systems or local energy trading approaches, often on the day-ahead or intraday level. A large majority of these approaches focuses on 15 or 60-minute time intervals for their operation, however, the question of how the planned solutions are realized within these time intervals is often left unanswered. Within this work, we aim to close this gap and propose a real-time balancing and control approach for a set of microgrids, which implements the day-ahead solutions. The approach is based on a three-step framework, in which the first step consists of ensuring the feasibility of devices within the microgrids. The second step focuses on the grid constraints of the connecting medium voltage grid using the DC power flow formulation due to the running time requirements of a real-time approach. The last step is to propagate the solution into the individual microgrids, where the allocated power needs to be distributed among the devices and households. Within a case study, we show that the proposed real-time control approach works as intended and is comparable to an optimal offline algorithm under some mild assumptions.

eess.SY