SearcharxivSearch

arXiv subjects

Claudia Czado

Publications and source records attributed to Claudia Czado.

At least 19 recordsLinked to original sources

Calibrating simplified vine copulas with a noise contrastive estimation approach

Vine copulas provide a flexible framework for modeling complex multivariate dependence structures using only bivariate building blocks. Their practical success relies heavily on the simplifying assumption, which restricts conditional pair copulas to be independent of the specific conditioning values. While this assumption greatly facilitates estimation, it may lead to model misspecification in applications with pronounced varying conditional dependence. We propose a novel calibration strategy for simplified vine copula models based on observation-specific correction factors. These factors are derived using noise contrastive estimation (NCE), a supervised learning technique for density estimation that reframes the problem as a binary classification task with an easily sampled noise distribution. Treating the fitted simplified vine copula as the noise model, the NCE approach yields corrected log-likelihood estimates for individual observations, thereby locally adjusting the simplified vine toward the underlying data-generating dependence structure. Simulation studies demonstrate that the proposed calibration provides sensible and effective adjustments, improving model accuracy when the simplifying assumption is violated while remaining neutral when the simplified model is adequate. Two real-data applications further illustrate the practical benefits of the method. The results highlight NCE-based calibration as a promising tool to enhance simplified vine copula models without abandoning their computational tractability.

stat.ME

Stepwise Variational Inference with Vine Copulas

We propose stepwise variational inference (VI) with vine copulas: a universal VI procedure that combines vine copulas with a novel stepwise estimation procedure of the variational parameters. Vine copulas consist of a nested sequence of trees built from copulas, where more complex latent dependence can be modeled with increasing number of trees. We propose to estimate the vine copula approximate posterior in a stepwise fashion, tree by tree along the vine structure. Further, we show that the usual backward Kullback-Leibler divergence cannot recover the correct parameters in the vine copula model, thus the evidence lower bound is defined based on the Rényi divergence. Finally, an intuitive stopping criterion for adding further trees to the vine eliminates the need to pre-define a complexity parameter of the variational distribution, as required for most other approaches. Thus, our method interpolates between mean-field VI (MFVI) and full latent dependence. In many applications, in particular sparse Gaussian processes, our method is parsimonious with parameters, while outperforming MFVI.

stat.ML

Bivariate Postprocessing of Wind Vectors

To quantify the uncertainty in numerical weather prediction (NWP) forecasts, ensemble prediction systems are utilized. Although NWP forecasts continuously improve, they suffer from systematic bias and dispersion errors. To obtain well calibrated and sharp predictive probability distributions, statistical postprocessing methods are applied to NWP output. Recent developments focus on multivariate postprocessing models incorporating dependencies directly into the model. We introduce three novel bivariate postprocessing approaches, and analyze their performance for joint postprocessing of bivariate wind vector components for 60 stations in Germany. Bivariate vine copula based models, a bivariate gradient boosted version of ensemble model output statistics (EMOS), and a bivariate distributional regression network (DRN) are compared to bivariate EMOS. The case study indicates that the novel bivariate methods improve over the bivariate EMOS approaches. The bivariate DRN and the most flexible version of the bivariate vine copula approach exhibit the best performance in terms of verification scores and calibration.

stat.AP

Sampling from Conditional Distributions of Simplified Vines

Simplified vine copulas are flexible tools over standard multivariate distributions for modeling and understanding different dependence properties in high-dimensional data. Their conditional distributions are of utmost importance, from statistical learning to graphical models. However, the conditional densities of vine copulas and, thus, vine distributions cannot be obtained in closed form without integration for all possible sets of conditioning variables. We propose a Markov Chain Monte Carlo based approach of using Hamiltonian Monte Carlo to sample from any conditional distribution of arbitrarily specified simplified vine copulas and thus vine distributions. We show its accuracy through simulation studies and analyze data of multiple maize traits such as flowering times, plant height, and vigor. Use cases from predicting traits to estimating conditional Kendall's tau are presented.

stat.ME

TVineSynth: A Truncated C-Vine Copula Generator of Synthetic Tabular Data to Balance Privacy and Utility

We propose TVineSynth, a vine copula based synthetic tabular data generator, which is designed to balance privacy and utility, using the vine tree structure and its truncation to do the trade-off. Contrary to synthetic data generators that achieve DP by globally adding noise, TVineSynth performs a controlled approximation of the estimated data generating distribution, so that it does not suffer from poor utility of the resulting synthetic data for downstream prediction tasks. TVineSynth introduces a targeted bias into the vine copula model that, combined with the specific tree structure of the vine, causes the model to zero out privacy-leaking dependencies while relying on those that are beneficial for utility. Privacy is here measured with membership (MIA) and attribute inference attacks (AIA). Further, we theoretically justify how the construction of TVineSynth ensures AIA privacy under a natural privacy measure for continuous sensitive attributes. When compared to competitor models, with and without DP, on simulated and on real-world data, TVineSynth achieves a superior privacy-utility balance.

cs.LG

Assessing univariate and bivariate risks of late-frost and drought using vine copulas: A historical study for Bavaria

In light of climate change's impacts on forests, including extreme drought and late-frost, leading to vitality decline and regional forest die-back, we assess univariate drought and late-frost risks and perform a joint risk analysis in Bavaria, Germany, from 1952 to 2020. Utilizing a vast dataset with 26 bioclimatic and topographic variables, we employ vine copula models due to the data's non-Gaussian and asymmetric dependencies. We use D-vine regression for univariate and Y-vine regression for bivariate analysis, and propose corresponding univariate and bivariate conditional probability risk measures. We identify "at-risk" regions, emphasizing the need for forest adaptation due to climate change.

stat.AP

Bivariate vine copula based regression, bivariate level and quantile curves

The statistical analysis of univariate quantiles is a well developed research topic. However, there is a need for research in multivariate quantiles. We construct bivariate (conditional) quantiles using the level curves of vine copula based bivariate regression model. Vine copulas are graph theoretical models identified by a sequence of linked trees, which allow for separate modelling of marginal distributions and the dependence structure. We introduce a novel graph structure model (given by a tree sequence) specifically designed for a symmetric treatment of two responses in a predictive regression setting. We establish computational tractability of the model and a straight forward way of obtaining different conditional distributions. Using vine copulas the typical shortfalls of regression, as the need for transformations or interactions of predictors, collinearity or quantile crossings are avoided. We illustrate the copula based bivariate level curves for different copula distributions and show how they can be adjusted to form valid quantile curves. We apply our approach to weather measurements from Seoul, Korea. This data example emphasizes the benefits of the joint bivariate response modelling in contrast to two separate univariate regressions or by assuming conditional independence, for bivariate response data set in the presence of conditional dependence.

stat.ME

High-dimensional sparse vine copula regression with application to genomic prediction

High-dimensional data sets are often available in genome-enabled predictions. Such data sets include nonlinear relationships with complex dependence structures. For such situations, vine copula based (quantile) regression is an important tool. However, the current vine copula based regression approaches do not scale up to high and ultra-high dimensions. To perform high-dimensional sparse vine copula based regression, we propose two methods. First, we show their superiority regarding computational complexity over the existing methods. Second, we define relevant, irrelevant, and redundant explanatory variables for quantile regression. Then we show our method's power in selecting relevant variables and prediction accuracy in high-dimensional sparse data sets via simulation studies. Next, we apply the proposed methods to the high-dimensional real data, aiming at the genomic prediction of maize traits. Some data-processing and feature extraction steps for the real data are further discussed. Finally, we show the advantage of our methods over linear models and quantile regression forests in simulation studies and real data applications.

stat.ME

Vine Copula based portfolio level conditional risk measure forecasting

Accurately estimating risk measures for financial portfolios is critical for both financial institutions and regulators. However, many existing models operate at the aggregate portfolio level and thus fail to capture the complex cross-dependencies between portfolio components. To address this, a new approach is presented that uses vine copulas in combination with univariate ARMA-GARCH models for marginal modelling to compute conditional portfolio-level risk measure estimates by simulating portfolio-level forecasts conditioned on a stress factor. A quantile-based approach is then presented to observe the behaviour of risk measures given a particular state of the conditioning asset(s). In a case study of Spanish equities with different stress factors, the results show that the portfolio is quite robust to a sharp downturn in the American market. At the same time, there is no evidence of this behaviour with respect to the European market.

q-fin.PM

On the Observability of Gaussian Models using Discrete Density Approximations

This paper proposes a novel method for testing observability in Gaussian models using discrete density approximations (deterministic samples) of (multivariate) Gaussians. Our notion of observability is defined by the existence of the maximum a posteriori estimator. In the first step of the proposed algorithm, the discrete density approximations are used to generate a single representative design observation vector to test for observability. In the second step, a number of carefully chosen design observation vectors are used to obtain information on the properties of the estimator. By using measures like the variance and the so-called local variance, we do not only obtain a binary answer to the question of observability but also provide a quantitative measure.

eess.SY

Statistical Dependence Analyses of Operational Flight Data Used for Landing Reconstruction Enhancement

The RTS smoother is widely used for state estimation and it is utilized here to increase the data quality with respect to physical coherence and to increase resolution. The purpose of this paper is to enhance the performance of the RTS smoother to reconstruct an aircraft landing using on board recorded data only. Thereby, errors and uncertainties of operational flight data (e.g. altitude, attitude, position, speed) recorded during flights of civil aircraft are minimized. These data can be used for subsequent analyses in terms of flight safety or efficiency, which is commonly referred to as Flight Data Monitoring (FDM). Statistical assumptions of the smoother theory are not always verified during application but (consciously or not) assumed to be fulfilled. These assumptions can hardly be verified prior to the smoother application, however, they can be verified using the results of an initial smoother iteration and modifications of specific smoother characteristics can be suggested. This project specifically verifies assumptions on the measurement noise characteristics. Variance and covariance of the measurement noise can be checked after the initial smoother application. It is discovered that these characteristics change over time and should be accounted for with a time varying covariance matrix. This sequence of matrices is estimated by kernel smoothing and replaces an initially assumed fixed and diagonal covariance matrix used for the first smoother run. The results of this second smoother iteration are mostly improved compared to the initial iteration, i.e. the errors are significantly reduced. Subsequently, the remaining dependence structures of the residuals of the second smoother iteration can be captured by copula models. Their interpretation is useful for a revision of the physical model utilized by the RTS smoother.

stat.AP

Environmental, Social, Governance scores and the Missing pillar -- Why does missing information matter?

Environmental, Social, and Governance (ESG) scores measure companies' performance concerning sustainability and societal impact and are organized on three pillars: Environmental (E), Social (S), and Governance (G). These complementary non-financial ESG scores should provide information about the ESG performance and risks of different companies. However, the extent of not yet published ESG information makes the reliability of ESG scores questionable. To explicitly denote the not yet published information on ESG category scores, a new pillar, the so-called Missing (M) pillar, is formulated. Environmental, Social, Governance, and Missing (ESGM) scores are introduced to consider the potential release of new information in the future. Furthermore, an optimization scheme is proposed to compute ESGM scores, linking them to the companies' riskiness. By relying on the data provided by Refinitiv, we show that the ESGM scores strengthen the companies' risk relationship. These new scores could benefit investors and practitioners as ESG exclusion strategies using only ESG scores might exclude assets with a low score solely because of their missing information and not necessarily because of a low ESG merit.

q-fin.RM

An Application of D-vine Regression for the Identification of Risky Flights in Runway Overrun

In aviation safety, runway overruns are of great importance because they are the most frequent type of landing accidents. Identification of factors which contribute to the occurrence of runway overruns can help mitigate the risk and prevent such accidents. Methods such as physics-based and statistical-based models were proposed in the past to estimate runway overrun probabilities. However, they are either costly or require experts' knowledge. We propose a statistical approach to quantify the risk probability of an aircraft to exceed a threshold at the speed of 80 knots given a set of influencing factors. This copula based D-vine regression approach is used because it allows for complex tail dependence and is computationally tractable. Data obtained from the Quick Access Recorder (QAR) for 711 flights are analyzed. We identify 41 flights with an estimated risk probability > 0.001 for a chosen threshold and rank the effects of each influencing factor for these flights. Also, the complex dependency patterns between some influencing factors for the 41 flights are shown to be non symmetric. The D-vine regression approach, compared to physics-based and statistical-based approaches, has an analytical solution, is not simulation based and can be used to estimate very small or large probabilities efficiently.

stat.AP

Analysis of an interventional protein experiment using a vine copula based structural equation model

While there is considerable effort to identify signaling pathways using linear Gaussian Bayesian networks from data, there is less emphasis of understanding and quantifying conditional densities and probabilities of nodes given its parents from the identifed Bayesian network. Most graphical models for continuous data assume a multivariate Gaussian distribution, which might be too restrictive. We re-analyse data from an experimental setting considered in Sachs et al. (2005) to illustrate the effects of such restrictions. For this we propose a novel non Gaussian nonlinear structural equation model based on vine copulas. In particular the D-vine regression approach of Kraus and Czado (2017) is adapted. We show that this model class is more suited to fit the data than the standard linear structural equation model based on the biological consent graph given in Sachs et al. (2005). The modelling approach also allows to study which pathway edges are supported by the data and which can be removed. For data experiment cd3cd28+aktinhib this approach identified three edges, which are no longer supported by the data. For each of these edges a plausible explanation based on underlying the experimental conditions could be found.

stat.AP

Nonparametric C- and D-vine based quantile regression

Quantile regression is a field with steadily growing importance in statistical modeling. It is a complementary method to linear regression, since computing a range of conditional quantile functions provides a more accurate modelling of the stochastic relationship among variables, especially in the tails. We introduce a non-restrictive and highly flexible nonparametric quantile regression approach based on C- and D-vine copulas. Vine copulas allow for separate modeling of marginal distributions and the dependence structure in the data, and can be expressed through a graph theoretical model given by a sequence of trees. This way we obtain a quantile regression model, that overcomes typical issues of quantile regression such as quantile crossings or collinearity, the need for transformations and interactions of variables. Our approach incorporates a two-step ahead ordering of variables, by maximizing the conditional log-likelihood of the tree sequence, while taking into account the next two tree levels. Further, we show that the nonparametric conditional quantile estimator is consistent. The performance of the proposed methods is evaluated in both low- and high-dimensional settings using simulated and real world data. The results support the superior prediction ability of the proposed models.

stat.ME

ESG, Risk, and (Tail) Dependence

While environmental, social, and governance (ESG) trading activity has been a distinctive feature of financial markets, the debate if ESG scores can also convey information regarding a company's riskiness remains open. Regulatory authorities, such as the European Banking Authority (EBA), have acknowledged that ESG factors can contribute to risk. Therefore, it is important to model such risks and quantify what part of a company's riskiness can be attributed to the ESG scores. This paper aims to question whether ESG scores can be used to provide information on (tail) riskiness. By analyzing the (tail) dependence structure of companies with a range of ESG scores, that is within an ESG rating class, using high-dimensional vine copula modelling, we are able to show that risk can also depend on and be directly associated with a specific ESG rating class. Empirical findings on real-world data show positive not negligible ESG risks determined by ESG scores, especially during the 2008 crisis.

q-fin.RM

Vine copula mixture models and clustering for non-Gaussian data

The majority of finite mixture models suffer from not allowing asymmetric tail dependencies within components and not capturing non-elliptical clusters in clustering applications. Since vine copulas are very flexible in capturing these types of dependencies, we propose a novel vine copula mixture model for continuous data. We discuss the model selection and parameter estimation problems and further formulate a new model-based clustering algorithm. The use of vine copulas in clustering allows for a range of shapes and dependency structures for the clusters. Our simulation experiments illustrate a significant gain in clustering accuracy when notably asymmetric tail dependencies or/and non-Gaussian margins within the components exist. The analysis of real data sets accompanies the proposed method. We show that the model-based clustering algorithm with vine copula mixture models outperforms the other model-based clustering techniques, especially for the non-Gaussian multivariate data.

stat.ME

Dependent censoring based on copulas

Consider a survival time T that is subject to random right censoring, and suppose that T is stochastically dependent on the censoring time C. We are interested in the marginal distribution of T. This situation is often encountered in practice. Consider for instance the case where T is the time to death of a patient suffering from a certain disease. Then, the censoring time C is for instance the time until the person leaves the study or the time until he/she dies from another disease. If the reason for leaving the study is related to the health condition of the patient or if he/she dies from a disease that has similar risk factors as the disease of interest, then T and C are likely dependent. In this paper we propose a new model that takes this dependence into account. The model is based on a parametric copula for the relationship between T and C, and on parametric marginal distributions for T and C. Unlike most other papers in the literature, we do not assume that the parameter defining the copula function is known. We give sufficient conditions on these parametric copula and marginals under which the bivariate distribution of (T;C) is identifed. These sufficient conditions are then checked for a wide range of common copulas and marginal distributions. We also study the estimation of the model, and carry out extensive simulations and the analysis of data on pancreas cancer to illustrate the proposed model and estimation procedure.

stat.ME