SearcharxivSearch

arXiv subjects

Carsten Limbach

Publications and source records attributed to Carsten Limbach.

4 recordsLinked to original sources

Dependence functions based on Chatterjee's rank correlation

We investigate a geometric and distributional reinterpretation of Chatterjee's $\xi$-coefficient, which measures functional dependence between a response variable $Y$ and a predictor vector $\mathbf{X}$. For this purpose, we analyze the Markov product $(Y,Y')$, where $Y'$ is a copy of $Y$ that is conditionally independent of $Y$ given $\mathbf{X}$. Based on this construction, we introduce and study two dependence functions, denoted by $\phi_{(Y,\mathbf{X})}$ and $\kappa_{(Y,\mathbf{X})}$. The proposed framework provides a geometric interpretation of the Markov product and extends Chatterjee's correlation coefficient to a richer and more interpretable object for the analysis of directed stochastic dependence. In particular, rather than only measuring how well $Y$ can be represented as a function of $\mathbf{X}$, the proposed dependence functions additionally quantify how strongly the corresponding Markov product is concentrated near the diagonal.

math.ST

On exact regions between measures of concordance and Chatterjee's rank correlation for lower semilinear copulas

We explore how the classical concordance measures - Kendall's $\tau$, Spearman's rank correlation $\rho$, and Spearman's footrule $\phi$ - relate to Chatterjee's rank correlation $\xi$ when restricted to lower semilinear copulas. First, we provide a complete characterization of the attainable $\tau$-$\rho$ region for this class, thus resolving the conjecture in [18]. Building on this result, we then derive the exact $\tau$-$\phi$ and $\phi$-$\rho$ regions, obtain a closed-form relationship between $\xi$ and $\tau$, and establish the exact $\tau$-$\xi$ region. In particular, we prove that $\xi$ never exceeds $\tau$, $\rho$, or $\phi$. Our results clarify the relationship between undirected and directed dependence measures and reveal novel insights into the dependence structures that result from lower semilinear copulas.

stat.ME

A dimension reduction for extreme types of directed dependence

In recent years, a variety of novel measures of dependence have been introduced being capable of characterizing diverse types of directed dependence, hence diverse types of how a number of predictor variables $\mathbf{X} = (X_1, \dots, X_p)$, $p \in \mathbb{N}$, may affect a response variable $Y$. This includes perfect dependence of $Y$ on $\mathbf{X}$ and independence between $\mathbf{X}$ and $Y$, but also less well-known concepts such as zero-explainability, stochastic comparability and complete separation. Certain such measures offer a representation in terms of the Markov product $(Y,Y')$, with $Y'$ being a conditionally independent copy of $Y$ given $\mathbf{X}$. This dimension reduction principle allows these measures to be estimated via the powerful nearest neighbor based estimation principle introduced in [4]. To achieve a deeper insight into the dimension reduction principle, this paper aims at translating the extreme variants of directed dependence, typically formulated in terms of the random vector $(\mathbf{X},Y)$, into the Markov product $(Y,Y')$.

math.ST

A new coefficient of separation

A coefficient is introduced that quantifies the extent of separation of a random variable $Y$ relative to a number of variables $\mathbf{X} = (X_1, \dots, X_p)$ by skillfully assessing the sensitivity of the relative effects of the conditional distributions. The coefficient is as simple as classical dependence coefficients such as Kendall's tau, also requires no distributional assumptions, and consistently estimates an intuitive and easily interpretable measure, which is $0$ if and only if $Y$ is stochastically comparable relative to $\mathbf{X}$, that is, the values of $Y$ show no location effect relative to $\mathbf{X}$, and $1$ if and only if $Y$ is completely separated relative to $\mathbf{X}$. As a true generalization of the classical relative effect, in applications such as medicine and the social sciences the coefficient facilitates comparing the distributions of any number of treatment groups or categories. It hence avoids the sometimes artificial grouping of variable values such as patient's age into just a few categories, which is known to cause inaccuracy and bias in the data analysis. The mentioned benefits are exemplified using synthetic and real data sets.

stat.ME