Searcharxiv⌕ Search

arXiv subjects

Bogdan Ćmiel

Publications and source records attributed to Bogdan Ćmiel.

11 recordsLinked to original sources

The post-hoc test for local dependence

The concept of independence plays a crucial role in probability theory and has been the subject of extensive research in recent years. Numerous approaches have been proposed to test for independence; however, most of them address the problem only at a global level. From a practical perspective, it is important not only to determine whether the data are dependent but also to identify where this dependence occurs and how strong it is. The graphical presentation of results is another essential aspect that should not be neglected, as it considerably enhances interpretability. The main objective of this work is to propose a solution that considers these aspects simultaneously. Relying on copula-based results, we introduce a novel method for testing global and local statistical independence using the quantile dependence function. Rather than assessing whether the value of the test statistic exceeds a single critical threshold and subsequently deciding whether to reject the independence hypothesis, we introduce so-called critical surfaces that guaranty a locally equal probability of exceeding them under independence. This approach enables a detailed examination of local discrepancies and an assessment of their statistical significance while preserving the overall significance level of the test.

stat.ME↗

Detecting dependence structure: visualization and inference

Identifying dependency between two random variables is a fundamental problem. The clear interpretability and ability of a procedure to provide information on the form of possible dependence is particularly important when exploring dependencies. In this paper, we introduce a novel method that employs a new estimator of the quantile dependence function and pertinent local acceptance regions. This leads to an insightful visualisation and a rigorous evaluation of the underlying dependence structure. We also propose a test of independence of two random variables, pertinent to this new estimator. Our procedures are based on ranks, and we derive a finite-sample theory that guarantees the inferential validity of our solutions at any given sample size. The procedures are simple to implement and computationally efficient. The large sample consistency of the proposed test is also proved. We show that, in terms of power, the new test is one of the best statistics for independence testing when considering a wide range of alternative models. Finally, we demonstrate the use of our approach to visualise dependence structure and to detect local departures from independence through analysing some real-world datasets.

stat.ME↗

Diagnostic tools for exploring differences in distributional properties between two samples: nonparametric approach

This paper reconsiders the problem of testing the equality of two unspecified continuous distributions. The framework, which we propose, allows for readable and insightful data visualisation and helps to understand and quantify how two groups of data differ. We consider a useful weighted rank empirical process on (0,1) and utilise a grid-based approach, based on diadic partitions of (0,1), to discretize the continuous process and construct local simultaneous acceptance regions. These regions help to identify statistically significant deviations from the null model. In addition, the form of the process and its dicretization lead to a highly interpretable visualisation of distributional differences. We also introduce a new two-sample test, explicitly related to the visualisation. Numerical studies show that the new test procedure performs very well. We illustrate the use and diagnostic capabilities of our approach by an application to a known set of neuroscience data.

stat.ME↗

Reproducibility Companion Paper: Describing Subjective Experiment Consistency by $p$-Value P-P Plot

In this paper we reproduce experimental results presented in our earlier work titled "Describing Subjective Experiment Consistency by $p$-Value P-P Plot" that was presented in the course of the 28th ACM International Conference on Multimedia. The paper aims at verifying the soundness of our prior results and helping others understand our software framework. We present artifacts that help reproduce tables, figures and all the data derived from raw subjective responses that were included in our earlier work. Using the artifacts we show that our results are reproducible. We invite everyone to use our software framework for subjective responses analyses going beyond reproducibility efforts.

cs.MM↗

Generalised Score Distribution: Underdispersed Continuation of the Beta-Binomial Distribution

A class of discrete probability distributions contains distributions with limited support. A typical example is some variant of a Likert scale, with response mapped to either the $\{1, 2, \ldots, 5\}$ or $\{-3, -2, \ldots, 2, 3\}$ set. An interesting subclass of discrete distributions with finite support are distributions limited to two parameters and having no more than one change in probability monotonicity. The main contribution of this paper is to propose a family of distributions fitting the above description, which we call the Generalised Score Distribution (GSD) class. The proposed GSD class covers the whole set of possible mean and variances, for any fixed and finite support. Furthermore, the GSD class can be treated as an underdispersed continuation of a reparametrized beta-binomial distribution. The GSD class parameters are intuitive and can be easily estimated by the method of moments. We also offer a Maximum Likelihood Estimation (MLE) algorithm for the GSD class and evidence that the class properly describes response distributions coming from 24 Multimedia Quality Assessment experiments. At last, we show that the GSD class can be represented as a sum of dichotomous zero-one random variables, which points to an interesting interpretation of the class.

stat.AP↗

Generalised Score Distribution: A Two-Parameter Discrete Distribution Accurately Describing Responses from Quality of Experience Subjective Experiments

Subjective responses from Multimedia Quality Assessment (MQA) experiments are conventionally analysed with methods not suitable for the data type these responses represent. Furthermore, obtaining subjective responses is resource intensive. A method allowing reuse of existing responses would be thus beneficial. Applying improper data analysis methods leads to difficult to interpret results. This encourages drawing erroneous conclusions. Building upon existing subjective responses is resource friendly and helps develop machine learning (ML) based visual quality predictors. We show that using a discrete model for analysis of responses from MQA subjective experiments is feasible. We indicate that our proposed Generalised Score Distribution (GSD) properly describes response distributions observed in typical MQA experiments. We highlight interpretability of GSD parameters and indicate that the GSD outperforms the approach based on sample empirical distribution when it comes to bootstrapping. We evidence that the GSD outcompetes the state-of-the-art model both in terms of goodness-of-fit and bootstrapping capabilities. To do all of that we analyse more than one million subjective responses from more than 30 subjective experiments. Furthermore, we make the code implementing the GSD model and related analyses available through our GitHub repository: https://github.com/Qub3k/subjective-exp-consistency-check

cs.MM↗

Describing Subjective Experiment Consistency by $p$-Value P-P Plot

There are phenomena that cannot be measured without subjective testing. However, subjective testing is a complex issue with many influencing factors. These interplay to yield either precise or incorrect results. Researchers require a tool to classify results of subjective experiment as either consistent or inconsistent. This is necessary in order to decide whether to treat the gathered scores as quality ground truth data. Knowing if subjective scores can be trusted is key to drawing valid conclusions and building functional tools based on those scores (e.g., algorithms assessing the perceived quality of multimedia materials). We provide a tool to classify subjective experiment (and all its results) as either consistent or inconsistent. Additionally, the tool identifies stimuli having irregular score distribution. The approach is based on treating subjective scores as a random variable coming from the discrete Generalized Score Distribution (GSD). The GSD, in combination with a bootstrapped G-test of goodness-of-fit, allows to construct $p$-value P-P plot that visualizes experiment's consistency. The tool safeguards researchers from using inconsistent subjective data. In this way, it makes sure that conclusions they draw and tools they build are more precise and trustworthy. The proposed approach works in line with expectations drawn solely on experiment design descriptions of 21 real-life multimedia quality subjective experiments.

cs.MM↗

Generalized Score Distribution

A class of discrete probability distributions contains distributions with limited support, i.e. possible argument values are limited to a set of numbers (typically consecutive). Examples of such data are results from subjective experiments utilizing the Absolute Category Rating (ACR) technique, where possible answers (argument values) are $\{1, 2, \cdots, 5\}$ or typical Likert scale $\{-3, -2, \cdots, 3\}$. An interesting subclass of those distributions are distributions limited to two parameters: describing the mean value and the spread of the answers, and having no more than one change in the probability monotonicity. In this paper we propose a general distribution passing those limitations called Generalized Score Distribution (GSD). The proposed GSD covers all spreads of the answers, from very small, given by the Bernoulli distribution, to the maximum given by a Beta Binomial distribution. We also show that GSD correctly describes subjective experiments scores from video quality evaluations with probability of 99.7\%. A Google Collaboratory website with implementation of the GSD estimation, simulation, and visualization is provided.

stat.ME↗

Intermediate efficiency of some weighted goodness-of-fit statistics

This paper compares the Anderson-Darling and some Eicker-Jaeschke statistics to the classical unweighted Kolmogorov-Smirnov statistic. The goal is to provide a quantitative comparison of such tests and to study real possibilities of using them to detect departures from the hypothesized distribution that occur in the tails. This contribution covers the case when under the alternative a moderately large portion of probability mass is allocated towards the tails. It is demonstrated that the approach allows for tractable, analytic comparison between the given test and the benchmark, and for reliable quantitative evaluation of weighted statistics. Finite sample results illustrate the proposed approach and confirm the theoretical findings. In the course of the investigation we also prove that a slight and natural modification of the solution proposed by Borovkov and Sycheva (1968) leads to a statistic which is a member of Eicker-Jaeschke class and can be considered an attractive competitor of the very popular supremum-type Anderson-Darling statistic.

math.ST↗

Multiresolution analysis and adaptive estimation on a sphere using stereographic wavelets

We construct an adaptive estimator of a density function on $d$ dimensional unit sphere $S^d$ ($d \geq 2 $), using a new type of spherical frames. The frames, or as we call them, stereografic wavelets are obtained by transforming a wavelet system, namely Daubechies, using some stereographic operators. We prove that our estimator achieves an optimal rate of convergence on some Besov type class of functions by adapting to unknown smoothness. Our new construction of stereografic wavelet system gives us a multiresolution approximation of $L^2(S^d)$ which can be used in many approximation and estimation problems. In this paper we also demonstrate how to implement the density estimator in $S^2$ and we present a finite sample behavior of that estimator in a numerical experiment.

math.ST↗

The smoothness test for a density function

The problem of testing hypothesis that a density function has no more than $μ$ derivatives versus it has more than $μ$ derivatives is considered. For a solution, the $L^2$ norms of wavelet orthogonal projections on some orthogonal "differences" of spaces from a multiresolution analysis is used. For the construction of the smoothness test an asymptotic distribution of a smoothness estimator is used. To analyze that asymptotic distribution, a new technique of enrichment procedure is proposed. The finite sample behaviour of the smoothness test is demonstrated in a numerical experiment in case of determination if a density function is continues or discontinues.

math.ST↗