SearcharxivSearch

arXiv subjects

Kouji Tahata

Publications and source records attributed to Kouji Tahata.

13 recordsLinked to original sources

Divergence-based Robust Generalised Bayesian Inference for Directional Data via von Mises-Fisher models

This paper focusses on robust estimation of location and concentration parameters of the von Mises-Fisher distribution in the Bayesian framework. The von Mises-Fisher (or Langevin) distribution has played a central role in directional statistics. Directional data have been investigated for many decades, and more recently, they have gained increasing attention in diverse areas such as bioinformatics and text data analysis. Although outliers can significantly affect the estimation results even for directional data, the treatment of outliers remains an unresolved and challenging problem. In the frequentist framework, numerous studies have developed robust estimation methods for directional data with outliers, but, in contrast, only a few robust estimation methods have been proposed in the Bayesian framework. In this paper, we propose Bayesian inference based on the density power divergence and the $\gamma$-divergence and establish their asymptotic properties and robustness. In addition, the Bayesian approach naturally provides a way to assess estimation uncertainty through the posterior distribution, which is particularly useful for small samples. Furthermore, to carry out the posterior computation, we develop the posterior computation algorithm based on the weighted Bayesian bootstrap for estimating parameters. The effectiveness of the proposed methods is demonstrated through simulation studies. Using two real datasets, we further show that the proposed method provides reliable and robust estimation even in the presence of outliers or data contamination.

stat.ME

Quasi-symmetry and geometric marginal homogeneity: A simplicial approach to square contingency tables

Square contingency tables are traditionally analyzed with a focus on the symmetric structure of the corresponding probability tables. We view probability tables as elements of a simplex equipped with the Aitchison geometry. This perspective allows us to present a novel approach to analyzing symmetric structure using a compositionally coherent framework. We present a geometric interpretation of quasi-symmetry as an e-flat subspace and introduce a new concept called geometric marginal homogeneity, which is also characterized as an e-flat structure. We prove that both quasi-symmetric tables and geometric marginal homogeneous tables form subspaces in the simplex, and demonstrate that the measure of skew-symmetry in Aitchison geometry can be orthogonally decomposed into measures of departure from quasi-symmetry and geometric marginal homogeneity. We illustrate the application and effectiveness of our proposed methodology using data on unaided distance vision from a sample of women.

math.ST

Association measures for two-way contingency tables based on multi-categorical proportional reduction in error

In two-way contingency tables under an asymmetric situation, where the row and column variables are defined as explanatory and response variables, respectively, quantifying the extent to which the explanatory variable contributes to predicting the response variable is important. One quantification method is the association measure, which indicates the degree of association in a range from $0$ to $1$. Among various measures that have been proposed, those based on proportional reduction in error (PRE) are particularly notable for their simplicity and intuitive interpretation. These measures, including Goodman-Kruskal's lambda proposed in 1954, are widely implemented in statistical software such as R and SAS and remain extensively used. However, a well-known limitation of PRE measures is their potential to return a value of $0$ despite no independence. This issue arises because the measures are constructed based solely on the maximum joint and marginal probabilities, failing to make full use of the information available in the contingency table. To address this problem, we propose an extension of PRE measures designed for the proportional reduction in error with multiple categories. The properties of the proposed measures are examined, and their utility is demonstrated through numerical experiments. The results suggest their potential as practical tools in applied statistics.

stat.ME

A measure of departure from symmetry via the Fisher-Rao distance for contingency tables

A measure of asymmetry is a quantification method that allows for the comparison of categorical evaluations before and after treatment effects or among different target populations, irrespective of sample size. We focus on square contingency tables that summarize survey results between two time points or cohorts, represented by the same categorical variables. We propose a measure to evaluate the degree of departure from a symmetry model using cosine similarity. This proposal is based on the Fisher-Rao distance, allowing asymmetry to be interpreted as a geodesic distance between two distributions. Various measures of asymmetry have been proposed, but visualizing the relationship of these quantification methods on a two-dimensional plane demonstrates that the proposed measure provides the geometrically simplest and most natural quantification. Moreover, the visualized figure indicates that the proposed method for measuring departures from symmetry is less affected by very few cells with extreme asymmetry. A simulation study shows that for square contingency tables with an underlying asymmetry model, our method can directly extract and quantify only the asymmetric structure of the model, and can more sensitively detect departures from symmetry than divergence-type measures.

stat.ME

Visualization for departures from symmetry with the power-divergence-type measure in two-way contingency tables

When the row and column variables consist of the same category in a two-way contingency table, it is specifically called a square contingency table. Since it is clear that the square contingency tables have an association structure, a primary objective is to examine symmetric relationships and transitions between variables. While various models and measures have been proposed to analyze these structures understanding changes between two variables in behavior at two-time points or cohorts, it is also necessary to require a detailed investigation of individual categories and their interrelationships, such as shifts in brand preferences. This paper proposes a novel approach to correspondence analysis (CA) for evaluating departures from symmetry in square contingency tables with nominal categories, using a power-divergence-type measure. The approach ensures that well-known divergences can also be visualized and, regardless of the divergence used, the CA plot consists of two principal axes with equal contribution rates. Additionally, the scaling is independent of sample size, making it well-suited for comparing departures from symmetry across multiple contingency tables. Confidence regions are also constructed to enhance the accuracy of the CA plot.

stat.ME

Optimal Bayesian predictive probability for delayed response in single-arm clinical trials with binary efficacy outcome

In oncology, phase II or multiple expansion cohort trials are crucial for clinical development plans. This is because they aid in identifying potent agents with sufficient activity to continue development and confirm the proof of concept. Typically, these clinical trials are single-arm trials, with the primary endpoint being short-term treatment efficacy. Despite the development of several well-designed methodologies, there may be a practical impediment in that the endpoints may be observed within a sufficient time such that adaptive go/no-go decisions can be made in a timely manner at each interim monitoring. Specifically, Response Evaluation Criteria in Solid Tumors guideline defines a confirmed response and necessitates it in non-randomized trials, where the response is the primary endpoint. However, obtaining the confirmed outcome from all participants entered at interim monitoring may be time-consuming as non-responders should be followed up until the disease progresses. Thus, this study proposed an approach to accelerate the decision-making process that incorporated the outcome without confirmation by discounting its contribution to the decision-making framework using the generalized Bayes' theorem. Further, the behavior of the proposed approach was evaluated through a simple simulation study. The results demonstrated that the proposed approach made appropriate interim go/no-go decisions.

stat.ME

Modeling asymmetry in multi-way contingency tables with ordinal categories via f-divergence

This study introduces a novel model that effectively captures asymmetric structures in multivariate contingency tables with ordinal categories. Leveraging the principle of maximum entropy, our approach employs f-divergence to provide a rational model under the presence of a ``prior guess.'' Inspired by the constraints used in the derivation of multivariate normal distributions, we demonstrate that the proposed model minimizes f-divergence from complete symmetry under specific constraints. The proposed model encompasses existing asymmetry models as special cases while offering remarkably high interpretability. By modifying divergence measures included in f-divergence, the model provides the flexibility to adapt to specific probabilistic structures of interest. Furthermore, we established theorems that show that a complete symmetry model can be decomposed into two or more models, each imposing less restrictive parameter constraints. We also investigated the properties of the goodness-of-fit statistics with an emphasis on the likelihood ratio and Wald test statistics. Extensive Monte Carlo simulations confirmed the nominal size, high power, and robustness of the choice of f-divergence. Finally, an application to real-world data highlights the practical utility of the proposed model for analyzing asymmetric structures in ordinal contingency tables.

stat.ME

A generalized ordinal quasi-symmetry model and its separability for analyzing multi-way tables

This paper addresses the challenge of modeling multi-way contingency tables for matched set data with ordinal categories. Although the complete symmetry and marginal homogeneity models are well established, they may not always provide a satisfactory fit to the data. To address this issue, we propose a generalized ordinal quasi-symmetry model that offers increased flexibility when the complete symmetry model fails to capture the underlying structure. We investigate the properties of this new model and provide an information-theoretic interpretation, elucidating its relationship to the ordinal quasi-symmetry model. Moreover, we revisit Agresti's findings and present a new necessary and sufficient condition for the complete symmetry model, proving that the proposed model and the marginal moment equality model are separable hypotheses. We demonstrate the practical application of our model through empirical studies on medical and public opinion datasets. Comprehensive simulation studies evaluate the proposed model under various scenarios, including model's performance for multivariate normal data and asymptotic behavior. It enables researchers to examine the symmetry structure in the data with greater precision, providing a more thorough understanding of the underlying patterns. This powerful framework equips researchers with the necessary tools to explore the complexities of ordinal variable relationships in matched data sets, paving the way for new discoveries and insights.

stat.ME

Bayesian predictive probability based on a bivariate index vector for single-arm phase II study with binary efficacy and safety endpoints

In oncology, phase II studies are crucial for clinical development plans as such studies identify potent agents with sufficient activity to continue development in the subsequent phase III trials. Traditionally, phase II studies are single-arm studies, with the primary endpoint being short-term treatment efficacy. However, drug safety is also an important consideration. In the context of such multiple-outcome designs, predictive probability-based Bayesian monitoring strategies have been developed to assess whether a clinical trial will provide enough evidence to continue with a phase III study at the scheduled end of the trial. Herein, we propose a new simple index vector for summarizing results that cannot be captured by existing strategies. Specifically, for each interim monitoring time point, we calculate the Bayesian predictive probability using our new index and use it to assign a go/no-go decision. Finally, simulation studies are performed to evaluate the operating characteristics of the design. The obtained results demonstrate that the proposed method makes appropriate interim go/no-go decisions.

stat.ME

Geometric Mean Type of Proportional Reduction in Variation Measure for Two-Way Contingency Tables

In a two-way contingency table analysis with explanatory and response variables, the analyst is interested in the independence of the two variables. However, if the test of independence does not show independence or clearly shows a relationship, the analyst is interested in the degree of their association. Various measures have been proposed to calculate the degree of their association, one of which is the proportional reduction in variation (PRV) measure which describes the PRV from the marginal distribution to the conditional distribution of the response. The conventional PRV measures can assess the association of the entire contingency table, but they can not accurately assess the association for each explanatory variable. In this paper, we propose a geometric mean type of PRV (geoPRV) measure that aims to sensitively capture the association of each explanatory variable to the response variable by using a geometric mean, and it enables analysis without underestimation when there is partial bias in cells of the contingency table. Furthermore, the geoPRV measure is constructed by using any functions that satisfy specific conditions, which has application advantages and makes it possible to express conventional PRV measures as geometric mean types in special cases.

stat.ME

Visualizing departures from marginal homogeneity for square contingency tables with ordered categories

Square contingency tables are a special case commonly used in various fields to analyze categorical data. Although several analysis methods have been developed to examine marginal homogeneity (MH) in these tables, existing measures are single-summary ones. To date, a visualization approach has yet to be proposed to intuitively depict the results of MH analysis. Current measures used to assess the degree of departure from MH are based on entropy such as the Kullback-Leibler divergence and do not satisfy distance postulates. Hence, the current measures are not conducive to visualization. Herein we present a measure utilizing the Matusita distance and introduce a visualization technique that employs sub-measures of categorical data. Through multiple examples, we demonstrate the meaningfulness of our visualization approach and validate its usefulness to provide insightful interpretations.

stat.ME

Diagonals-parameter symmetry model and its property for square contingency tables with ordinal categories

Previously, the diagonals-parameter symmetry model based on $f$-divergence (denoted by DPS[$f$]) was reported to be equivalent to the diagonals-parameter symmetry model regardless of the function $f$, but the proof was omitted. Here, we derive the DPS[$f$] model and the proof of the relation between the two models. We can obtain various interpretations of the diagonals-parameter symmetry model from the result. Additionally, the necessary and sufficient conditions for symmetry and property between test statistics for goodness of fit are discussed.

stat.ME

Test for symmetry in $2 \times 2$ contingency tables with nonignorable nonresponses

The McNemar test evaluates the hypothesis that two correlated proportion is common in $2 \times 2$ contingency tables with the same categories. This study discusses a test for symmetry in $2 \times 2$ contingency tables with nonignorable nonresponses. The proposed method is based on Takai and Kano (2008), which discusses a test for independence because a dependency assumption between the two observed outcomes is required to obtain an identification. Here, we focus on three models and propose a test for symmetry in $2 \times 2$ contingency tables with nonignorable nonresponses.

stat.ME