Searcharxiv⌕ Search

arXiv subjects

Hiroshi Yadohisa

Publications and source records attributed to Hiroshi Yadohisa.

12 recordsLinked to original sources

Principal component-guided sparse reduced-rank regression

Reduced-rank regression estimates regression coefficients by imposing a low-rank constraint on the matrix of regression coefficients, thereby accounting for correlations among response variables. To further improve predictive accuracy and model interpretability, several regularized reduced-rank regression methods have been proposed. However, these existing methods cannot bias the regression coefficients toward the leading principal component directions while accounting for the correlation structure among explanatory variables. In addition, when the explanatory variables exhibit a group structure, the correlation structure within each group cannot be adequately incorporated. To overcome these limitations, we propose a new method that introduces pcLasso into the reduced-rank regression framework. The proposed method improves predictive accuracy by accounting for the correlation among response variables while strongly biasing the matrix of regression coefficients toward principal component directions with large variance. Furthermore, even in settings where the explanatory variables possess a group structure, the proposed method is capable of explicitly incorporating this structure into the estimation process. Finally, we illustrate the effectiveness of the proposed method through numerical simulations and real data application.

stat.ME↗

Regularized Sparse Optimal Discriminant Clustering

We propose a new method based on sparse optimal discriminant clustering (SODC), incorporating a penalty term into the scoring matrix based on convex clustering. With the addition of this penalty term, it is expected to improve the accuracy of cluster identification by pulling points within the same cluster closer together and points from different clusters further apart. When the estimation results are visualized, the clustering structure can be depicted more clearly. Moreover, we develop a novel algorithm to derive the updated formula of this scoring matrix using a majorizing function. The scoring matrix is updated using the alternating direction method of multipliers (ADMM), which is often employed to calculate the parameters of the objective function in the convex clustering. In the proposed method, as in the conventional SODC, the scoring matrix is subject to an orthogonal constraint. Therefore, it is necessary to satisfy the orthogonal constraint on the scoring matrix while maintaining the clustering structure. Using a majorizing function, we adress the challenge of enforcing both orthogonal constraint and the clustering structure within the scoring matrix. We demonstrate numerical simulations and an application to real data to assess the performance of the proposed method.

stat.ME↗

Multinomial probit model based on joint quantile regression

The multinomial probit model is a typical statistical model for multiple-choice data applied in many research areas. When we are interested in some quantiles of relative utilities for understanding the distribution of these utilities, the multinomial probit model is unsuitable because we only interpret the expectation of relative utilities based on it. We thus propose quantile regression analysis methods for multinomial choice data based on joint quantile regression and multinomial probit models to compare relative utilities with some quantiles. Using a joint quantile regression model allows us to consider the conditional quantile points of relative utilities and explicitly describe the correlation structure in the latent variables. We derive the full conditional distribution under several prior distributions and estimate the model's parameters from the posterior distribution by Gibbs sampling. The ability to calculate by Gibbs sampling is not only computationally less expensive than the Metropolis--Hastings method, but also easier to implement. We also apply the proposed model to several datasets. Consequently, we obtain interpretable results about different parameters by quantile.

stat.ME↗

Wilcoxon-type Multivariate Cluster Elastic Net

We propose a method for high dimensional multivariate regression that is robust to random error distributions that are heavy-tailed or contain outliers, while preserving estimation accuracy in normal random error distributions. We extend the Wilcoxon-type regression to a multivariate regression model as a tuning-free approach to robustness. Furthermore, the proposed method regularizes the L1 and L2 terms of the clustering based on k-means, which is extended from the multivariate cluster elastic net. The estimation of the regression coefficient and variable selection are produced simultaneously. Moreover, considering the relationship among the correlation of response variables through the clustering is expected to improve the estimation performance. Numerical simulation demonstrates that our proposed method overperformed the multivariate cluster method and other methods of multiple regression in the case of heavy-tailed error distribution and outliers. It also showed stability in normal error distribution. Finally, we confirm the efficacy of our proposed method using a data example for the gene associated with breast cancer.

stat.ME↗

Estimating heterogeneous treatment effects by W-MCM based on Robust reduced rank regression

Recently, from the personalized medicine perspective, there has been an increased demand to identify subgroups of subjects for whom treatment is effective. Consequently, the estimation of heterogeneous treatment effects (HTE) has been attracting attention. While various estimation methods have been developed for a single outcome, there are still limited approaches for estimating HTE for multiple outcomes. Accurately estimating HTE remains a challenge especially for datasets where there is a high correlation between outcomes or the presence of outliers. Therefore, this study proposes a method that uses a robust reduced-rank regression framework to estimate treatment effects and identify effective subgroups. This approach allows the consideration of correlations between treatment effects and the estimation of treatment effects with an accurate low-rank structure. It also provides robust estimates for outliers. This study demonstrates that, when treatment effects are estimated using the reduced rank regression framework with an appropriate rank, the expected value of the estimator equals the treatment effect. Finally, we illustrate the effectiveness and interpretability of the proposed method through simulations and real data examples.

stat.ME↗

Extension of W-method and A-learner for multiple binary outcomes

In this study, we compared two groups, in which subjects were assigned to either the treatment or the control group. In such trials, if the efficacy of the treatment cannot be demonstrated in a population that meets the eligibility criteria, identifying the subgroups for which the treatment is effective is desirable. Such subgroups can be identified by estimating heterogeneous treatment effects (HTE). In recent years, methods for estimating HTE have increasingly relied on complex models. Although these models improve the estimation accuracy, they often sacrifice interpretability. Despite significant advancements in the methods for continuous or univariate binary outcomes, methods for multiple binary outcomes are less prevalent, and existing interpretable methods, such as the W-method and A-learner, while capable of estimating HTE for a single binary outcome, still fail to capture the correlation structure when applied to multiple binary outcomes. We thus propose two methods for estimating HTE for multiple binary outcomes: one based on the W-method and the other based on the A-learner. We also demonstrate that the conventional A-learner introduces bias in the estimation of the treatment effect. The proposed method employs a framework based on reduced-rank regression to capture the correlation structure among multiple binary outcomes. We correct for the bias inherent in the A-learner estimates and investigate the impact of this bias through numerical simulations. Finally, we demonstrate the effectiveness of the proposed method using a real data application.

stat.ME↗

Quantile Outcome Adaptive Lasso: Covariate Selection for Inverse Probability Weighting Estimator of Quantile Treatment Effects

When using the propensity score method to estimate the treatment effects, it is important to select the covariates to be included in the propensity score model. The inclusion of covariates unrelated to the outcome in the propensity score model led to bias and large variance in the estimator of treatment effects. Many data-driven covariate selection methods have been proposed for selecting covariates related to outcomes. However, most of them assume an average treatment effect estimation and may not be designed to estimate quantile treatment effects (QTE), which is the effect of treatment on the quantiles of outcome distribution. In QTE estimation, we consider two relation types with the outcome as the expected value and quantile point. To achieve this, we propose a data-driven covariate selection method for propensity score models that allows for the selection of covariates related to the expected value and quantile of the outcome for QTE estimation. Assuming the quantile regression model as an outcome regression model, covariate selection was performed using a regularization method with the partial regression coefficients of the quantile regression model as weights. The proposed method was applied to artificial data and a dataset of mothers and children born in King County, Washington, to compare the performance of existing methods and QTE estimators. As a result, the proposed method performs well in the presence of covariates related to both the expected value and quantile of the outcome.

stat.ME↗

Bayesian Geographically Weighted Regression using Fused Lasso Prior

A main purpose of spatial data analysis is to predict the objective variable for the unobserved locations. Although Geographically Weighted Regression (GWR) is often used for this purpose, estimation instability proves to be an issue. To address this issue, Bayesian Geographically Weighted Regression (BGWR) has been proposed. In BGWR, by setting the same prior distribution for all locations, the coefficients' estimation stability is improved. However, when observation locations' density is spatially different, these methods do not sufficiently consider the similarity of coefficients among locations. Moreover, the prediction accuracy of these methods becomes worse. To solve these issues, we propose Bayesian Geographically Weighted Sparse Regression (BGWSR) that uses Bayesian Fused Lasso for the prior distribution of the BGWR coefficients. Constraining the parameters to have the same values at adjacent locations is expected to improve the prediction accuracy at locations with a low number of adjacent locations. Furthermore, from the predictive distribution, it is also possible to evaluate the uncertainty of the predicted value of the objective variable. By examining numerical studies, we confirmed that BGWSR has better prediction performance than the existing methods (GWR and BGWR) when the density of observation locations is spatial difference. Finally, the BGWSR is applied to land price data in Tokyo. Thus, the results suggest that BGWSR has better prediction performance and smaller uncertainty than existing methods.

stat.ME↗

Causal rule ensemble method for estimating heterogeneous treatment effect with consideration of main effects

This study proposes a novel framework based on the RuleFit method to estimate Heterogeneous Treatment Effect (HTE) in a randomized clinical trial. To achieve this, we adopted S-learner of the metaalgorithm for our proposed framework. The proposed method incorporates a rule term for the main effect and treatment effect, which allows HTE to be interpretable form of rule. By including a main effect term in the proposed model, the selected rule is represented as an HTE that excludes other effects. We confirmed a performance equivalent to that of another ensemble learning methods through numerical simulation and demonstrated the interpretation of the proposed method from a real data application.

stat.ME↗

Estimation and visualization of treatment effects for multiple outcomes

We consider a randomized controlled trial between two groups. The objective is to identify a population with characteristics such that the test therapy is more effective than the control therapy. Such a population is called a subgroup. This identification can be made by estimating the treatment effect and identifying interactions between treatments and covariates. To date, many methods have been proposed to identify subgroups for a single outcome. There are also multiple outcomes, but they are difficult to interpret and cannot be applied to outcomes other than continuous values. In this paper, we propose a multivariate regression method that introduces latent variables to estimate the treatment effect on multiple outcomes simultaneously. The proposed method introduces latent variables and adds Lasso sparsity constraints to the estimated loadings to facilitate the interpretation of the relationship between outcomes and covariates. The framework of the generalized linear model makes it applicable to various types of outcomes. Interpretation of subgroups is made by visualizing treatment effects and latent variables. This allows us to identify subgroups with characteristics that make the test therapy more effective for multiple outcomes. Simulation and real data examples demonstrate the effectiveness of the proposed method.

stat.ME↗

F-measure Maximizing Logistic Regression

Logistic regression is a widely used method in several fields. When applying logistic regression to imbalanced data, for which majority classes dominate over minority classes, all class labels are estimated as `majority class.' In this article, we use an F-measure optimization method to improve the performance of logistic regression applied to imbalanced data. While many F-measure optimization methods adopt a ratio of the estimators to approximate the F-measure, the ratio of the estimators tends to have more bias than when the ratio is directly approximated. Therefore, we employ an approximate F-measure for estimating the relative density ratio. In addition, we define a relative F-measure and approximate the relative F-measure. We show an algorithm for a logistic regression weighted approximated relative to the F-measure. The experimental results using real world data demonstrated that our proposed method is an efficient algorithm to improve the performance of logistic regression applied to imbalanced data.

stat.ME↗

Bi-clustering for time-varying relational count data analysis

Relational count data are often obtained from sources such as simultaneous purchase in online shops and social networking service information. Bi-clustering such relational count data reveals the latent structure of the relationship between objects such as household items or people. When relational count data observed at multiple time points are available, it is worthwhile incorporating the time structure into the bi-clustering result to understand how objects move between the cluster over time. In this paper, we propose two bi-clustering methods for analyzing time-varying relational count data. The first model, the dynamic Poisson infinite relational model (dPIRM), handles time-varying relational count data. In the second model, which we call the dynamic zero-inflated Poisson infinite relational model, we further extend the dPIRM so that it can handle zero-inflated data. Proposing both two models is important as zero-inflated data are often encountered, especially when the time intervals are short. In addition, by explicitly deriving the relevant full conditional distributions, we describe the features of the estimated parameters and, in turn, the relationship between the two models. We show the effectiveness of both models through a simulation study and a real data example.

stat.ME↗