Statistical inference with win statistics in cluster-randomized trials with hierarchical composite outcomes
Win statistics have become increasingly popular for analyzing hierarchical composite endpoints in clinical trials. The win ratio, win odds, net benefit, and desirability of outcome ranking (DOOR) share a pairwise-comparison framework and provide complementary summaries of treatment benefit. Despite recent progress in individually randomized trials, statistical inference for these measures in cluster-randomized trials (CRTs) remains underdeveloped. We provide a unified development and comparison of six testing procedures for all four win measures in parallel-arm CRTs: three Wald tests, two randomization-based tests, and a jackknife empirical likelihood ratio test. Through simulations with hierarchical semi-competing risks outcomes, we evaluate type I error rate and power across varying design parameters. With 20 clusters, the randomization-based procedures provided the most stable type I error rate control, while the clustered rank-sum Wald test with a t-distribution was the most reliable among the Wald tests, occasionally carrying a conservative test size. For the win ratio and win odds, the permutation test also yielded higher power than the clustered rank-sum Wald test. With 100 clusters, differences in type I error rate and power were small. These findings favor permutation inference in small CRTs, with the clustered rank-sum Wald test providing an analytic alternative, particularly for net benefit and DOOR. On the other hand, the analytic Wald procedures offer a computationally convenient choice for CRTs with a large number of clusters. We illustrate the methods by reanalyzing the STRIDE trial and implement all procedures in the WinsCRT R package.