SearcharxivSearch

arXiv subjects

Jiahua Chen

Publications and source records attributed to Jiahua Chen.

At least 19 recordsLinked to original sources

Byzantine-tolerant distributed learning of finite mixture models under partial corruptions

Finite mixture models characterize heterogeneous populations and are increasingly fitted to distributed data using split-and-conquer procedures that aggregate local mixture estimates at a central server. The aggregation step can be seriously compromised when transmitted local mixture estimates are partially or completely corrupted. To guard against Byzantine failures, existing robust aggregation methods have been developed for settings in which a local mixture estimate is either entirely authentic or entirely corrupted. Such methods can discard useful information when only some component estimates are corrupted. We consider component-wise Byzantine failure, in which some component estimates may be corrupted, but for each mixture component, a majority of the corresponding local estimates remain authentic. We propose component-wise filtered mixture reduction (CFMR), which selects a data-driven anchor, aligns transmitted components, filters each aligned cluster by a majority radius, and aggregates the retained estimates through mixture reduction. By filtering out only unreliable components, CFMR preserves authentic information from partially corrupted machines without requiring knowledge of the failure rates. We establish an adaptive convergence bound for CFMR and show that, under suitable conditions, it attains the oracle rate that would be achieved if the authentic component estimates were known in advance. Simulations and a real-data application show that CFMR remains close to the component-level oracle, whereas whole-machine filtering and unprotected aggregation can deteriorate substantially.

stat.ME

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning

Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing approaches typically mitigate this limitation either through computationally expensive post-training procedures on limited 3D datasets or through rigid tool-calling mechanisms that lack explicit geometric understanding and viewpoint flexibility. To address these challenges, we propose a \textit{training-free} framework that introduces a Visual Chain-of-Thought mechanism grounded in explicit 3D reconstruction. The proposed pipeline first reconstructs a high-fidelity 3D mesh from a single image using MLLM-guided keyword extraction and mask generation at multiple granularities. Subsequently, the framework leverages an external knowledge base to iteratively compute optimal camera extrinsic parameters and synthesize novel views, thereby emulating human perspective-taking. Extensive experiments demonstrate that the proposed approach significantly enhances spatial comprehension. Specifically, the framework outperforms specialized spatial models and general-purpose MLLMs, including \textit{GPT-5.2} and \textit{Gemini-2.5-Flash}, on major benchmarks such as 3DSRBench and Rel3D.

cs.CV

Bootstrap Consistency for Empirical Likelihood in Density Ratio Models

We establish the validity of bootstrap methods for empirical likelihood (EL) inference under the density ratio model (DRM). In particular, we prove that the bootstrap maximum EL estimators share the same limiting distribution as their population counterparts, both at the parameter level and for distribution functionals. Our results extend existing pointwise convergence theory to weak convergence of processes, which in turn justifies bootstrap inference for quantiles and dominance indices within the DRM framework. These theoretical guarantees close an important gap in the literature, providing rigorous foundations for resampling-based confidence intervals and hypothesis tests. Simulation studies further demonstrate the accuracy and practical value of the proposed approach.

math.ST

Byzantine-tolerant distributed learning of finite mixture models

Traditional statistical methods need to be updated to work with modern distributed data storage paradigms. A common approach is the split-and-conquer framework, which involves learning models on local machines and averaging their parameter estimates. However, this does not work for the important problem of learning finite mixture models, because subpopulation indices on each local machine may be arbitrarily permuted (the "label switching problem"). Zhang and Chen (2022) proposed Mixture Reduction (MR) to address this issue, but MR remains vulnerable to Byzantine failure, whereby a fraction of local machines may transmit arbitrarily erroneous information. This paper introduces Distance Filtered Mixture Reduction (DFMR), a Byzantine tolerant adaptation of MR that is both computationally efficient and statistically sound. DFMR leverages the densities of local estimates to construct a robust filtering mechanism. By analysing the pairwise L2 distances between local estimates, DFMR identifies and removes severely corrupted local estimates while retaining the majority of uncorrupted ones. We provide theoretical justification for DFMR, proving its optimal convergence rate and asymptotic equivalence to the global maximum likelihood estimate under standard assumptions. Numerical experiments on simulated and real-world data validate the effectiveness of DFMR in achieving robust and accurate aggregation in the presence of Byzantine failure.

stat.ME

Gaussian Mixture Reduction with Composite Transportation Divergence

Gaussian mixtures are widely used for approximating density functions in various applications such as density estimation, belief propagation, and Bayesian filtering. These applications often utilize Gaussian mixtures as initial approximations that are updated recursively. A key challenge in these recursive processes stems from the exponential increase in the mixture's order, resulting in intractable inference. To overcome the difficulty, the Gaussian mixture reduction (GMR), which approximates a high order Gaussian mixture by one with a lower order, can be used. Although existing clustering-based methods are known for their satisfactory performance and computational efficiency, their convergence properties and optimal targets remain unknown. In this paper, we propose a novel optimization-based GMR method based on composite transportation divergence (CTD). We develop a majorization-minimization algorithm for computing the reduced mixture and establish its theoretical convergence under general conditions. Furthermore, we demonstrate that many existing clustering-based methods are special cases of ours, effectively bridging the gap between optimization-based and clustering-based techniques. Our unified framework empowers users to select the most appropriate cost function in CTD to achieve superior performance in their specific applications. Through extensive empirical experiments, we demonstrate the efficiency and effectiveness of our proposed method, showcasing its potential in various domains.

stat.ML

Optimal Estimation under a Semiparametric Density Ratio Model

In many statistical and econometric applications, we gather individual samples from various interconnected populations that undeniably exhibit common latent structures. Utilizing a model that incorporates these latent structures for such data enhances the efficiency of inferences. Recently, many researchers have been adopting the semiparametric density ratio model (DRM) to address the presence of latent structures. The DRM enables estimation of each population distribution using pooled data, resulting in statistically more efficient estimations in contrast to nonparametric methods that analyze each sample in isolation. In this article, we investigate the limit of the efficiency improvement attainable through the DRM. We focus on situations where one population's sample size significantly exceeds those of the other populations. In such scenarios, we demonstrate that the DRM-based inferences for populations with smaller sample sizes achieve the highest attainable asymptotic efficiency as if a parametric model is assumed. The estimands we consider include the model parameters, distribution functions, and quantiles. We use simulation experiments to support the theoretical findings with a specific focus on quantile estimation. Additionally, we provide an analysis of real revenue data from U.S. collegiate sports to illustrate the efficacy of our contribution.

stat.ME

Global Consistency of Empirical Likelihood

This paper develops several interesting, significant, and interconnected approaches to nonparametric or semi-parametric statistical inferences. The overwhelmingly favoured maximum likelihood estimator (MLE) under parametric model is renowned for its strong consistency and optimality generally credited to Cramer. These properties, however, falter when the model is not regular or not completely accurate. In addition, their applicability is limited to local maxima close to the unknown true parameter value. One must therefore ascertain that the global maximum of the likelihood is strongly consistent under generic conditions (Wald, 1949). Global consistency is also a vital research problem in the context of empirical likelihood (Owen, 2001). The EL is a ground-breaking platform for nonparametric statistical inference. A subsequent milestone is achieved by placing estimating functions under the EL umbrella (Qin and Lawless, 1994). The resulting profile EL function possesses many nice properties of parametric likelihood but also shares the same shortcomings. These properties cannot be utilized unless we know the local maximum at hand is close to the unknown true parameter value. To overcome this obstacle, we first put forward a clean set of conditions under which the global maximum is consistent. We then develop a global maximum test to ascertain if the local maximum at hand is in fact a global maximum. Furthermore, we invent a global maximum remedy to ensure global consistency by expanding the set of estimating functions under EL. Our simulation experiments on many examples from the literature firmly establish that the proposed approaches work as predicted. Our approaches also provide superior solutions to problems of their parametric counterparts investigated by DeHaan (1981), Veall (1991), and Gan and Jiang (1999).

math.ST

Zero-Shot Aspect-Based Sentiment Analysis

Aspect-based sentiment analysis (ABSA) typically requires in-domain annotated data for supervised training/fine-tuning. It is a big challenge to scale ABSA to a large number of new domains. This paper aims to train a unified model that can perform zero-shot ABSA without using any annotated data for a new domain. We propose a method called contrastive post-training on review Natural Language Inference (CORN). Later ABSA tasks can be cast into NLI for zero-shot transfer. We evaluate CORN on ABSA tasks, ranging from aspect extraction (AE), aspect sentiment classification (ASC), to end-to-end aspect-based sentiment analysis (E2E ABSA), which show ABSA can be conducted without any human annotated ABSA data.

cs.CL

Empirical likelihood ratio test on quantiles under a density ratio model

Population quantiles are important parameters in many applications. Enthusiasm for the development of effective statistical inference procedures for quantiles and their functions has been high for the past decade. In this article, we study inference methods for quantiles when multiple samples from linked populations are available. The research problems we consider have a wide range of applications. For example, to study the evolution of the economic status of a country, economists monitor changes in the quantiles of annual household incomes, based on multiple survey datasets collected annually. Even with multiple samples, a routine approach would estimate the quantiles of different populations separately. Such approaches ignore the fact that these populations are linked and share some intrinsic latent structure. Recently, many researchers have advocated the use of the density ratio model (DRM) to account for this latent structure and have developed more efficient procedures based on pooled data. The nonparametric empirical likelihood (EL) is subsequently employed. Interestingly, there has been no discussion in this context of the EL-based likelihood ratio test (ELRT) for population quantiles. We explore the use of the ELRT for hypotheses concerning quantiles and confidence regions under the DRM. We show that the ELRT statistic has a chi-square limiting distribution under the null hypothesis. Simulation experiments show that the chi-square distributions approximate the finite-sample distributions well and lead to accurate tests and confidence regions. The DRM helps to improve statistical efficiency. We also give a real-data example to illustrate the efficiency of the proposed method.

math.ST

Distributed Learning of Finite Gaussian Mixtures

Advances in information technology have led to extremely large datasets that are often kept in different storage centers. Existing statistical methods must be adapted to overcome the resulting computational obstacles while retaining statistical validity and efficiency. Split-and-conquer approaches have been applied in many areas, including quantile processes, regression analysis, principal eigenspaces, and exponential families. We study split-and-conquer approaches for the distributed learning of finite Gaussian mixtures. We recommend a reduction strategy and develop an effective MM algorithm. The new estimator is shown to be consistent and retains root-n consistency under some general conditions. Experiments based on simulated and real-world data show that the proposed split-and-conquer approach has comparable statistical performance with the global estimator based on the full dataset, if the latter is feasible. It can even slightly outperform the global estimator if the model assumption does not match the real-world data. It also has better statistical and computational performance than some existing methods.

stat.ME

A Knowledge-Driven Approach to Classifying Object and Attribute Coreferences in Opinion Mining

Classifying and resolving coreferences of objects (e.g., product names) and attributes (e.g., product aspects) in opinionated reviews is crucial for improving the opinion mining performance. However, the task is challenging as one often needs to consider domain-specific knowledge (e.g., iPad is a tablet and has aspect resolution) to identify coreferences in opinionated reviews. Also, compiling a handcrafted and curated domain-specific knowledge base for each domain is very time consuming and arduous. This paper proposes an approach to automatically mine and leverage domain-specific knowledge for classifying objects and attribute coreferences. The approach extracts domain-specific knowledge from unlabeled review data and trains a knowledgeaware neural coreference classification model to leverage (useful) domain knowledge together with general commonsense knowledge for the task. Experimental evaluation on realworld datasets involving five domains (product types) shows the effectiveness of the approach.

cs.CL

Minimum Wasserstein Distance Estimator under Finite Location-scale Mixtures

When a population exhibits heterogeneity, we often model it via a finite mixture: decompose it into several different but homogeneous subpopulations. Contemporary practice favors learning the mixtures by maximizing the likelihood for statistical efficiency and the convenient EM-algorithm for numerical computation. Yet the maximum likelihood estimate (MLE) is not well defined for the most widely used finite normal mixture in particular and for finite location-scale mixture in general. We hence investigate feasible alternatives to MLE such as minimum distance estimators. Recently, the Wasserstein distance has drawn increased attention in the machine learning community. It has intuitive geometric interpretation and is successfully employed in many new applications. Do we gain anything by learning finite location-scale mixtures via a minimum Wasserstein distance estimator (MWDE)? This paper investigates this possibility in several respects. We find that the MWDE is consistent and derive a numerical solution under finite location-scale mixtures. We study its robustness against outliers and mild model mis-specifications. Our moderate scaled simulation study shows the MWDE suffers some efficiency loss against a penalized version of MLE in general without noticeable gain in robustness. We reaffirm the general superiority of the likelihood based learning strategies even for the non-regular finite location-scale mixtures.

stat.ML

Density ratio model with data-adaptive basis function

In many applications, we collect independent samples from interconnected populations. These population distributions share some latent structure, so it is advantageous to jointly analyze the samples. One effective way to connect the distributions is the semiparametric density ratio model (DRM). A key ingredient in the DRM is that the log density ratios are linear combinations of prespecified functions; the vector formed by these functions is called the basis function. A sensible basis function can often be chosen based on knowledge of the context, and DRM-based inference is effective even if the basis function is imperfect. However, a data-adaptive approach to the choice of basis function remains an interesting and important research problem. We propose an approach based on the classical functional principal component analysis (FPCA). Under some conditions, we show that this approach leads to consistent basis function estimation. Our simulation results show that the proposed adaptive choice leads to an efficiency gain. We use a real-data example to demonstrate the efficiency gain and the ease of our approach.

stat.ME

Consistency of the MLE under a two-parameter gamma mixture model with a structural shape parameter

The finite Gamma mixture model is often used to describe randomness in income data, insurance data, and data from other applications. The popular likelihood approach, however, does not work for this model because the likelihood function is unbounded, and the maximum likelihood estimator is therefore not well defined. There has been much research into ways to ensure the consistent estimation of the mixing distribution, including placing an upper bound on the shape parameter or adding a penalty to the log-likelihood function. In this paper, we show that if the shape parameter in the finite Gamma mixture model is structural, then the maximum likelihood estimator of the mixing distribution is well defined and strongly consistent. We also present simulation results demonstrating the consistency of the estimator. We illustrate the application of the model with a structural scale parameter to household income data. The fitted mixture distribution leads to several possible subpopulation structures in terms of the level of disposable income.

math.ST

Permutation tests under a rotating sampling plan with clustered data

Consider a population consisting of clusters of sampling units, evolving temporally, spatially, or according to other dynamics. We wish to monitor the evolution of its means, medians, or other parameters. For administrative convenience and informativeness, clustered data are often collected via a rotating plan. Under rotating plans, the observations in the same clusters are correlated, and observations on the same unit collected on different occasions are also correlated. Ignoring this correlation structure may lead to invalid inference procedures. Accommodating cluster structure in parametric models is difficult or will have a high level of misspecification risk. In this paper, we explore exchangeability in clustered data collected via a rotating sampling plan to develop a permutation scheme for testing various hypotheses of interest. We also introduce a semiparametric density ratio model to facilitate the multiple population structure in rotating sampling plans. The combination ensures the validity of the inference methods while extracting maximum information from the sampling plan. A simulation study indicates that the proposed tests firmly control the type I error whether or not the data are clustered. The use of the density ratio model improves the power of the tests.

stat.ME

Minority Voter Distributions and Partisan Gerrymandering

Many people believe that it is disadvantageous for members aligning with a minority party to cluster in cities, as this makes it easier for the majority party to gerrymander district boundaries to diminish the representation of the minority. We examine this effect by exhaustively computing the average representation for every possible $5\times 5$ grid of population placement and district boundaries. We show that, in fact, it is advantageous for the minority to arrange themselves in clusters, as it is positively correlated with representation. We extend this result to more general cases by considering the dual graph of districts, and we also propose and analyze metaheuristic algorithms that allow us to find strong lower bounds for maximum expected representation.

cs.CY

Test for homogeneity with unordered paired observations

In some applications, an experimental unit is composed of two distinct but related subunits. The response from such a unit is $(X_{1}, X_{2})$ but we observe only $Y_1 = \min\{X_{1},X_{2}\}$ and $Y_2 = \max\{X_{1},X_{2}\}$, i.e., the subunit identities are not observed. We call $(Y_1, Y_2)$ unordered paired observations. Based on unordered paired observations $\{(Y_{1i}, Y_{2i})\}_{i=1}^n$, we are interested in whether the marginal distributions for $X_1$ and $X_2$ are identical. Testing methods are available in the literature under the assumptions that $Var(X_1) = Var(X_2)$ and $Cov(X_1, X_2) = 0$. However, by extensive simulation studies, we observe that when one or both assumptions are violated, these methods have inflated type I errors or much lower powers. In this paper, we study the likelihood ratio test statistics for various scenarios and explore their limiting distributions without these restrictive assumptions. Furthermore, we develop Bartlett correction formulae for these statistics to enhance their precision when the sample size is not large. Simulation studies and real-data examples are used to illustrate the efficacy of the proposed methods.

math.ST

Homogeneity testing under finite location-scale mixtures

The testing problem for the order of finite mixture models has a long history and remains an active research topic. Since Ghosh and Sen (1985) revealed the hard-to-manage asymptotic properties of the likelihood ratio test, there has been marked progress. The most successful attempts include the modified likelihood ratio test and the EM-test, which lead to neat solutions for finite mixtures of univariate normal distributions, finite mixtures of single-parameter distributions, and several mixture-like models. The problem remains challenging, and there is still no generic solution for location-scale mixtures. In this paper, we provide an EM-test solution for homogeneity for finite mixtures of location-scale family distributions. This EM-test has nonstandard limiting distributions, but we are able to find the critical values numerically. We use computer experiments to obtain appropriate values for the tuning parameters. A simulation study shows that the fine-tuned EM-test has close to nominal type I errors and very good power properties. Two application examples are included to demonstrate the performance of the EM-test.

stat.ME