SearcharxivSearch

arXiv subjects

Qihua Wang

Publications and source records attributed to Qihua Wang.

15 recordsLinked to original sources

Performance of a SuperCDMS HVeV Detector with Sub-eV Energy Resolution and Single Charge-sensitivity

We present a detailed characterization of a new generation of athermal-phonon single-charge sensitive Si HVeV detectors, the best of which achieved 589~meV~$\pm$~5~meV baseline resolution. Our sub-eV energy resolution enables precise measurements of single-photon events and reveal consistent energy losses of 0.92~eV~$\pm$~0.02~eV per charge excitation across two facilities. We demonstrate that the noise for these detectors is well described using a standard Transition Edge Sensor noise model. We also measure a nominal phonon collection efficiency of 45\%~$\pm$~3\%~(stat.)~$\pm$~6\%~(syst.) in the best performing device, establishing these detectors as the most efficient athermal phonon detectors to date, limited only by intrinsic limitations of quasiparticle generation.

physics.ins-det

A Moment-assisted Approach for Improving Subsampling-based MLE with Large-scale data

The maximum likelihood estimation is computationally demanding for large datasets, particularly when the likelihood function includes integrals. Subsampling can reduce the computational burden, but it often results in efficiency loss.This paper proposes a moment-assisted subsampling (MAS) method that can improve the estimation efficiency of existing subsampling-based maximum likelihood estimators.The motivation behind this approach stems from the fact that sample moments can be efficiently computed even if the sample size of the whole data set is huge.Through the generalized method of moments, the proposed method incorporates informative sample moments of the whole data. The MAS estimator can be computed rapidly and is asymptotically normal with a smaller asymptotic variance than the corresponding estimator without incorporating sample moments of the whole data. The asymptotic variance of the proposed estimator depends on the specific sample moments incorporated. We derive the optimal moment that minimizes the resulting asymptotic variance in terms of Loewner order. The proposed MAS estimator can achieve the same estimation efficiency as the whole data-based estimator when the optimal moment is incorporated. Numerical results demonstrate the promising performance of the proposed method in both estimation and computational efficiency compared with existing subsampling methods.

stat.ME

A maximin optimal approach for sampling designs in two-phase studies

Data collection costs can vary widely across variables in data science tasks. Two-phase designs can be employed to save data collection costs. This paper considers the two-phase studies where inexpensive variables are collected for all subjects in the first phase, and expensive variables are measured for a subsample of subjects in the second phase based on a predetermined sampling rule. The estimation efficiency under two-phase designs relies heavily on the sampling rule. Existing literature primarily focuses on designing sampling rules for estimating a scalar parameter in some parametric models or specific estimating problems. However, real-world scenarios are usually model-unknown and involve two-phase designs for model-free estimation of a scalar or multi-dimensional parameter. This paper proposes a maximin criterion to design an optimal sampling rule based on semiparametric efficiency bounds. The proposed method is model-free and applicable to general estimating problems. The resulting sampling rule can minimize the semiparametric efficiency bound when the parameter is scalar and improve the bound for every component when the parameter is multi-dimensional. Simulation studies demonstrate that the proposed designs reduce the variance of the resulting estimator in various settings. The implementation of the proposed design is illustrated in a real data analysis.

stat.ME

Distributed Empirical Likelihood Inference With or Without Byzantine Failures

Empirical likelihood is a very important nonparametric approach which is of wide application. However, it is hard and even infeasible to calculate the empirical log-likelihood ratio statistic with massive data. The main challenge is the calculation of the Lagrange multiplier. This motivates us to develop a distributed empirical likelihood method by calculating the Lagrange multiplier in a multi-round distributed manner. It is shown that the distributed empirical log-likelihood ratio statistic is asymptotically standard chi-squared under some mild conditions. The proposed algorithm is communication-efficient and achieves the desired accuracy in a few rounds. Further, the distributed empirical likelihood method is extended to the case of Byzantine failures. A machine selection algorithm is developed to identify the worker machines without Byzantine failures such that the distributed empirical likelihood method can be applied. The proposed methods are evaluated by numerical simulations and illustrated with an analysis of airline on-time performance study and a surface climate analysis of Yangtze River Economic Belt.

stat.ME

Empirical Likelihood Inference over Decentralized Networks

As a nonparametric statistical inference approach, empirical likelihood has been found very useful in numerous occasions. However, it encounters serious computational challenges when applied directly to the modern massive dataset. This article studies empirical likelihood inference over decentralized distributed networks, where the data are locally collected and stored by different nodes. To fully utilize the data, this article fuses Lagrange multipliers calculated in different nodes by employing a penalization technique. The proposed distributed empirical log-likelihood ratio statistic with Lagrange multipliers solved by the penalized function is asymptotically standard chi-squared under regular conditions even for a divergent machine number. Nevertheless, the optimization problem with the fused penalty is still hard to solve in the decentralized distributed network. To address the problem, two alternating direction method of multipliers (ADMM) based algorithms are proposed, which both have simple node-based implementation schemes. Theoretically, this article establishes convergence properties for proposed algorithms, and further proves the linear convergence of the second algorithm in some specific network structures. The proposed methods are evaluated by numerical simulations and illustrated with analyses of census income and Ford gobike datasets.

stat.ME

Updatable Estimation in Generalized Linear Models with Missing Response

This paper develops an updatable inverse probability weighting (UIPW) estimation for the generalized linear models with response missing at random in streaming data sets. A two-step online updating algorithm is provided for the proposed method. In the first step we construct an updatable estimator for the parameter in propensity function and hence obtain an updatable estimator of the propensity function; in the second step we propose an UIPW estimator with the inverse of the updating propensity function value at each observation as the weight for estimating the parameter of interest. The UIPW estimation is universally applicable due to its relaxation on the constraint on the number of data batches. It is shown that the proposed estimator is consistent and asymptotically normal with the same asymptotic variance as that of the oracle estimator, and hence the oracle property is obtained. The finite sample performance of the proposed estimator is illustrated by the simulation and real data analysis. All numerical studies confirm that the UIPW estimator performs as well as the batch learner.

math.ST

A robust fusion-extraction procedure with summary statistics in the presence of biased sources

Information from various data sources is increasingly available nowadays. However, some of the data sources may produce biased estimation due to commonly encountered biased sampling, population heterogeneity, or model misspecification. This calls for statistical methods to combine information in the presence of biased sources. In this paper, a robust data fusion-extraction method is proposed. The method can produce a consistent estimator of the parameter of interest even if many of the data sources are biased. The proposed estimator is easy to compute and only employs summary statistics, and hence can be applied to many different fields, e.g. meta-analysis, Mendelian randomisation and distributed system. Moreover, the proposed estimator is asymptotically equivalent to the oracle estimator that only uses data from unbiased sources under some mild conditions. Asymptotic normality of the proposed estimator is also established. In contrast to the existing meta-analysis methods, the theoretical properties are guaranteed even if both the number of data sources and the dimension of the parameter diverge as the sample size increases, which ensures the performance of the proposed method over a wide range. The robustness and oracle property is also evaluated via simulation studies. The proposed method is applied to a meta-analysis data set to evaluate the surgical treatment for the moderate periodontal disease, and a Mendelian randomization data set to study the risk factors of head and neck cancer.

stat.ME

Distributed nonparametric regression imputation for missing response problems with large-scale data

Nonparametric regression imputation is commonly used in missing data analysis. However, it suffers from the ``curse of dimension". The problem can be alleviated by the explosive sample size in the era of big data, while the large-scale data size presents some challenges on the storage of data and the calculation of estimators. These challenges make the classical nonparametric regression imputation methods no longer applicable. This motivates us to develop two distributed nonparametric regression imputation methods. One is based on kernel smoothing and the other on the sieve method. The kernel-based distributed imputation method has extremely low communication cost and the sieve-based distributed imputation method can accommodate more local machines. To illustrate the proposed imputation methods, response mean estimation is considered. Two distributed nonparametric regression imputation estimators are proposed for the response mean, which are proved to be asymptotically normal with asymptotic variances achieving the semiparametric efficiency bound. The proposed methods are evaluated through simulation studies and are illustrated by a real data analysis.

stat.ME

Sharp bounds for variance of treatment effect estimators in the finite population in the presence of covariates

In a completely randomized experiment, the variances of treatment effect estimators in the finite population are usually not identifiable and hence not estimable. Although some estimable bounds of the variances have been established in the literature, few of them are derived in the presence of covariates. In this paper, the difference-in-means estimator and the Wald estimator are considered in the completely randomized experiment with perfect compliance and noncompliance, respectively. Sharp bounds for the variances of these two estimators are established when covariates are available. Furthermore, consistent estimators for such bounds are obtained, which can be used to shorten the confidence intervals and improve the power of tests. Confidence intervals are constructed based on the consistent estimators of the upper bounds, whose coverage rates are uniformly asymptotically guaranteed. Simulations were conducted to evaluate the proposed methods. The proposed methods are also illustrated with two real data analyses.

math.ST

A Convex Programming Solution Based Debiased Estimator for Quantile with Missing Response and High-dimensional Covariables

This paper is concerned with the estimating problem of response quantile with high dimensional covariates when response is missing at random. Some existing methods define root-n consistent estimators for the response quantile. But these methods require correct specifications of both the conditional distribution of response given covariates and the selection probability function. In this paper, a debiased method is proposed by solving a convex programming. The estimator obtained by the proposed method is asymptotically normal given a correctly specified parametric model for the condition distribution function, without the requirement to specify and estimate the selection probability function. Moreover, the proposed estimator is asymptotically more efficient than the existing estimators. The proposed method is evaluated by a simulation study and is illustrated by a real data example.

stat.ME

Determination and estimation of optimal quarantine duration for infectious diseases with application to data analysis of COVID-19

Quarantine measure is a commonly used non-pharmaceutical intervention during the outbreak of infectious diseases. A key problem for implementing quarantine measure is to determine the duration of quarantine. In this paper, a policy with optimal quarantine duration is developed. The policy suggests different quarantine durations for every individual with different characteristic. The policy is optimal in the sense that it minimizes the average quarantine duration of uninfected people with the constraint that the probability of symptom presentation for infected people attains the given value closing to 1. The optimal solution for the quarantine duration is obtained and estimated by some statistic methods with application to analyzing COVID-19 data.

stat.AP

Simultaneous confidence bands for nonparametric regression with partially missing covariates

In this paper, we consider a weighted local linear estimator based on the inverse selection probability for nonparametric regression with missing covariates at random. The asymptotic distribution of the maximal deviation between the estimator and the true regression function is derived and an asymptotically accurate simultaneous confidence band is constructed. The estimator for the regression function is shown to be oracally efficient in the sense that it is uniformly indistinguishable from that when the selection probabilities are known. Finite sample performance is examined via simulation studies which support our asymptotic theory. The proposed method is demonstrated via an analysis of a data set from the Canada 2010/2011 Youth Student Survey.

stat.ME

Finite sample breakdown point of Tukey's halfspace median

Tukey's halfspace median ($\HM$), servicing as the {multivariate} counterpart of the univariate median, has been introduced and extensively studied in the literature. It is supposed and expected to preserve robustness property (the most outstanding property) of the univariate median. One of prevalent quantitative assessments of robustness is finite sample breakdown point (FSBP). Indeed, the FSBP of many multivariate medians have been identified, except for the most prevailing one---the Tukey's halfspace median. This paper presents a precise result on FSBP for Tukey's halfspace median. The result here depicts the complete prospect of the global robustness of $\HM$ in the \emph{finite sample} practical scenario, revealing the dimension effect on the breakdown point robustness and complimenting the existing \emph{asymptotic} breakdown point result.

math.ST

Empirical likelihood for single-index varying-coefficient models

In this paper, we develop statistical inference techniques for the unknown coefficient functions and single-index parameters in single-index varying-coefficient models. We first estimate the nonparametric component via the local linear fitting, then construct an estimated empirical likelihood ratio function and hence obtain a maximum empirical likelihood estimator for the parametric component. Our estimator for parametric component is asymptotically efficient, and the estimator of nonparametric component has an optimal convergence rate. Our results provide ways to construct the confidence region for the involved unknown parameter. We also develop an adjusted empirical likelihood ratio for constructing the confidence regions of parameters of interest. A simulation study is conducted to evaluate the finite sample behaviors of the proposed methods.

math.ST

Bias-corrected GEE estimation and smooth-threshold GEE variable selection for single-index models with clustered data

In this paper, we present a generalized estimating equations based estimation approach and a variable selection procedure for single-index models when the observed data are clustered. Unlike the case of independent observations, bias-correction is necessary when general working correlation matrices are used in the estimating equations. Our variable selection procedure based on smooth-threshold estimating equations \citep{Ueki-2009} can automatically eliminate irrelevant parameters by setting them as zeros and is computationally simpler than alternative approaches based on shrinkage penalty. The resulting estimator consistently identifies the significant variables in the index, even when the working correlation matrix is misspecified. The asymptotic property of the estimator is the same whether or not the nonzero parameters are known (in both cases we use the same estimating equations), thus achieving the oracle property in the sense of \cite{Fan-Li-2001}. The finite sample properties of the estimator are illustrated by some simulation examples, as well as a real data application.

stat.ME