SearcharxivSearch

arXiv subjects

Mucyo Karemera

Publications and source records attributed to Mucyo Karemera.

12 recordsLinked to original sources

An accurate percentile method for parametric inference based on asymptotically biased estimators

Inference methods for computing confidence intervals in parametric settings usually rely on consistent estimators of the parameter of interest. However, it may be computationally and/or analytically burdensome to obtain such estimators in various parametric settings, for example when the data exhibit certain features such as censoring, misclassification errors or outliers. To address these challenges, we propose a simulation-based inferential method, called the implicit bootstrap, that remains valid regardless of the potential asymptotic bias of the estimator on which the method is based. We demonstrate that this method allows for the construction of asymptotically valid percentile confidence intervals of the parameter of interest. Additionally, we show that these confidence intervals can also achieve second-order accuracy. We also show that the method is exact in three instances where the standard bootstrap fails. Using simulation studies, we illustrate the coverage accuracy of the method in three examples where standard parametric bootstrap procedures are computationally intensive and less accurate in finite samples.

stat.ME

Just Identified Indirect Inference Estimator: Accurate Inference through Bias Correction

An important challenge in statistical analysis lies in controlling the estimation bias when handling the ever-increasing data size and model complexity of modern data settings. In this paper, we propose a reliable estimation and inference approach for parametric models based on the Just Identified iNdirect Inference estimator (JINI). The key advantage of our approach is that it allows to construct a consistent estimator in a simple manner, while providing strong bias correction guarantees that lead to accurate inference. Our approach is particularly useful for complex parametric models, as it allows to bypass the analytical and computational difficulties (e.g., due to intractable estimating equation) typically encountered in standard procedures. The properties of JINI (including consistency, asymptotic normality, and its bias correction property) are also studied when the parameter dimension is allowed to diverge, which provide the theoretical foundation to explain the advantageous performance of JINI in increasing dimensional covariates settings. Our simulations and an alcohol consumption data analysis highlight the practical usefulness and excellent performance of JINI when data present features (e.g., misclassification, rounding) as well as in robust estimation.

stat.ME

Accounting for Vibration Noise in Stochastic Measurement Errors

The measurement of data over time and/or space is of utmost importance in a wide range of domains from engineering to physics. Devices that perform these measurements therefore need to be extremely precise to obtain correct system diagnostics and accurate predictions, consequently requiring a rigorous calibration procedure which models their errors before being employed. While the deterministic components of these errors do not represent a major modelling challenge, most of the research over the past years has focused on delivering methods that can explain and estimate the complex stochastic components of these errors. This effort has allowed to greatly improve the precision and uncertainty quantification of measurement devices but has this far not accounted for a significant stochastic noise that arises for many of these devices: vibration noise. Indeed, having filtered out physical explanations for this noise, a residual stochastic component often carries over which can drastically affect measurement precision. This component can originate from different sources, including the internal mechanics of the measurement devices as well as the movement of these devices when placed on moving objects or vehicles. To remove this disturbance from signals, this work puts forward a modelling framework for this specific type of noise and adapts the Generalized Method of Wavelet Moments to estimate these models. We deliver the asymptotic properties of this method when applied to processes that include vibration noise and show the considerable practical advantages of this approach in simulation and applied case studies.

stat.ME

$\mathfrak{sl}_3$ Matrix Dilogarithm as a $6j$-Symbol

We construct quantum invariants of 3-manifolds based on a $\mathfrak{sl}_3$ matrix dilogarithm proposed by Kashaev. This matrix dilogarithm is an $\mathfrak{sl}_3$ analogue of the (cyclic) quantum dilogarithm used to define Kashaev's invariants as well as Baseilhac and Benedetti's quantum hyperbolic invariants. % In this article, we show that the $\mathfrak{sl}_3$ matrix dilogarithm can be considered as a 6$j$-symbol associated to modules of a quantum group related to $U_q(\mathfrak{sl}_3)$. Moreover, we show that the quantum invariants aforementioned allow to define a $\mathfrak{sl}_3$ version of Kashaev's invariants, opening a route to define a $\mathfrak{sl}_3$ version of Baseilhac and Benedetti's quantum hyperbolic invariants.

math.QA

SWAG: A Wrapper Method for Sparse Learning

The majority of machine learning methods and algorithms give high priority to prediction performance which may not always correspond to the priority of the users. In many cases, practitioners and researchers in different fields, going from engineering to genetics, require interpretability and replicability of the results especially in settings where, for example, not all attributes may be available to them. As a consequence, there is the need to make the outputs of machine learning algorithms more interpretable and to deliver a library of "equivalent" learners (in terms of prediction performance) that users can select based on attribute availability in order to test and/or make use of these learners for predictive/diagnostic purposes. To address these needs, we propose to study a procedure that combines screening and wrapper approaches which, based on a user-specified learning method, greedily explores the attribute space to find a library of sparse learners with consequent low data collection and storage costs. This new method (i) delivers a low-dimensional network of attributes that can be easily interpreted and (ii) increases the potential replicability of results based on the diversity of attribute combinations defining strong learners with equivalent predictive power. We call this algorithm "Sparse Wrapper AlGorithm" (SWAG).

stat.ML

A General Approach for Simulation-based Bias Correction in High Dimensional Settings

An important challenge in statistical analysis lies in controlling the bias of estimators due to the ever-increasing data size and model complexity. Approximate numerical methods and data features like censoring and misclassification often result in analytical and/or computational challenges when implementing standard estimators. As a consequence, consistent estimators may be difficult to obtain, especially in complex and/or high dimensional settings. In this paper, we study the properties of a general simulation-based estimation framework that allows to construct bias corrected consistent estimators. We show that the considered approach leads, under more general conditions, to stronger bias correction properties compared to alternative methods. Besides its bias correction advantages, the considered method can be used as a simple strategy to construct consistent estimators in settings where alternative methods may be challenging to apply. Moreover, the considered framework can be easily implemented and is computationally efficient. These theoretical results are highlighted with simulation studies of various commonly used models, including the negative binomial regression (with and without censoring) and the logistic regression (with and without misclassification errors). Additional numerical illustrations are provided in the supplementary materials.

math.ST

Asymptotically Optimal Bias Reduction for Parametric Models

An important challenge in statistical analysis concerns the control of the finite sample bias of estimators. This problem is magnified in high-dimensional settings where the number of variables $p$ diverges with the sample size $n$, as well as for nonlinear models and/or models with discrete data. For these complex settings, we propose to use a general simulation-based approach and show that the resulting estimator has a bias of order $\mathcal{O}(0)$, hence providing an asymptotically optimal bias reduction. It is based on an initial estimator that can be slightly asymptotically biased, making the approach very generally applicable. This is particularly relevant when classical estimators, such as the maximum likelihood estimator, can only be (numerically) approximated. We show that the iterative bootstrap of Kuk (1995) provides a computationally efficient approach to compute this bias reduced estimator. We illustrate our theoretical results in simulation studies for which we develop new bias reduced estimators for the logistic regression, with and without random effects. These estimators enjoy additional properties such as robustness to data contamination and to the problem of separability.

math.ST

Wavelet-Based Moment-Matching Techniques for Inertial Sensor Calibration

The task of inertial sensor calibration has required the development of various techniques to take into account the sources of measurement error coming from such devices. The calibration of the stochastic errors of these sensors has been the focus of increasing amount of research in which the method of reference has been the so-called "Allan variance slope method" which, in addition to not having appropriate statistical properties, requires a subjective input which makes it prone to mistakes. To overcome this, recent research has started proposing "automatic" approaches where the parameters of the probabilistic models underlying the error signals are estimated by matching functions of the Allan variance or Wavelet Variance with their model-implied counterparts. However, given the increased use of such techniques, there has been no study or clear direction for practitioners on which approach is optimal for the purpose of sensor calibration. This paper formally defines the class of estimators based on this technique and puts forward theoretical and applied results that, comparing with estimators in this class, suggest the use of the Generalized Method of Wavelet Moments as an optimal choice.

stat.ME

Phase Transition Unbiased Estimation in High Dimensional Settings

An important challenge in statistical analysis concerns the control of the finite sample bias of estimators. For example, the maximum likelihood estimator has a bias that can result in a significant inferential loss. This problem is typically magnified in high-dimensional settings where the number of variables $p$ is allowed to diverge with the sample size $n$. However, it is generally difficult to establish whether an estimator is unbiased and therefore its asymptotic order is a common approach used (in low-dimensional settings) to quantify the magnitude of the bias. As an alternative, we introduce a new and stronger property, possibly for high-dimensional settings, called phase transition unbiasedness. An estimator satisfying this property is unbiased for all $n$ greater than a finite sample size $n^\ast$. Moreover, we propose a phase transition unbiased estimator built upon the idea of matching an initial estimator computed on the sample and on simulated data. It is not required for this initial estimator to be consistent and thus it can be chosen for its computational efficiency and/or for other desirable properties such as robustness. This estimator can be computed using a suitable simulation based algorithm, namely the iterative bootstrap, which is shown to converge exponentially fast. In addition, we demonstrate the consistency and the limiting distribution of this estimator in high-dimensional settings. Finally, as an illustration, we use our approach to develop new estimators for the logistic regression model, with and without random effects, that also enjoy other properties such as robustness to data contamination and are also not affected by the problem of separability. In a simulation exercise, the theoretical results are confirmed in settings where the sample size is relatively small compared to the model dimension.

math.ST

Multivariate Signal Modelling with Applications to Inertial Sensor Calibration

The common approach to inertial sensor calibration for navigation purposes has been to model the stochastic error signals of individual sensors independently, whether as components of a single inertial measurement unit (IMU) in different directions or arrayed in the same direction for redundancy. These signals usually have an extremely complex spectral structure that is often described using latent (or composite) models composed by a sum of underlying models. A large amount of research in this domain has been focused on the latter aspect through the proposal of various methods that have been able to improve the estimation of these models both from a computational and a statistical point of view. However, the separate calibration of the individual sensors is still unable to take into account the dependence between each of them which can have an important impact on the precision of the navigation systems. In this paper, we develop a new approach to simultaneously model both the individual signals and the dependence between them by studying the quantity called Wavelet Cross-Covariance and using it to extend the application of the Generalized Method of Wavelet Moments. This new method can be used in other settings for time series modeling, especially in cases where the dependence among signals may be hard to detect. Moreover, in the field of inertial sensor calibration, this approach can deliver important contributions among which the possibility to test dependence between sensors, integrate their dependence within the navigation filter and construct an optimal virtual sensor that can be used to simplify and improve navigation accuracy. The advantages of this method and its usefulness for inertial sensor calibration are highlighted through a simulation study and an applied example with a small array of XSens MTi-G IMUs.

stat.ME

A simple recipe for making accurate parametric inference in finite sample

Constructing tests or confidence regions that control over the error rates in the long-run is probably one of the most important problem in statistics. Yet, the theoretical justification for most methods in statistics is asymptotic. The bootstrap for example, despite its simplicity and its widespread usage, is an asymptotic method. There are in general no claim about the exactness of inferential procedures in finite sample. In this paper, we propose an alternative to the parametric bootstrap. We setup general conditions to demonstrate theoretically that accurate inference can be claimed in finite sample.

stat.ME

On the Properties of Simulation-based Estimators in High Dimensions

Considering the increasing size of available data, the need for statistical methods that control the finite sample bias is growing. This is mainly due to the frequent settings where the number of variables is large and allowed to increase with the sample size bringing standard inferential procedures to incur significant loss in terms of performance. Moreover, the complexity of statistical models is also increasing thereby entailing important computational challenges in constructing new estimators or in implementing classical ones. A trade-off between numerical complexity and statistical properties is often accepted. However, numerically efficient estimators that are altogether unbiased, consistent and asymptotically normal in high dimensional problems would generally be ideal. In this paper, we set a general framework from which such estimators can easily be derived for wide classes of models. This framework is based on the concepts that underlie simulation-based estimation methods such as indirect inference. The approach allows various extensions compared to previous results as it is adapted to possibly inconsistent estimators and is applicable to discrete models and/or models with a large number of parameters. We consider an algorithm, namely the Iterative Bootstrap (IB), to efficiently compute simulation-based estimators by showing its convergence properties. Within this framework we also prove the properties of simulation-based estimators, more specifically the unbiasedness, consistency and asymptotic normality when the number of parameters is allowed to increase with the sample size. Therefore, an important implication of the proposed approach is that it allows to obtain unbiased estimators in finite samples. Finally, we study this approach when applied to three common models, namely logistic regression, negative binomial regression and lasso regression.

math.ST