SearcharxivSearch

arXiv subjects

Alessandra R. Brazzale

Publications and source records attributed to Alessandra R. Brazzale.

6 recordsLinked to original sources

Advances in Ontology--Based Mining of Adverse Drug Reactions

Post--marketing pharmacovigilance is essential for identifying adverse drug reactions (ADRs) that elude detection during pre--marketing clinical trials. This study explores a novel approach that integrates an adverse event (AE) ontology into a zero--inflated negative binomial model to improve ADR detection. By accounting for the biological similarities among correlated AEs and addressing the excess of zero counts, this method more effectively disentangles AE associations. Statistical significance is evaluated using a permutation--based maximum statistic that preserves AE correlations within individual reports. Simulations and an application to real data from the Veneto drug safety database demonstrate that the ontology--based model consistently outperforms classical models such as the Gamma--Poisson Shrinker (GPS). For post--selection inference, we furthermore explore a data thinning technique for convolution--closed families, enabling the creation of independent training and validation datasets while retaining all drug--AE pairs. This approach is compared with conventional random train/test splitting, which may leave some drugs or AEs absent from one subset, and stratified splitting, which requires expanding aggregated counts into individual instances. The data--thinning technique and stratified splitting yield very similar results, with stratified splitting showing a slight benefit, and both clearly outperform random splitting in ensuring reliable and consistent model evaluation.

stat.ME

Likelihood Asymptotics in Nonregular Settings: A Review with Emphasis on the Likelihood Ratio

This paper reviews the most common situations where one or more regularity conditions which underlie classical likelihood-based parametric inference fail. We identify three main classes of problems: boundary problems, indeterminate parameter problems -- which include non-identifiable parameters and singular information matrices -- and change-point problems. The review focuses on the large-sample properties of the likelihood ratio statistic. We emphasize analytical solutions and acknowledge software implementations where available. We furthermore give summary insight about the possible tools to derivate the key results. Other approaches to hypothesis testing and connections to estimation are listed in the annotated bibliography of the Supplementary Material.

math.ST

Locating γ-Ray Sources on the Celestial Sphere via Modal Clustering

Searching for as yet undetected gamma-ray sources is a major target of the Fermi LAT Collaboration. We present an algorithm capable of identifying such type of sources by non-parametrically clustering the directions of arrival of the high-energy photons detected by the telescope onboard the Fermi spacecraft. n particular, the sources will be identified using a von Mises-Fisher kernel estimate of the photon count density on the unit sphere via an adjustment of the mean-shift algorithm to account for the directional nature of data. This choice entails a number of desirable benefits. It allows us to by-pass the difficulties inherent on the borders of any projection of the photon directions onto a 2-dimensional plane, while guaranteeing high flexibility. The smoothing parameter will be chosen adaptively, by combining scientific input with optimal selection guidelines, as known from the literature. Using statistical tools from hypothesis testing and classification, we furthermore present an automatic way to skim off sound candidate sources from the gamma-ray emitting diffuse background and to quantify their significance. The algorithm was calibrated on simulated data provided by the Fermi LAT Collaboration and will be illustrated on a real Fermi LAT case-study.

astro-ph.IM

Identification of high-energy astrophysical point sources via hierarchical Bayesian nonparametric clustering

The light we receive from distant astrophysical objects carries information about their origins and the physical mechanisms that power them. The study of these signals, however, is complicated by the fact that observations are often a mixture of the light emitted by multiple localized sources situated in a spatially-varying background. A general algorithm to achieve robust and accurate source identification in this case remains an open question in astrophysics. This paper focuses on high-energy light (such as X-rays and gamma-rays), for which observatories can detect individual photons (quanta of light), measuring their incoming direction, arrival time, and energy. Our proposed Bayesian methodology uses both the spatial and energy information to identify point sources, that is, separate them from the spatially-varying background, to estimate their number, and to compute the posterior probabilities that each photon originated from each identified source. This is accomplished via a Dirichlet process mixture while the background is simultaneously reconstructed via a flexible Bayesian nonparametric model based on B-splines. Our proposed method is validated with a suite of simulation studies and illustrated with an application to a complex region of the sky observed by the \emph{Fermi} Gamma-ray Space Telescope.

stat.AP

Margin-free classification and new class detection using finite Dirichlet mixtures

We present a margin-free finite mixture model which allows us to simultaneously classify objects into known classes and to identify possible new object types using a set of continuous attributes. This application is motivated by the needs of identifying and possibly detecting new types of a particular kind of stars known as variable stars. We first suitably transform the physical attributes of the stars onto the simplex to achieve scale invariance while maintaining their dependence structure. This allows us to compare data collected by different sky surveys which can have different scales. The model hence combines a mixture of Dirichlet mixtures to represent the known classes with the semi-supervised classification strategy of Vatanen et al. (2012) for outlier detection. In line with previous work on semiparametric model-based clustering, the single Dirichlet distributions can be seen as providing the baseline pattern of the data. These are then combined to effectively model the complex distributions of the attributes for the different classes. The model is estimated using a hierarchical two-step procedure which combines a suitably adapted version of the Expectation-Maximization (EM) algorithm with Bayes' rule. We validate our model on a reliable sample of periodic variable stars available in the literature (Dubath et al., 2011) achieving an overall classification accuracy of 71.95%, a sensitivity of 86.11% and a specificity of 99.79% for new class detection.

stat.AP

Accurate Parametric Inference for Small Samples

We outline how modern likelihood theory, which provides essentially exact inferences in a variety of parametric statistical problems, may routinely be applied in practice. Although the likelihood procedures are based on analytical asymptotic approximations, the focus of this paper is not on theory but on implementation and applications. Numerical illustrations are given for logistic regression, nonlinear models, and linear non-normal models, and we describe a sampling approach for the third of these classes. In the case of logistic regression, we argue that approximations are often more appropriate than `exact' procedures, even when these exist.

stat.ME