SearcharxivSearch

arXiv subjects

Malgorzata Bogdan

Publications and source records attributed to Malgorzata Bogdan.

At least 19 recordsLinked to original sources

Hierarchical Bayesian Estimation of Covariance Matrices

We develop a hierarchical Bayesian framework for covariance matrix estimation built on a key observation: while equivariance under the full general linear group GL(p) is well known, it is an extremely restrictive property -- estimators equivariant to GL(p) are limited to scalar multiples of the sample covariance matrix and carry considerably larger risks than shrinkage estimators. By contrast, commonly used shrinkage estimators, including the Haff empirical Bayes estimator, and the Ledoit--Wolf estimators, are all equivariant under the smaller orthogonal group O(p). Exploiting this structure, we establish that the Haar measure Bayes rule in an oracle eigenvalue model is the minimum risk estimator within the class of O(p)-equivariant estimators, and derive oracle Bayes rules for the covariance and precision matrices under the squared Frobenius, Stein, and squared Stein loss functions. These oracle rules serve as theoretical benchmarks that dominate all commonly used estimators. To approximate them when the true eigenvalues are unknown, we introduce a hierarchical Bayes model that places a finite P'olya tree prior on the eigenvalue distribution and uses Gibbs sampling to generate posterior draws, yielding both shrinkage estimates for the eigenvalues and approximations to the oracle Bayes rules. Simulations suggest that the finite P'olya tree prior is able to recover the general form of the distribution of the eigenvalues, and confirm that the resulting estimators closely approach oracle performance, substantially outperforming classical competitors for both covariance and precision matrix estimation.

stat.ME

Redshift Classification of Optical Gamma-Ray Bursts using Supervised Learning

Gamma-ray bursts (GRBs) are among the most luminous explosions in the Universe and serve as powerful probes of the early cosmos. However, the rapid fading of their afterglows and the scarcity of spectroscopic measurements make photometric classification crucial for timely high-redshift identification. We present an ensemble machine learning framework for redshift classification of GRBs based solely on their optical plateau and prompt emission properties. Our dataset comprises 171 long GRBs observed by the Swift UVOT and more than 450 ground-based telescopes. The analysis pipeline integrates robust statistical techniques, including M-estimator outlier rejection, multivariate imputation using Multiple Imputation by Chained Equations, and Least Absolute Shrinkage and Selection Operator feature selection, followed by a SuperLearner ensemble combining parametric, semi-parametric, and non-parametric algorithms. The optimal model, trained on raw optical data with outlier removal at a redshift threshold of z equals 2.0, achieves a true positive rate of 74 percent and an area under the curve of 0.84, maintaining balanced generalization between training and test sets. At higher thresholds, such as z equals 3.0, the classifier sustains strong discriminative power with an area under the curve of 0.88. Validation on an independent GRB sample yields 97 percent overall accuracy, perfect specificity, and an ensemble area under the curve of 0.93. Compared to previous prompt- and X-ray-based classifiers, our optical framework offers enhanced sensitivity to high-redshift events, improved robustness against data incompleteness, and greater applicability to ground-based follow-up. We also publicly release a web application that enables real-time redshift classification, facilitating rapid identification of candidate high-redshift GRBs for cosmological studies.

astro-ph.HE

Efficient Solvers for SLOPE in R, Python, Julia, and C++

We present a suite of packages in R, Python, Julia, and C++ that efficiently solve the Sorted L-One Penalized Estimation (SLOPE) problem. The packages feature a highly efficient hybrid coordinate descent algorithm that fits generalized linear models (GLMs) and supports a variety of loss functions, including Gaussian, binomial, Poisson, and multinomial logistic regression. Our implementation is designed to be fast, memory-efficient, and flexible. The packages support a variety of data structures (dense, sparse, and out-of-memory matrices) and are designed to efficiently fit the full SLOPE path as well as handle cross-validation of SLOPE models, including the relaxed SLOPE. We present examples of how to use the packages and benchmarks that demonstrate the performance of the packages on both real and simulated data and show that our packages outperform existing implementations of SLOPE in terms of speed.

stat.CO

Gamma-ray Bursts as Distance Indicators by a Statistical Learning Approach

Gamma-ray bursts (GRBs) can be probes of the early universe, but currently, only 26% of GRBs observed by the Neil Gehrels Swift Observatory GRBs have known redshifts ($z$) due to observational limitations. To address this, we estimated the GRB redshift (distance) via a supervised statistical learning model that uses optical afterglow observed by Swift and ground-based telescopes. The inferred redshifts are strongly correlated (a Pearson coefficient of 0.93) with the observed redshifts, thus proving the reliability of this method. The inferred and observed redshifts allow us to estimate the number of GRBs occurring at a given redshift (GRB rate) to be 8.47-9 $yr^{-1} Gpc^{-1}$ for $1.9<z<2.3$. Since GRBs come from the collapse of massive stars, we compared this rate with the star formation rate highlighting a discrepancy of a factor of 3 at $z<1$.

astro-ph.HE

GRB Redshift Estimation using Machine Learning and the Associated Web-App

Context. Gamma-ray bursts (GRBs), observed at redshifts as high as 9.4, could serve as valuable probes for investigating the distant Universe. However, this necessitates an increase in the number of GRBs with determined redshifts, as currently, only 12% of GRBs have known redshifts due to observational biases. Aims. We aim to address the shortage of GRBs with measured redshifts, enabling us to fully realize their potential as valuable cosmological probes Methods. Following Dainotti et al. (2024c), we have taken a second step to overcome this issue by adding 30 more GRBs to our ensemble supervised machine learning training sample, an increase of 20%, which will help us obtain better redshift estimates. In addition, we have built a freely accessible and user-friendly web app that infers the redshift of long GRBs (LGRBs) with plateau emission using our machine learning model. The web app is the first of its kind for such a study and will allow the community to obtain redshift estimates by entering the GRB parameters in the app. Results. Through our machine learning model, we have successfully estimated redshifts for 276 LGRBs using X-ray afterglow parameters detected by the Neil Gehrels Swift Observatory and increased the sample of LGRBs with known redshifts by 110%. We also perform Monte Carlo simulations to demonstrate the future applicability of this research. Conclusions. The results presented in this research will enable the community to increase the sample of GRBs with known redshift estimates. This can help address many outstanding issues, such as GRB formation rate, luminosity function, and the true nature of low-luminosity GRBs, and enable the application of GRBs as standard candles

astro-ph.HE

GRB Redshift Classifier to Follow-up High-Redshift GRBs Using Supervised Machine Learning

Gamma-ray bursts (GRBs) are intense, short-lived bursts of gamma-ray radiation observed up to a high redshift ($z \sim 10$) due to their luminosities. Thus, they can serve as cosmological tools to probe the early Universe. However, we need a large sample of high$-z$ GRBs, currently limited due to the difficulty in securing time at the large aperture Telescopes. Thus, it is painstaking to determine quickly whether a GRB is high$z$ or low$-z$, which hampers the possibility of performing rapid follow-up observations. Previous efforts to distinguish between high$-$ and low$-z$ GRBs using GRB properties and machine learning (ML) have resulted in limited sensitivity. In this study, we aim to improve this classification by employing an ensemble ML method on 251 GRBs with measured redshifts and plateaus observed by the Neil Gehrels Swift Observatory. Incorporating the plateau phase with the prompt emission, we have employed an ensemble of classification methods to enhance the sensitivity unprecedentedly. Additionally, we investigate the effectiveness of various classification methods using different redshift thresholds, $z_{threshold}$=$z_t$ at $z_{t}=$ 2.0, 2.5, 3.0, and 3.5. We achieve a sensitivity of 87\% and 89\% with a balanced sampling for both $z_{t}=3.0$ and $z_{t}=3.5$, respectively, representing a 9\% and 11\% increase in the sensitivity over Random Forest used alone. Overall, the best results are at $z_{t} = 3.5$, where the difference between the sensitivity of the training set and the test set is the smallest. This enhancement of the proposed method paves the way for new and intriguing follow-up observations of high$-z$ GRBs.

astro-ph.HE

Reduced uncertainties up to 43\% on the Hubble constant and the matter density with the SNe Ia with a new statistical analysis

Type Ia Supernovae (SNe Ia) are considered the most reliable \textit{standard candles} and they have played an invaluable role in cosmology since the discovery of the Universe's accelerated expansion. During the last decades, the SNe Ia samples have been improved in number, redshift coverage, calibration methodology, and systematics treatment. These efforts led to the most recent \textit{``Pantheon"} (2018) and \textit{``Pantheon +"} (2022) releases, which enable to constrain cosmological parameters more precisely than previous samples. In this era of precision cosmology, the community strives to find new ways to reduce uncertainties on cosmological parameters. To this end, we start our investigation even from the likelihood assumption of Gaussianity, implicitly used in this domain. Indeed, the usual practise involves constraining parameters through a Gaussian distance moduli likelihood. This method relies on the implicit assumption that the difference between the distance moduli measured and the ones expected from the cosmological model is Gaussianly distributed. In this work, we test this hypothesis for both the \textit{Pantheon} and \textit{Pantheon +} releases. We find that in both cases this requirement is not fulfilled and the actual underlying distributions are a logistic and a Student's t distribution for the \textit{Pantheon} and \textit{Pantheon +} data, respectively. When we apply these new likelihoods fitting a flat $Λ$CDM model, we significantly reduce the uncertainties on $Ω_M$ and $H_0$ of $\sim 40 \%$. This boosts the SNe Ia power in constraining cosmological parameters, thus representing a huge step forward to shed light on the current debated tensions in cosmology.

astro-ph.CO

Inferring the redshift of more than 150 GRBs with a Machine Learning Ensemble model

Gamma-Ray Bursts (GRBs), due to their high luminosities are detected up to redshift 10, and thus have the potential to be vital cosmological probes of early processes in the universe. Fulfilling this potential requires a large sample of GRBs with known redshifts, but due to observational limitations, only 11\% have known redshifts ($z$). There have been numerous attempts to estimate redshifts via correlation studies, most of which have led to inaccurate predictions. To overcome this, we estimated GRB redshift via an ensemble supervised machine learning model that uses X-ray afterglows of long-duration GRBs observed by the Neil Gehrels Swift Observatory. The estimated redshifts are strongly correlated (a Pearson coefficient of 0.93) and have a root mean square error, namely the square root of the average squared error $\langleΔz^2\rangle$, of 0.46 with the observed redshifts showing the reliability of this method. The addition of GRB afterglow parameters improves the predictions considerably by 63\% compared to previous results in peer-reviewed literature. Finally, we use our machine learning model to infer the redshifts of 154 GRBs, which increase the known redshifts of long GRBs with plateaus by 94\%, a significant milestone for enhancing GRB population studies that require large samples with redshift.

astro-ph.CO

Shedding new light on the Hubble constant tension through Supernovae Ia

The standard cosmological model, the $Λ$CDM model, is the most suitable description for our universe. This framework can explain the accelerated expansion phase of the universe but still is not immune to open problems when it comes to the comparison with observations. One of the most critical issues is the so-called Hubble constant ($H_0$) tension, namely, the difference of about $5σ$ as an average between the value of $H_0$ estimated locally and the cosmological value measured from the Last Scattering Surface. The value of this tension changes from 4 to 6 $σ$ according to the data used. The current analysis explores the $H_0$ tension in the \textit{Pantheon} sample (PS) of SNe Ia. Through the division of the PS in 3 and 4 bins, the value of $H_0$ is estimated for each bin and all the values are fitted with a decreasing function of the redshift ($z$). Remarkably, $H_0$ undergoes a slow decreasing evolution with $z$, having an evolutionary coefficient compatible with zero up to $5.8σ$. If this trend is not caused by hidden astrophysical biases or $z$-selection effects, then the $f(R)$ modified theories of gravity represent a valid model for explaining such a trend.

astro-ph.CO

Fermi LAT AGN classification using supervised machine learning

Classifying Active Galactic Nuclei (AGN) is a challenge, especially for BL Lac Objects (BLLs), which are identified by their weak emission line spectra. To address the problem of classification, we use data from the 4th Fermi Catalog, Data Release 3. Missing data hinders the use of machine learning to classify AGN. A previous paper found that Multiple Imputation by Chain Equations (MICE) imputation is useful for estimating missing values. Since many AGN have missing redshift and the highest energy, we use data imputation with MICE and K-nearest neighbor (kNN) algorithm to fill in these missing variables. Then, we classify AGN into the BLLs or the Flat Spectrum Radio Quasars (FSRQs) using the SuperLearner, an ensemble method that includes several classification algorithms like logistic regression, support vector classifiers, Random Forests, Ranger Random Forests, multivariate adaptive regression spline (MARS), Bayesian regression, Extreme Gradient Boosting. We find that a SuperLearner model using MARS regression and Random Forests algorithms is 91.1% accurate for kNN imputed data and 91.2% for MICE imputed data. Furthermore, the kNN-imputed SuperLearner model predicts that 892 of the 1519 unclassified blazars are BLLs and 627 are Flat Spectrum Radio Quasars (FSRQs), while the MICE-imputed SuperLearner model predicts 890 BLLs and 629 FSRQs in the unclassified set. Thus, we can conclude that both imputation methods work efficiently and with high accuracy and that our methodology ushers the way for using SuperLearner as a novel classification method in the AGN community and, in general, in the astrophysics community.

astro-ph.HE

Sparse Graphical Modelling via the Sorted L$_1$-Norm

Sparse graphical modelling has attained widespread attention across various academic fields. We propose two new graphical model approaches, Gslope and Tslope, which provide sparse estimates of the precision matrix by penalizing its sorted L1-norm, and relying on Gaussian and T-student data, respectively. We provide the selections of the tuning parameters which provably control the probability of including false edges between the disjoint graph components and empirically control the False Discovery Rate for the block diagonal covariance matrices. In extensive simulation and real world analysis, the new methods are compared to other state-of-the-art sparse graphical modelling approaches. The results establish Gslope and Tslope as two new effective tools for sparse network estimation, when dealing with both Gaussian, t-student and mixture data.

stat.ME

Using Multivariate Imputation by Chained Equations to Predict Redshifts of Active Galactic Nuclei

Redshift measurement of active galactic nuclei (AGNs) remains a time-consuming and challenging task, as it requires follow up spectroscopic observations and detailed analysis. Hence, there exists an urgent requirement for alternative redshift estimation techniques. The use of machine learning (ML) for this purpose has been growing over the last few years, primarily due to the availability of large-scale galactic surveys. However, due to observational errors, a significant fraction of these data sets often have missing entries, rendering that fraction unusable for ML regression applications. In this study, we demonstrate the performance of an imputation technique called Multivariate Imputation by Chained Equations (MICE), which rectifies the issue of missing data entries by imputing them using the available information in the catalog. We use the Fermi-LAT Fourth Data Release Catalog (4LAC) and impute 24% of the catalog. Subsequently, we follow the methodology described in Dainotti et al. (2021) and create an ML model for estimating the redshift of 4LAC AGNs. We present results which highlight positive impact of MICE imputation technique on the machine learning models performance and obtained redshift estimation accuracy.

astro-ph.IM

On the evolution of the Hubble constant with the SNe Ia Pantheon Sample and Baryon Acoustic Oscillations: a feasibility study for GRB-cosmology in 2030

The difference from 4 to 6 $σ$ in the Hubble constant ($H_0$) between the values observed with the local (Cepheids and Supernovae Ia, SNe Ia) and the high-z probes (CMB obtained by the Planck data) still challenges the astrophysics and cosmology community. Previous analysis has shown that there is an evolution in the Hubble constant that scales as $f(z)=\mathcal{H}_{0}/(1+z)^η$, where $\mathcal{H}_0$ is $H_{0}(z=0)$ and $η$ is the evolutionary parameter. Here, we investigate if this evolution still holds by using the SNe Ia gathered in the Pantheon sample and the BAOs. We assume $H_{0}=70 \textrm{km s}^{-1}\textrm{Mpc}^{-1}$ as the local value and divide the Pantheon into 3 bins ordered in increasing values of redshift. Similar to our previous analysis but varying two cosmological parameters contemporaneously ($H_0$, $Ω_{0m}$ in the $Λ$CDM model and $H_0$, $w_a$ in the $w_{0}w_{a}$CDM model), for each bin we implement a MCMC analysis obtaining the value of $H_0$ $[...]$. Subsequently, the values of $H_0$ are fitted with the model $f(z)$. Our results show that a decreasing trend with $η\sim10^{-2}$ is still visible in this sample. The $η$ coefficient reaches zero in 2.0 $σ$ for the $Λ$CDM model up to 5.8 $σ$ for $w_{0}w_{a}$CDM model. This trend, if not due to statistical fluctuations, could be explained through a hidden astrophysical bias, such as the effect of stretch evolution, or it requires new theoretical models, a possible proposition is the modified gravity theories, $f(R)$ $[...]$. This work is also a preparatory to understand how the combined probes still show an evolution of the $H_0$ by redshift and what is the current status of simulations on GRB cosmology to obtain the uncertainties on the $Ω_{0m}$ comparable with the ones achieved through SNe Ia.

astro-ph.CO

Predicting the redshift of gamma-ray loud AGNs using Supervised Machine Learning: Part 2

Measuring the redshift of active galactic nuclei (AGNs) requires the use of time-consuming and expensive spectroscopic analysis. However, obtaining redshift measurements of AGNs is crucial as it can enable AGN population studies, provide insight into the star formation rate, the luminosity function, and the density rate evolution. Hence, there is a requirement for alternative redshift measurement techniques. In this project, we aim to use the Fermi gamma-ray space telescope's 4LAC Data Release (DR2) catalog to train a machine learning model capable of predicting the redshift reliably. In addition, this project aims at improving and extending with the new 4LAC Catalog the predictive capabilities of the machine learning (ML) methodology published in Dainotti et al. (2021). Furthermore, we implement feature engineering to expand the parameter space and a bias correction technique to our final results. This study uses additional machine learning techniques inside the ensemble method, the SuperLearner, previously used in Dainotti et al.(2021). Additionally, we also test a novel ML model called Sorted L-One Penalized Estimation (SLOPE). Using these methods we provide a catalog of estimated redshift values for those AGNs that do not have a spectroscopic redshift measurement. These estimates can serve as a redshift reference for the community to verify as updated Fermi catalogs are released with more redshift measurements.

astro-ph.HE

On the sign recovery by LASSO, thresholded LASSO and thresholded Basis Pursuit Denoising

Basis Pursuit (BP), Basis Pursuit DeNoising (BPDN), and LASSO are popular methods for identifying important predictors in the high-dimensional linear regression model, i.e. when the number of rows of the design matrix X is smaller than the number of columns. By definition, BP uniquely recovers the vector of regression coefficients b if there is no noise and the vector b has the smallest L1 norm among all vectors s such that Xb=Xs (identifiability condition). Furthermore, LASSO can recover the sign of b only under a much stronger irrepresentability condition. Meanwhile, it is known that the model selection properties of LASSO can be improved by hard-thresholding its estimates. This article supports these findings by proving that thresholded LASSO, thresholded BPDN and thresholded BP recover the sign of b in both the noisy and noiseless cases if and only if b is identifiable and large enough. In particular, if X has iid Gaussian entries and the number of predictors grows linearly with the sample size, then these thresholded estimators can recover the sign of b when the signal sparsity is asymptotically below the Donoho-Tanner transition curve. This is in contrast to the regular LASSO, which asymptotically recovers the sign of b only when the signal sparsity tends to 0. Numerical experiments show that the identifiability condition, unlike the irrepresentability condition, does not seem to be affected by the structure of the correlations in the $X$ matrix.

stat.ME

Predicting the redshift of gamma-ray loud AGNs using supervised machine learning

AGNs are very powerful galaxies characterized by extremely bright emissions coming out from their central massive black holes. Knowing the redshifts of AGNs provides us with an opportunity to determine their distance to investigate important astrophysical problems such as the evolution of the early stars, their formation along with the structure of early galaxies. The redshift determination is challenging because it requires detailed follow-up of multi-wavelength observations, often involving various astronomical facilities. Here, we employ machine learning algorithms to estimate redshifts from the observed gamma-ray properties and photometric data of gamma-ray loud AGN from the Fourth Fermi-LAT Catalog. The prediction is obtained with the Superlearner algorithm, using LASSO selected set of predictors. We obtain a tight correlation, with a Pearson Correlation Coefficient of 71.3% between the inferred and the observed redshifts, an average Δz_norm = 11.6 x 10^-4. We stress that notwithstanding the small sample of gamma-ray loud AGNs, we obtain a reliable predictive model using Superlearner, which is an ensemble of several machine learning models.

astro-ph.HE

VARCLUST: clustering variables using dimensionality reduction

VARCLUST algorithm is proposed for clustering variables under the assumption that variables in a given cluster are linear combinations of a small number of hidden latent variables, corrupted by the random noise. The entire clustering task is viewed as the problem of selection of the statistical model, which is defined by the number of clusters, the partition of variables into these clusters and the 'cluster dimensions', i.e. the vector of dimensions of linear subspaces spanning each of the clusters. The optimal model is selected using the approximate Bayesian criterion based on the Laplace approximations and using a non-informative uniform prior on the number of clusters. To solve the problem of the search over a huge space of possible models we propose an extension of the ClustOfVar algorithm which was dedicated to subspaces of dimension only 1, and which is similar in structure to the $K$-centroid algorithm. We provide a complete methodology with theoretical guarantees, extensive numerical experimentations, complete data analyses and implementation. Our algorithm assigns variables to appropriate clusterse based on the consistent Bayesian Information Criterion (BIC), and estimates the dimensionality of each cluster by the PEnalized SEmi-integrated Likelihood Criterion (PESEL), whose consistency we prove. Additionally, we prove that each iteration of our algorithm leads to an increase of the Laplace approximation to the model posterior probability and provide the criterion for the estimation of the number of clusters. Numerical comparisons with other algorithms show that VARCLUST may outperform some popular machine learning tools for sparse subspace clustering. We also report the results of real data analysis including TCGA breast cancer data and meteorological data. The proposed method is implemented in the publicly available R package varclust.

stat.CO

Identifying important predictors in large data bases -- multiple testing and model selection

This is a chapter of the forthcoming Handbook of Multiple Testing. We consider a variety of model selection strategies in a high-dimensional setting, where the number of potential predictors p is large compared to the number of available observations n. In particular modifications of information criteria which are suitable in case of p > n are introduced and compared with a variety of penalized likelihood methods, in particular SLOPE and SLOBE. The focus is on methods which control the FDR in terms of model identification. Theoretical results are provided both with respect to model identification and prediction and various simulation results are presented which illustrate the performance of the different methods in different situations.

stat.ME