SearcharxivSearch

arXiv subjects

Jessi Cisewski-Kehe

Publications and source records attributed to Jessi Cisewski-Kehe.

At least 19 recordsLinked to original sources

Improving the Precision of Line-by-Line Radial Velocities: A Data-Driven Iterative Algorithm for Spectral Line Selection

Independent analysis of individual spectral lines, or line-by-line (LBL) analyses, can improve upon standard cross-correlation function (CCF) methods for measuring radial velocities (RVs) because they preserve critical information about individual line shape changes that can be caused by stellar activity. In this work, we measure LBL RVs of 3,830 spectral lines across 383 days of NEID solar observations. Our LBL approach achieves an RV RMS of $2.012~\mathrm{m\,s^{-1}}$, which is slightly lower than the $2.129~\mathrm{m\,s^{-1}}$ achieved by a CCF approach using a shared line list. Then, we describe and benchmark several methods for selecting line lists based on line properties such as depth and intrinsic RV scatter. We find that these subsets have a lower RV RMS compared to either the full line list or random subsets of equal size. Motivated by these results, we present FLARES (Filtering Lines for Accurate Radial-velocity Exoplanet Search), an iterative line-selection algorithm. FLARES selects candidate spectral lines with extreme values of multiple line metrics and properties such as depth, signal-to-noise ratio, and detector position, and preferentially rejects lines whose removal produces the largest decrease in the weighted RV scatter. FLARES achieves an RV RMS of $1.122~\mathrm{m\,s^{-1}}$ using just 24 lines and performs better than the benchmark methods. We perform Monte Carlo simulations and show FLARES is robust and reproducible. Comparisons to alternative line lists chosen to have properties similar to the best FLARES-selected lines demonstrate that FLARES is successfully identifying line properties that lead to effective line lists for future extreme-precision RV measurements.

astro-ph.EP

GJ 523b is a Massive, 170 Myr-old Mega-Earth, Likely on a Polar Orbit

We use WIYN/NEID radial velocity measurements to confirm the planetary nature and measure the mass of the TESS transiting exoplanet candidate around the mid-K dwarf GJ 523 ($V=9.23$, $K=6.525$). We find that GJ 523b is on a 17.75 day orbit and has a radius of $2.55\pm0.15\,R_\oplus$, a mass of $23.5\pm3.3\,M_\oplus$, and a zero-albedo equilibrium temperature of 538 K. GJ 523b's high bulk density of $7.8\pm1.8$ g cm$^{-3}$ and position on a mass-radius diagram implies a surprising low atmospheric mass fraction despite its relatively large mass. Additionally, we determine that the system has an age of $169^{+100}_{-48}$ Myr through a gyrochronological analysis of GJ 523 and its comoving companions. We also use the SED-derived stellar radius, the photometric rotation period, and the spectroscopic $v\sin i_\star$ to derive a stellar inclination of $17.6\pm5.0$ degrees, implying that GJ 523b has a minimum orbital obliquity of $71.4_{-5.0}^{+4.7}$ degrees. GJ 523b's high mass, apparent lack of a gas envelope, young age, and high orbital obliquity present a challenge to typical planet formation pathways, and at the moment there is not enough data on the system to definitively determine how GJ 523b formed. Finally, we present a new observational classification for ultra-dense, sub-Neptune-sized exoplanets similar to GJ 523b: the mega-Earths, planets with $R_p \geq2.1\,R_\oplus$ and $\rho_p \geq 5.5$ g cm$^{-3}$.

astro-ph.EP

Tensor Computation of Euler Characteristic Functions and Transforms

The weighted Euler characteristic transform (WECT) and Euler characteristic function (ECF) have proven to be useful tools in a variety of applications. However, current methods for computing these functions are either not optimized for GPU computation or do not scale to higher-dimensional settings. In this work, we present a tensor-based framework for computing such topological descriptors which is highly optimized for GPU architectures and works in full generality across simplicial and cubical complexes of arbitrary dimension. Experimentally, the framework demonstrates significant speedups over existing methods when computing the WECT and ECF across a variety of two- and three-dimensional datasets. Computation of these transforms is implemented in a publicly available Python package called pyECT.

cs.CG

Tracking Temporal Evolution of Topological Features in Image Data

Topological Data Analysis (TDA) can be used to detect and characterize holes in an image, such as zero-dimensional holes (connected components) or one-dimensional holes (loops). However, there is currently no widely accepted statistical framework for modeling spatiotemporal dependence in the evolution of topological features, such as holes, within a time series of images. We propose a hypothesis testing framework to identify statistically significant topological features of images in space and time, simultaneously. This addition of time may induce higher-dimensional topological features which can be used to establish temporal connections between the lower-dimensional features at each point in time. The temporal evolution of these lower-dimensional features is then represented on a zigzag persistence diagram, as a topological summary statistic focused on time dynamics. We demonstrate that the method effectively captures the emergence and progression of topological features in a study of a series of images of a wounded cell as it repairs. The proposed method outperforms a current approach in a simulation study that includes features of the wound healing process. Since, the wounded cell images exhibit nonlinear, dynamic, spatial, and temporal structures during single-cell repair, they provide a good application for this method.

stat.ME

MaxTDA: Robust Statistical Inference for Maximal Persistence in Topological Data Analysis

Persistent homology is an area within topological data analysis (TDA) that can uncover different dimensional holes (connected components, loops, voids, etc.) in data. The holes are characterized, in part, by how long they persist across different scales. Noisy data can result in many additional holes that are not true topological signal. Various robust TDA techniques have been proposed to reduce the number of noisy holes, however, these robust methods have a tendency to also reduce the topological signal. This work introduces Maximal TDA (MaxTDA), a statistical framework addressing a limitation in TDA wherein robust inference techniques systematically underestimate the persistence of significant homological features. MaxTDA combines kernel density estimation with level-set thresholding via rejection sampling to generate consistent estimators for the maximal persistence features that minimizes bias while maintaining robustness to noise and outliers. We establish the consistency of the sampling procedure and the stability of the maximal persistence estimator. The framework also enables statistical inference on topological features through rejection bands, constructed from quantiles that bound the estimator's deviation probability. MaxTDA is particularly valuable in applications where precise quantification of statistically significant topological features is essential for revealing underlying structural properties in complex datasets. Numerical simulations across varied datasets, including an example from exoplanet astronomy, highlight the effectiveness of MaxTDA in recovering true topological signals.

stat.ME

Searching for Low-Mass Exoplanets Amid Stellar Variability with a Fixed Effects Linear Model of Line-by-Line Shape Changes

The radial velocity (RV) method, also known as Doppler spectroscopy, is a powerful technique for exoplanet discovery and characterization. In recent years, progress has been made thanks to the improvements in the quality of spectra from new extreme precision RV spectrometers. However, detecting the RV signals of Earth-like exoplanets remains challenging, as the spectroscopic signatures of low-mass planets can be obscured or confused with intrinsic stellar variability. Changes in the shapes of spectral lines across time can provide valuable information for disentangling stellar activity from true Doppler shifts caused by low-mass exoplanets. In this work, we present a fixed effects linear model to estimate RV signals that controls for changes in line shapes by aggregating information from hundreds of spectral lines. Our methodology incorporates a wild-bootstrap approach for modeling uncertainty and cross-validation to control for overfitting. We evaluate the model's ability to remove stellar activity using solar observations from the NEID spectrograph, as the sun's true center-of-mass motion is precisely known. Including line shape-change covariates reduces the RV root-mean-square errors by approximately 70% (from 1.919 m s$^{-1}$ to 0.575 m s$^{-1}$) relative to using only the line-by-line Doppler shifts. The magnitude of the residuals is significantly less than that from traditional CCF-based RV estimators and comparable to other state-of-the-art methods for mitigating stellar variability.

astro-ph.EP

A Subsequence Approach to Topological Data Analysis for Irregularly-Spaced Time Series

A time-delay embedding (TDE), grounded in the framework of Takens's Theorem, provides a mechanism to represent and analyze the inherent dynamics of time-series data. Recently, topological data analysis (TDA) methods have been applied to study this time series representation mainly through the lens of persistent homology. Current literature on the fusion of TDE and TDA are adept at analyzing uniformly-spaced time series observations. This work introduces a novel {\em subsequence} embedding method for irregularly-spaced time-series data. We show that this method preserves the original state space topology while reducing spurious homological features. Theoretical stability results and convergence properties of the proposed method in the presence of noise and varying levels of irregularity in the spacing of the time series are established. Numerical studies and an application to real data illustrates the performance of the proposed method.

stat.ME

A Divide-and-Conquer Approach to Persistent Homology

Persistent homology is a tool of topological data analysis that has been used in a variety of settings to characterize different dimensional holes in data. However, persistent homology computations can be memory intensive with a computational complexity that does not scale well as the data size becomes large. In this work, we propose a divide-and-conquer (DaC) method to mitigate these issues. The proposed algorithm efficiently finds small, medium, and large-scale holes by partitioning data into sub-regions and uses a Vietoris-Rips filtration. Furthermore, we provide theoretical results that quantify the bottleneck distance between DaC and the true persistence diagram and the recovery probability of holes in the data. We empirically verify that the rate coincides with our theoretical rate, and find that the memory and computational complexity of DaC outperforms an alternative method that relies on a clustering preprocessing step to reduce the memory and computational complexity of the persistent homology computations. Finally, we test our algorithm using spatial data of the locations of lakes in Wisconsin, where the classical persistent homology is computationally infeasible.

math.AT

High-energy Neutrino Source Cross-correlations with Nearest-neighbor Distributions

The astrophysical origins of the majority of the IceCube neutrinos remain unknown. Effectively characterizing the spatial distribution of the neutrino samples and associating the events with astrophysical source catalogs can be challenging given the large atmospheric neutrino background and underlying non-Gaussian spatial features in the neutrino and source samples. In this paper, we investigate a framework for identifying and statistically evaluating the cross-correlations between IceCube data and an astrophysical source catalog based on the $k$-nearest-neighbor cumulative distribution functions ($k$NN-CDFs). We propose a maximum likelihood estimation procedure for inferring the true proportions of astrophysical neutrinos in the point-source data. We conduct a statistical power analysis of an associated likelihood ratio test with estimations of its sensitivity and discovery potential with synthetic neutrino data samples and a WISE-2MASS galaxy sample. We apply the method to IceCube's public ten-year point-source data and find no statistically significant evidence for spatial cross-correlations with the selected galaxy sample. We discuss possible extensions to the current method and explore the method's potential to identify the cross-correlation signals in data sets with different sample sizes.

astro-ph.HE

Confidence regions for a persistence diagram of a single image with one or more loops

Topological data analysis (TDA) uses persistent homology to quantify loops and higher-dimensional holes in data, making it particularly relevant for examining the characteristics of images of cells in the field of cell biology. In the context of a cell injury, as time progresses, a wound in the form of a ring emerges in the cell image and then gradually vanishes. Performing statistical inference on this ring-like pattern in a single image is challenging due to the absence of repeated samples. In this paper, we develop a novel framework leveraging TDA to estimate underlying structures within individual images and quantify associated uncertainties through confidence regions. Our proposed method partitions the image into the background and the damaged cell regions. Then pixels within the affected cell region are used to establish confidence regions in the space of persistence diagrams (topological summary statistics). The method establishes estimates on the persistence diagrams which correct the bias of traditional TDA approaches. A simulation study is conducted to evaluate the coverage probabilities of the proposed confidence regions in comparison to an alternative approach is proposed in this paper. We also illustrate our methodology by a real-world example provided by cell repair.

stat.ME

The Weighted Euler Characteristic Transform for Image Shape Classification

The weighted Euler characteristic transform (WECT) is a new tool for extracting shape information from data equipped with a weight function. Image data may benefit from the WECT where the intensity of the pixels are used to define the weight function. In this work, an empirical assessment of the WECT's ability to distinguish shapes on images with different pixel intensity distributions is considered, along with visualization techniques to improve the intuition and understanding of what is captured by the WECT. Additionally, the expected weighted Euler characteristic and the expected WECT are derived.

cs.CG

Sidestepping the inversion of the weak-lensing covariance matrix with Approximate Bayesian Computation

Weak gravitational lensing is one of the few direct methods to map the dark-matter distribution on large scales in the Universe, and to estimate cosmological parameters. We study a Bayesian inference problem where the data covariance $\mathbf{C}$, estimated from a number $n_{\textrm{s}}$ of numerical simulations, is singular. In a cosmological context of large-scale structure observations, the creation of a large number of such $N$-body simulations is often prohibitively expensive. Inference based on a likelihood function often includes a precision matrix, $Ψ= \mathbf{C}^{-1}$. The covariance matrix corresponding to a $p$-dimensional data vector is singular for $p \ge n_{\textrm{s}}$, in which case the precision matrix is unavailable. We propose the likelihood-free inference method Approximate Bayesian Computation (ABC) as a solution that circumvents the inversion of the singular covariance matrix. We present examples of increasing degree of complexity, culminating in a realistic cosmological scenario of the determination of the weak-gravitational lensing power spectrum for the upcoming European Space Agency satellite Euclid. While we found the ABC parameter estimate variances to be mildly larger compared to likelihood-based approaches, which are restricted to settings with $p < n_{\textrm{s}}$, we obtain unbiased parameter estimates with ABC even in extreme cases where $p / n_{\textrm{s}} \gg 1$. The code has been made publicly available to ensure the reproducibility of the results.

astro-ph.CO

Practical Guidance for Bayesian Inference in Astronomy

In the last two decades, Bayesian inference has become commonplace in astronomy. At the same time, the choice of algorithms, terminology, notation, and interpretation of Bayesian inference varies from one sub-field of astronomy to the next, which can lead to confusion to both those learning and those familiar with Bayesian statistics. Moreover, the choice varies between the astronomy and statistics literature, too. In this paper, our goal is two-fold: (1) provide a reference that consolidates and clarifies terminology and notation across disciplines, and (2) outline practical guidance for Bayesian inference in astronomy. Highlighting both the astronomy and statistics literature, we cover topics such as notation, specification of the likelihood and prior distributions, inference using the posterior distribution, and posterior predictive checking. It is not our intention to introduce the entire field of Bayesian data analysis -- rather, we present a series of useful practices for astronomers who already have an understanding of the Bayesian "nuts and bolts" and wish to increase their expertise and extend their knowledge. Moreover, as the field of astrostatistics and astroinformatics continues to grow, we hope this paper will serve as both a helpful reference and as a jumping off point for deeper dives into the statistics and astrostatistics literature.

astro-ph.IM

Differentiating small-scale subhalo distributions in CDM and WDM models using persistent homology

The spatial distribution of galaxies at sufficiently small scales will encode information about the identity of the dark matter. We develop a novel description of the halo distribution using persistent homology summaries, in which collections of points are decomposed into clusters, loops and voids. We apply these methods, together with a set of hypothesis tests, to dark matter haloes in MW-analog environment regions of the cold dark matter (CDM) and warm dark matter (WDM) Copernicus Complexio $N$-body cosmological simulations. The results of the hypothesis tests find statistically significant differences (p-values $\leq$ 0.001) between the CDM and WDM structures, and the functional summaries of persistence diagrams detect differences at scales that are distinct from the comparison spatial point process functional summaries considered (including the two-point correlation function). The differences between the models are driven most strongly at filtration scales $\sim100$~kpc, where CDM generates larger numbers of unconnected halo clusters while WDM instead generates loops. This study was conducted on dark matter haloes generally; future work will involve applying the same methods to realistic galaxy catalogues.

astro-ph.IM

The EXPRES Stellar Signals Project II. State of the Field in Disentangling Photospheric Velocities

Measured spectral shifts due to intrinsic stellar variability (e.g., pulsations, granulation) and activity (e.g., spots, plages) are the largest source of error for extreme precision radial velocity (EPRV) exoplanet detection. Several methods are designed to disentangle stellar signals from true center-of-mass shifts due to planets. The EXPRES Stellar Signals Project (ESSP) presents a self-consistent comparison of 22 different methods tested on the same extreme-precision spectroscopic data from EXPRES. Methods derived new activity indicators, constructed models for mapping an indicator to the needed RV correction, or separated out shape- and shift-driven RV components. Since no ground truth is known when using real data, relative method performance is assessed using the total and nightly scatter of returned RVs and agreement between the results of different methods. Nearly all submitted methods return a lower RV RMS than classic linear decorrelation, but no method is yet consistently reducing the RV RMS to sub-meter-per-second levels. There is a concerning lack of agreement between the RVs returned by different methods. These results suggest that continued progress in this field necessitates increased interpretability of methods, high-cadence data to capture stellar signals at all timescales, and continued tests like the ESSP using consistent data sets with more advanced metrics for method performance. Future comparisons should make use of various well-characterized data sets -- such as solar data or data with known injected planetary and/or stellar signals -- to better understand method performance and whether planetary signals are preserved.

astro-ph.EP

A Stellar Activity F-statistic for Exoplanet Surveys (SAFE)

In the search for planets orbiting distant stars the presence of stellar activity in the atmospheres of observed stars can obscure the radial velocity signal used to detect such planets. Furthermore, this stellar activity contamination is set by the star itself and cannot simply be avoided with better instrumentation. Various stellar activity indicators have been developed that may correlate with this contamination. We introduce a new stellar activity indicator called the Stellar Activity F-statistic for Exoplanet surveys (SAFE) that has higher statistical power (i.e., probability of detecting a true stellar activity signal) than many traditional stellar activity indicators in a simulation study of an active region on a Sun-like star with moderate to high signal-to-noise. Also through simulation, the SAFE is demonstrated to be associated with the projected area on the visible side of the star covered by active regions. We also demonstrate that the SAFE detects statistically significant stellar activity in most of the spectra for HD 22049, a star known to have high stellar variability. Additionally, the SAFE is calculated for recent observations of the three low-variability stars HD 34411, HD 10700, and HD 3651, the latter of which is known to have a planetary companion. As expected, the SAFE for these three only occasionally detects activity. Furthermore, initial exploration appears to indicate that the SAFE may be useful for disentangling stellar activity signals from planet-induced Doppler shifts.

astro-ph.EP

The Role of Machine Learning in the Next Decade of Cosmology

In recent years, machine learning (ML) methods have remarkably improved how cosmologists can interpret data. The next decade will bring new opportunities for data-driven cosmological discovery, but will also present new challenges for adopting ML methodologies and understanding the results. ML could transform our field, but this transformation will require the astronomy community to both foster and promote interdisciplinary research endeavors.

astro-ph.IM

Approximate Bayesian Computation for Finite Mixture Models

Finite mixture models are used in statistics and other disciplines, but inference for mixture models is challenging due, in part, to the multimodality of the likelihood function and the so-called label switching problem. We propose extensions of the Approximate Bayesian Computation-Population Monte Carlo (ABC-PMC) algorithm as an alternative framework for inference on finite mixture models. There are several decisions to make when implementing an ABC-PMC algorithm for finite mixture models, including the selection of the kernels used for moving the particles through the iterations, how to address the label switching problem, and the choice of informative summary statistics. Examples are presented to demonstrate the performance of the proposed ABC-PMC algorithm for mixture modeling. The performance of the proposed method is evaluated in a simulation study and for the popular recessional velocity galaxy data.

stat.ME