SearcharxivSearch

arXiv subjects

Marta Dembska

Publications and source records attributed to Marta Dembska.

7 recordsLinked to original sources

Evaluation of Provenance Serialisations for Astronomical Provenance

Provenance data from astronomical pipelines are instrumental in establishing trust and reproducibility in the data processing and products. In addition, astronomers can query their provenance to answer questions routed in areas such as anomaly detection, recommendation, and prediction. The next generation of astronomical survey telescopes such as the Vera Rubin Observatory or Square Kilometre Array, are capable of producing peta to exabyte scale data, thereby amplifying the importance of even small improvements to the efficiency of provenance storage or querying. In order to determine how astronomers should store and query their provenance data, this paper reports on a comparison between the turtle and JSON provenance serialisations. The triple store Apache Jena Fuseki and the graph database system Neo4j were selected as representative database management systems (DBMS) for turtle and JSON, respectively. Simulated provenance data was uploaded to and queried over each DBMS and the metrics measured for comparison were the accuracy and timing of the queries as well as the data upload times. It was found that both serialisations are competent for this purpose, and both have similar query accuracy. The turtle provenance was found to be more efficient at storing and uploading the data. Regarding queries, for small datasets ($<$5MB) and simple information retrieval queries, the turtle serialisation was also found to be more efficient. However, queries for JSON serialised provenance were found to be more efficient for more complex queries which involved matching patterns across the DBMS, this effect scaled with the size of the queried provenance.

astro-ph.IM

Pipeline Provenance for Analysis, Evaluation, Trust or Reproducibility

Data volumes and rates of research infrastructures will continue to increase in the upcoming years and impact how we interact with their final data products. Little of the processed data can be directly investigated and most of it will be automatically processed with as little user interaction as possible. Capturing all necessary information of such processing ensures reproducibility of the final results and generates trust in the entire process. We present PRAETOR, a software suite that enables automated generation, modelling, and analysis of provenance information of Python pipelines. Furthermore, the evaluation of the pipeline performance, based upon a user defined quality matrix in the provenance, enables the first step of machine learning processes, where such information can be fed into dedicated optimisation procedures.

astro-ph.IM

Astronomical Pipeline Provenance: A Use Case Evaluation

In this decade astronomy is undergoing a paradigm shift to handle data from next generation observatories such as the Square Kilometre Array (SKA) or the Vera C. Rubin Observatory (LSST). Producing real time data streams of up to 10 TB/s and data products of the order of 600 Pbytes/year, the SKA will be the biggest civil data producing machine of the world that demands novel solutions on how these data volumes can be stored and analysed. Through the use of complex, automated pipelines the provenance of this real time data processing is key to establish confidence within the system, its final data products, and ultimately its scientific results. The intention of this paper is to lay the foundation for making an automated provenance generation tool for astronomical/data-processing pipelines. We therefore present a use case analysis, specific to the astronomical needs which addresses the issues of trust and reproducibility as well as other ulterior use cases which are of interest to astronomers. This analysis is subsequently used as the basis to discuss the requirements, challenges, and opportunities involved in designing both the tool and the associated provenance model.

astro-ph.IM

Time variation in the low frequency spectrum of Vela-like pulsar B1800-21

We report the flux measurement of the Vela like pulsar B1800-21 at the low radio frequency regime over multiple epochs spanning several years. The spectrum shows a turnover around the GHz frequency range and represents a typical example of gigahertz-peaked spectrum (GPS) pulsar. Our observations revealed that the pulsar spectrum show a significant evolution during the observing period with the low frequency part of the spectrum becoming steeper, with a higher turnover frequency, for a period of several years before reverting back to the initial shape during the latest measurements. The spectral change over times spanning several years requires dense structures, with free electron densities around 1000--20000 cm$^{-3}$ and physical dimensions ~220 AU, in the interstellar medium (ISM) traversing across the pulsar line of sight. We look into the possible sites of such structures in the ISM and likely mechanisms particularly the thermal free-free absorption as possible explanations for the change.

astro-ph.HE

Pulse broadening analysis for several new pulsars and anomalous scattering

We show the results of our analysis of the pulse broadening phenomenon in 25 pulsars at several frequencies using the data gathered with GMRT and Effelsberg radiotelescopes. Twenty two of these pulsars were not studied in that regard before and our work has increased the total number of pulsars with multi-frequency scattering measurements to almost 50, basically doubling the amount available so far. The majority of the pulsars we observed have high to very-high dispersion measures (DM>200) and our results confirm the suggestion of Loehmer et al.(2001, 2004) that the scatter time spectral indices for high-DM pulsars deviate from the value predicted by a single thin screen model with Kolmogorov's distribution of the density fluctuations. In this paper we discuss the possible explanations for such deviations.

astro-ph.HE

Binary pulsar B1259-63 spectrum evolution: detailed study

We studied the radio spectrum of PSR B1259-63 in an unique binary with Be star LS 2883 and showed that the shape of the spectrum depends on the orbital phase. We proposed a qualitative model which explains this evolution. We considered two mechanisms that might influence the observed radio emission: free-free absorption and cyclotron resonance. Recently published results have revealed a new aspect in pulsar radio spectra. There were found objects with turnover at high frequencies in spectra, called gigahertz-peaked spectra (GPS) pulsars. Most of them adjoin such interesting environments as HII regions or compact pulsar wind nebulae (PWN). Thus, it is suggested that the turnover phenomenon is associated with the environment than being related intrinsically to the radio emission mechanism. Having noticed the apparent resemblance between the B1259-63 spectrum and the GPS, we suggest that the same mechanisms should be responsible for both cases. Therefore, the case of B1259-63 can be treated as a key factor to explain the GPS phenomenon observed for the solitary pulsars with interesting environments and also another types of spectra (e.g. with break).

astro-ph.SR

Food-chain competition influences gene's size

We have analysed an effect of the Bak-Sneppen predator-prey food-chain self-organization on nucleotide content of evolving species. In our model, genomes of the species under consideration have been represented by their nucleotide genomic fraction and we have applied two-parameter Kimura model of substitutions to include the changes of the fraction in time. The initial nucleotide fraction and substitution rates were decided with the help of random number generator. Deviation of the genomic nucleotide fraction from its equilibrium value was playing the role of the fitness parameter, $B$, in Bak-Sneppen model. Our finding is, that the higher is the value of the threshold fitness, during the evolution course, the more frequent are large fluctuations in number of species with strongly differentiated nucleotide content; and it is more often the case that the oldest species, which survive the food-chain competition, might have specific nucleotide fraction making possible generating long genes

q-bio.PE