SearcharxivSearch

arXiv subjects

Ciro Donalek

Publications and source records attributed to Ciro Donalek.

At least 19 recordsLinked to original sources

XR and Hybrid Data Visualization Spaces for Enhanced Data Analytics

The growing complexity and information content of data, together with the need to understand both the complex structures, relationships, and phenomena present in these data spaces, compounded with the emerging need to understand the results produced by AI tools used to analyze the data, requires development of novel, effective data visualization tools. Much of the growing complexity is reflected in the increasing dimensionality of data spaces, where extended reality (XR) naturally emerges as a candidate to help extend our capability for higher dimensional understanding. However, humans often understand lower dimensionality representations more effectively. Still, XR offers an opportunity for a seamless integration of simulated traditional data displays within the 3-dimensional virtual data spaces, leading to more intuitive and more effective data analytics. In this paper we present an overview of the benefits of seamlessly integrated 2-dimensional and 3-dimensional interactive visual representations embedded in XR spaces, and present three case studies that leverage these approaches for more efficient data analytics.

cs.HC

Deep Co-Added Sky from Catalina Sky Survey Images

A number of synoptic sky surveys are underway or being planned. Typically they are done with small telescopes and relatively short exposure times. A search for transient or variable sources involves comparison with deeper baseline images, ideally obtained through the same telescope and camera. With that in mind we have stacked images from the 0.68~m Schmidt telescope on Mt. Bigelow taken over ten years as part of the Catalina Sky Survey. In order to generate deep reference images for the Catalina Real-time Transient Survey, close to 0.8 million images over 8000 fields and covering over 27000~sq.~deg. have gone into the deep stack that goes up to 3 magnitudes deeper than individual images. CRTS system does not use a filter in imaging, hence there is no standard passband in which the optical magnitude is measured. We estimate depth by comparing these wide-band unfiltered co-added images with images in the $g$-band and find that the image depth ranges from 22.0--24.2 across the sky, with a 200-image stack attaining an equivalent AB magnitude sensitivity of 22.8. We compared various state-of-the-art software packages for co-adding astronomical images and have used SWarp for the stacking. We describe here the details of the process adopted. This methodology may be useful in other panoramic imaging applications, and to other surveys as well. The stacked images are available through a server at Inter-University Centre for Astronomy and Astrophysics (IUCAA).

astro-ph.IM

Long-term Periodicities of Cataclysmic Variables with Synoptic Surveys

A systematic study on the long-term periodicities of known Galactic cataclysmic variables (CVs) was conducted. Among 1580 known CVs, 344 sources were matched and extracted from the Palomar Transient Factory (PTF) data repository. The PTF light curves were combined with the Catalina Real-Time Transient Survey (CRTS) light curves and analyzed. Ten targets were found to exhibit long-term periodic variability, which is not frequently observed in the CV systems. These long-term variations are possibly caused by various mechanisms, such as the precession of the accretion disk, hierarchical triple star system, magnetic field change of the companion star, and other possible mechanisms. We discuss the possible mechanisms in this study. If the long-term period is less than several tens of days, the disk precession period scenario is favored. However, the hierarchical triple star system or the variations in magnetic field strengths are most likely the predominant mechanisms for longer periods.

astro-ph.SR

Extreme Variability in a Broad Absorption Line Quasar

CRTS J084133.15+200525.8 is an optically bright quasar at z=2.345 that has shown extreme spectral variability over the past decade. Photometrically, the source had a visual magnitude of V~17.3 between 2002 and 2008. Then, over the following five years, the source slowly brightened by approximately one magnitude, to V~16.2. Only ~1 in 10,000 quasars show such extreme variability, as quantified by the extreme parameters derived for this quasar assuming a damped random walk model. A combination of archival and newly acquired spectra reveal the source to be an iron low-ionization broad absorption line (FeLoBAL) quasar with extreme changes in its absorption spectrum. Some absorption features completely disappear over the 9 years of optical spectra, while other features remain essentially unchanged. We report the first definitive redshift for this source, based on the detection of broad H-alpha in a Keck/MOSFIRE spectrum. Absorption systems separated by several 1000 km/s in velocity show coordinated weakening in the depths of their troughs as the continuum flux increases. We interpret the broad absorption line variability to be due to changes in photoionization, rather than due to motion of material along our line of sight. This source highlights one sort of rare transition object that astronomy will now be finding through dedicated time-domain surveys.

astro-ph.GA

An analysis of feature relevance in the classification of astronomical transients with machine learning methods

The exploitation of present and future synoptic (multi-band and multi-epoch) surveys requires an extensive use of automatic methods for data processing and data interpretation. In this work, using data extracted from the Catalina Real Time Transient Survey (CRTS), we investigate the classification performance of some well tested methods: Random Forest, MLPQNA (Multi Layer Perceptron with Quasi Newton Algorithm) and K-Nearest Neighbors, paying special attention to the feature selection phase. In order to do so, several classification experiments were performed. Namely: identification of cataclysmic variables, separation between galactic and extra-galactic objects and identification of supernovae.

astro-ph.IM

A systematic search for close supermassive black hole binaries in the Catalina Real-Time Transient Survey

Hierarchical assembly models predict a population of supermassive black hole (SMBH) binaries. These are not resolvable by direct imaging but may be detectable via periodic variability (or nanohertz frequency gravitational waves). Following our detection of a 5.2 year periodic signal in the quasar PG 1302-102 (Graham et al. 2015), we present a novel analysis of the optical variability of 243,500 known spectroscopically confirmed quasars using data from the Catalina Real-time Transient Survey (CRTS) to look for close (< 0.1 pc) SMBH systems. Looking for a strong Keplerian periodic signal with at least 1.5 cycles over a baseline of nine years, we find a sample of 111 candidate objects. This is in conservative agreement with theoretical predictions from models of binary SMBH populations. Simulated data sets, assuming stochastic variability, also produce no equivalent candidates implying a low likelihood of spurious detections. The periodicity seen is likely attributable to either jet precession, warped accretion disks or periodic accretion associated with a close SMBH binary system. We also consider how other SMBH binary candidates in the literature appear in CRTS data and show that none of these are equivalent to the identified objects. Finally, the distribution of objects found is consistent with that expected from a gravitational wave-driven population. This implies that circumbinary gas is present at small orbital radii and is being perturbed by the black holes. None of the sources is expected to merge within at least the next century. This study opens a new unique window to study a population of close SMBH binaries that must exist according to our current understanding of galaxy and SMBH evolution.

astro-ph.GA

A possible close supermassive black-hole binary in a quasar with optical periodicity

Quasars have long been known to be variable sources at all wavelengths. Their optical variability is stochastic, can be due to a variety of physical mechanisms, and is well-described statistically in terms of a damped random walk model. The recent availability of large collections of astronomical time series of flux measurements (light curves) offers new data sets for a systematic exploration of quasar variability. Here we report on the detection of a strong, smooth periodic signal in the optical variability of the quasar PG 1302-102 with a mean observed period of 1,884 $\pm$ 88 days. It was identified in a search for periodic variability in a data set of light curves for 247,000 known, spectroscopically confirmed quasars with a temporal baseline of $\sim9$ years. While the interpretation of this phenomenon is still uncertain, the most plausible mechanisms involve a binary system of two supermassive black holes with a subparsec separation. Such systems are an expected consequence of galaxy mergers and can provide important constraints on models of galaxy formation and evolution.

astro-ph.GA

Immersive and Collaborative Data Visualization Using Virtual Reality Platforms

Effective data visualization is a key part of the discovery process in the era of big data. It is the bridge between the quantitative content of the data and human intuition, and thus an essential component of the scientific path from data into knowledge and understanding. Visualization is also essential in the data mining process, directing the choice of the applicable algorithms, and in helping to identify and remove bad data from the analysis. However, a high complexity or a high dimensionality of modern data sets represents a critical obstacle. How do we visualize interesting structures and patterns that may exist in hyper-dimensional data spaces? A better understanding of how we can perceive and interact with multi dimensional information poses some deep questions in the field of cognition technology and human computer interaction. To this effect, we are exploring the use of immersive virtual reality platforms for scientific data visualization, both as software and inexpensive commodity hardware. These potentially powerful and innovative tools for multi dimensional data visualization can also provide an easy and natural path to a collaborative data visualization and exploration, where scientists can interact with their data and their colleagues in the same visual space. Immersion provides benefits beyond the traditional desktop visualization tools: it leads to a demonstrably better perception of a datascape geometry, more intuitive data understanding, and a better retention of the perceived relationships in the data.

cs.HC

DAMEWARE: A web cyberinfrastructure for astrophysical data mining

Astronomy is undergoing through a methodological revolution triggered by an unprecedented wealth of complex and accurate data. The new panchromatic, synoptic sky surveys require advanced tools for discovering patterns and trends hidden behind data which are both complex and of high dimensionality. We present DAMEWARE (DAta Mining & Exploration Web Application REsource): a general purpose, web-based, distributed data mining environment developed for the exploration of large datasets, and finely tuned for astronomical applications. By means of graphical user interfaces, it allows the user to perform classification, regression or clustering tasks with machine learning methods. Salient features of DAMEWARE include its capability to work on large datasets with minimal human intervention, and to deal with a wide variety of real problems such as the classification of globular clusters in the galaxy NGC1399, the evaluation of photometric redshifts and, finally, the identification of candidate Active Galactic Nuclei in multiband photometric surveys. In all these applications, DAMEWARE allowed to achieve better results than those attained with more traditional methods. With the aim of providing potential users with all needed information, in this paper we briefly describe the technological background of DAMEWARE, give a short introduction to some relevant aspects of data mining, followed by a summary of some science cases and, finally, we provide a detailed description of a template use case.

astro-ph.IM

A novel variability-based method for quasar selection: evidence for a rest frame ~54 day characteristic timescale

We compare quasar selection techniques based on their optical variability using data from the Catalina Real-time Transient Survey (CRTS). We introduce a new technique based on Slepian wavelet variance (SWV) that shows comparable or better performance to structure functions and damped random walk models but with fewer assumptions. Combining these methods with WISE mid-IR colors produces a highly efficient quasar selection technique which we have validated spectroscopically. The SWV technique also identifies characteristic timescales in a time series and we find a characteristic rest frame timescale of ~54 days, confirmed in the light curves of ~18000 quasars from CRTS, SDSS and MACHO data, and anticorrelated with absolute magnitude. This indicates a transition between a damped random walk and $P(f) \propto f^{-1/3}$ behaviours and is the first strong indication that a damped random walk model may be too simplistic to describe optical quasar variability.

astro-ph.CO

Feature Selection Strategies for Classifying High Dimensional Astronomical Data Sets

The amount of collected data in many scientific fields is increasing, all of them requiring a common task: extract knowledge from massive, multi parametric data sets, as rapidly and efficiently possible. This is especially true in astronomy where synoptic sky surveys are enabling new research frontiers in the time domain astronomy and posing several new object classification challenges in multi dimensional spaces; given the high number of parameters available for each object, feature selection is quickly becoming a crucial task in analyzing astronomical data sets. Using data sets extracted from the ongoing Catalina Real-Time Transient Surveys (CRTS) and the Kepler Mission we illustrate a variety of feature selection strategies used to identify the subsets that give the most information and the results achieved applying these techniques to three major astronomical problems.

astro-ph.IM

A comparison of period finding algorithms

This paper presents a comparison of popular period finding algorithms applied to the light curves of variable stars from the Catalina Real-time Transient Survey (CRTS), MACHO and ASAS data sets. We analyze the accuracy of the methods against magnitude, sampling rates, quoted period, quality measures (signal-to-noise and number of observations), variability, and object classes. We find that measure of dispersion-based techniques - analysis-of-variance with harmonics and conditional entropy - consistently give the best results but there are clear dependencies on object class and light curve quality. Period aliasing and identifying a period harmonic also remain significant issues. We consider the performance of the algorithms and show that a new conditional entropy-based algorithm is the most optimal in terms of completeness and speed. We also consider a simple ensemble approach and find that it performs no better than individual algorithms.

astro-ph.IM

Using conditional entropy to identify periodicity

This paper presents a new period finding method based on conditional entropy that is both efficient and accurate. We demonstrate its applicability on simulated and real data. We find that it has comparable performance to other information-based techniques with simulated data but is superior with real data, both for finding periods and just identifying periodic behaviour. In particular, it is robust against common aliasing issues found with other period-finding algorithms.

astro-ph.IM

Machine-assisted discovery of relationships in astronomy

High-volume feature-rich data sets are becoming the bread-and-butter of 21st century astronomy but present significant challenges to scientific discovery. In particular, identifying scientifically significant relationships between sets of parameters is non-trivial. Similar problems in biological and geosciences have led to the development of systems which can explore large parameter spaces and identify potentially interesting sets of associations. In this paper, we describe the application of automated discovery systems of relationships to astronomical data sets, focussing on an evolutionary programming technique and an information-theory technique. We demonstrate their use with classical astronomical relationships - the Hertzsprung-Russell diagram and the fundamental plane of elliptical galaxies. We also show how they work with the issue of binary classification which is relevant to the next generation of large synoptic sky surveys, such as LSST. We find that comparable results to more familiar techniques, such as decision trees, are achievable. Finally, we consider the reality of the relationships discovered and how this can be used for feature selection and extraction.

astro-ph.IM

The MICA Experiment: Astrophysics in Virtual Worlds

We describe the work of the Meta-Institute for Computational Astrophysics (MICA), the first professional scientific organization based in virtual worlds. MICA was an experiment in the use of this technology for science and scholarship, lasting from the early 2008 to June 2012, mainly using the Second Life and OpenSimulator as platforms. We describe its goals and activities, and our future plans. We conducted scientific collaboration meetings, professional seminars, a workshop, classroom instruction, public lectures, informal discussions and gatherings, and experiments in immersive, interactive visualization of high-dimensional scientific data. Perhaps the most successful of these was our program of popular science lectures, illustrating yet again the great potential of immersive VR as an educational and outreach platform. While the members of our research groups and some collaborators found the use of immersive VR as a professional telepresence tool to be very effective, we did not convince a broader astrophysics community to adopt it at this time, despite some efforts; we discuss some possible reasons for this non-uptake. On the whole, we conclude that immersive VR has a great potential as a scientific and educational platform, as the technology matures and becomes more broadly available and accepted.

astro-ph.IM

Data challenges of time domain astronomy

Astronomy has been at the forefront of the development of the techniques and methodologies of data intensive science for over a decade with large sky surveys and distributed efforts such as the Virtual Observatory. However, it faces a new data deluge with the next generation of synoptic sky surveys which are opening up the time domain for discovery and exploration. This brings both new scientific opportunities and fresh challenges, in terms of data rates from robotic telescopes and exponential complexity in linked data, but also for data mining algorithms used in classification and decision making. In this paper, we describe how an informatics-based approach-part of the so-called "fourth paradigm" of scientific discovery-is emerging to deal with these. We review our experiences with the Palomar-Quest and Catalina Real-Time Transient Sky Surveys; in particular, addressing the issue of the heterogeneity of data associated with transient astronomical events (and other sensor networks) and how to manage and analyze it.

astro-ph.IM

Connecting the time domain community with the Virtual Astronomical Observatory

The time domain has been identified as one of the most important areas of astronomical research for the next decade. The Virtual Observatory is in the vanguard with dedicated tools and services that enable and facilitate the discovery, dissemination and analysis of time domain data. These range in scope from rapid notifications of time-critical astronomical transients to annotating long-term variables with the latest modeling results. In this paper, we will review the prior art in these areas and focus on the capabilities that the VAO is bringing to bear in support of time domain science. In particular, we will focus on the issues involved with the heterogeneous collections of (ancillary) data associated with astronomical transients, and the time series characterization and classification tools required by the next generation of sky surveys, such as LSST and SKA.

astro-ph.IM

Extracting Knowledge From Massive Astronomical Data Sets

The exponential growth of astronomical data collected by both ground based and space borne instruments has fostered the growth of Astroinformatics: a new discipline laying at the intersection between astronomy, applied computer science, and information and computation (ICT) technologies. At the very heart of Astroinformatics is a complex set of methodologies usually called Data Mining (DM) or Knowledge Discovery in Data Bases (KDD). In the astronomical domain, DM/KDD are still in a very early usage stage, even though new methods and tools are being continuously deployed in order to cope with the Massive Data Sets (MDS) that can only grow in the future. In this paper, we briefly outline some general problems encountered when applying DM/KDD methods to astrophysical problems, and describe the DAME (DAta Mining & Exploration) web application. While specifically tailored to work on MDS, DAME can be effectively applied also to smaller data sets. As an illustration, we describe two application of DAME to two different problems: the identification of candidate globular clusters in external galaxies, and the classification of active galactic nuclei (AGN). We believe that tools and services of this nature will become increasingly necessary for the data-intensive astronomy (and indeed all sciences) in the 21st century.

astro-ph.IM