SearcharxivSearch

arXiv subjects

Jan Novotný

Publications and source records attributed to Jan Novotný.

13 recordsLinked to original sources

Dark matter halos modeled by polytropic spheres influenced by the relict cosmological constant and trapping polytropes forming supermassive black holes

We study dark matter halos modeled by general relativistic polytropic spheres in spacetimes with the repulsive cosmological constant representing vacuum energy density, governed by a polytropic index $n$ and a relativistic (cosmological) parameter $σ$ ($λ$) determining the ratio of central pressure (vacuum energy density) and central energy density of the fluid. To give mapping of the polytrope parameters for matching extension and mass of large dark matter halos, we study properties of the polytropic spheres and introduce an effective potential of the geodesic motion in their internal spacetime. Circular geodesics enable us to find the limits of the trapping polytropes with central regions containing trapped null geodesics; supermassive black holes can be formed due to the instability of the central region against gravitational perturbations. Stability of the polytropic spheres relative to radial perturbations is determined. We match extension and mass of the polytropes to those of dark matter halos related to large galaxies or galaxy clusters, with extension $100 < \ell/\mathrm{kpc} < 5000$ and gravitational mass $10^{12} < M/M_\odot < 5 \times 10^{15}$. The observed velocity profiles simulated by the phenomenological dark matter halo density profiles can be well matched also by the velocity profiles of the exact polytrope spacetimes. The matching is possible by the non-relativistic polytropes for each value of $n$, with relativistic parameter $σ\leq 10^{-4}$ and very low central energy density. Surprisingly, the matching works for ``spread'' relativistic polytropes with $n > 3.3$ and $σ\geq 0.1$ when the central density can be much larger. The trapping polytropes forming supermassive black holes must have $n > 3.8$ and $σ> 0.667$.

gr-qc

Accelerating Dedispersion using Many-Core Architectures

Astrophysical radio signals are excellent probes of extreme physical processes that emit them. However, to reach Earth, electromagnetic radiation passes through the ionised interstellar medium (ISM), introducing a frequency-dependent time delay (dispersion) to the emitted signal. Removing dispersion enables searches for transient signals like Fast Radio Bursts (FRB) or repeating signals from isolated pulsars or those in orbit around other compact objects. The sheer volume and high resolution of data that next generation radio telescopes will produce require High-Performance Computing (HPC) solutions and algorithms to be used in time-domain data processing pipelines to extract scientifically valuable results in real-time. This paper presents a state-of-the-art implementation of brute force incoherent dedispersion on NVIDIA GPUs, and on Intel and AMD CPUs. We show that our implementation is 4x faster (8-bit 8192 channels input) than other available solutions and demonstrate, using 11 existing telescopes, that our implementation is at least 20 faster than real-time. This work is part of the AstroAccelerate package.

astro-ph.IM

A Survey of Feature detection methods for localisation of plain sections of Axial Brain Magnetic Resonance Imaging

Matching MRI brain images between patients or mapping patients' MRI slices to the simulated atlas of a brain is key to the automatic registration of MRI of a brain. The ability to match MRI images would also enable such applications as indexing and searching MRI images among multiple patients or selecting images from the region of interest. In this work, we have introduced robustness, accuracy and cumulative distance metrics and methodology that allows us to compare different techniques and approaches in matching brain MRI of different patients or matching MRI brain slice to a position in the brain atlas. To that end, we have used feature detection methods AGAST, AKAZE, BRISK, GFTT, HardNet, and ORB, which are established methods in image processing, and compared them on their resistance to image degradation and their ability to match the same brain MRI slice of different patients. We have demonstrated that some of these techniques can correctly match most of the brain MRI slices of different patients. When matching is performed with the atlas of the human brain, their performance is significantly lower. The best performing feature detection method was a combination of SIFT detector and HardNet descriptor that achieved 93% accuracy in matching images with other patients and only 52% accurately matched images when compared to atlas.

eess.IV

Efficiency Near the Edge: Increasing the Energy Efficiency of FFTs on GPUs for Real-time Edge Computing

The Square Kilometre Array (SKA) is an international initiative for developing the world's largest radio telescope with a total collecting area of over a million square meters. The scale of the operation, combined with the remote location of the telescope, requires the use of energy-efficient computational algorithms. This, along with the extreme data rates that will be produced by the SKA and the requirement for a real-time observing capability, necessitates in-situ data processing in an edge style computing solution. More generally, energy efficiency in the modern computing landscape is becoming of paramount concern. Whether it be the power budget that can limit some of the world's largest supercomputers, or the limited power available to the smallest Internet-of-Things devices. In this paper, we study the impact of hardware frequency scaling on the energy consumption and execution time of the Fast Fourier Transform (FFT) on NVIDIA GPUs using the cuFFT library. The FFT is used in many areas of science and it is one of the key algorithms used in radio astronomy data processing pipelines. Through the use of frequency scaling, we show that we can lower the power consumption of the NVIDIA V100 GPU when computing the FFT by up to 60% compared to the boost clock frequency, with less than a 10% increase in the execution time. Furthermore, using one common core clock frequency for all tested FFT lengths, we show on average a 50% reduction in power consumption compared to the boost core clock frequency with an increase in the execution time still below 10%. We demonstrate how these results can be used to lower the power consumption of existing data processing pipelines. These savings, when considered over years of operation, can yield significant financial savings, but can also lead to a significant reduction of greenhouse gas emissions.

cs.PF

Implementing CUDA Streams into AstroAccelerate -- A Case Study

To be able to run tasks asynchronously on NVIDIA GPUs a programmer must explicitly implement asynchronous execution in their code using the syntax of CUDA streams. Streams allow a programmer to launch independent concurrent execution tasks, providing the ability to utilise different functional units on the GPU asynchronously. For example, it is possible to transfer the results from a previous computation performed on input data n-1, over the PCIe bus whilst computing the result for input data n, by placing different tasks in different CUDA streams. The benefit of such an approach is that the time taken for the data transfer between the host and device can be hidden with computation. This case study deals with the implementation of CUDA streams into AstroAccelerate. AstroAccelerate is a GPU accelerated real-time signal processing pipeline for time-domain radio astronomy.

astro-ph.IM

Polytropic spheres modelling dark matter halos of dwarf galaxies

Dwarf galaxies and their dark matter (DM) halos have the velocity curves of a different character than those in large galaxies. They are modelled by a simple pseudo iso-thermal model containing only two parameters that do not allow to obtain insight into physics of the DM halo. We would like to obtain some insight into the physical conditions in DM halos of dwarf galaxies by using a simple physically based model of DM halos. In order to treat a diversity of the dwarf galaxy velocity profiles in a unifying framework, we apply the polytropic spheres characterised by the polytropic index $n$ and the relativistic parameter $σ$ as a model of dwarf-galaxy DM halos and match the velocity of circular geodesics of the polytropes to the velocity curves observed in the dwarf galaxies from the LITTLE THINGS ensemble. We introduce three classes of the LITTLE THINGS dwarf galaxies in accord with the polytrope models, due to the different character of the velocity profile. The first class corresponds to polytropes having $n < 1$ with linearly increasing velocity along with the whole profile, the second class has $1 < n < 2$ and the velocity profile becomes flat in the external region, the third class has $n > 2$ and the velocity profile reaches a maximum and demonstrated a decline in the external region. The $σ$ parameter has to be strongly non-relativistic ($σ< 10^{-8}$) for all dwarf galaxy models -- it varies for the models of each class, but these variations have a negligible influence on the character of the velocity profile. Our results indicate the possibility that at least two different kinds of dark matter are behind the composition of DM halos. The matches of the observational velocity curves are of the same quality as those obtained by the pseudo-isothermal, core-like models of dwarf galaxy DM halos.

astro-ph.CO

Development of production-ready GPU data processing pipeline software for AstroAccelerate

Upcoming large scale telescope projects such as the Square Kilometre Array (SKA) will see high data rates and large data volumes; requiring tools that can analyse telescope event data quickly and accurately. In modern radio telescopes, analysis software forms a core part of the data read out, and long-term software stability and maintainability are essential. AstroAccelerate is a many core accelerated software package that uses NVIDIA(R) GPUs to perform realtime analysis of radio telescope data, and it has been shown to be substantially faster than realtime at processing simulated SKA-like data. AstroAccelerate contains optimised GPU implementations of signal processing tools used in radio astronomy including dedispersion, Fourier domain acceleration search, single pulse detection, and others. This article describes the transformation of AstroAccelerate from a C-like prototype code to a production-ready software library with a C++ API and a Python interface; while preserving compatibility with legacy software that is implemented in C. The design of the software library interfaces, refactoring aspects, and coding techniques are discussed.

astro-ph.IM

Searching for pulsars in extreme orbits -- GPU acceleration of the Fourier domain 'jerk' search

Binary pulsars are an important target for radio surveys because they present a natural laboratory for a wide range of astrophysics for example testing general relativity, including detection of gravitational waves. The orbital motion of a pulsar which is locked in a binary system causes a frequency shift (a Doppler shift) in their normally very periodic pulse emissions. These shifts cause a reduction in the sensitivity of traditional periodicity searches. To correct this smearing Ransom [2001], Ransom et al. [2002] developed the Fourier domain acceleration search (FDAS) which uses a matched filtering technique. This method is however limited to a constant pulsar acceleration. Therefore, Andersen and Ransom [2018] broadened the Fourier domain acceleration search to account also for a linear change in the acceleration by implementing the Fourier domain "jerk" search into the PRESTO software package. This extension increases the number of matched filters used significantly. We have implemented the Fourier domain "jerk" search (JERK) on GPUs using CUDA. We have achieved 90x performance increase when compared to the parallel implementation of JERK in PRESTO. This work is part of the AstroAccelerate project Armour et al. [2019], a many-core accelerated time-domain signal processing library for radio astronomy.

astro-ph.IM

Gravitational instability of polytropic spheres containing region of trapped null geodesics: a possible explanation of central supermassive black holes in galactic halos

We study behaviour of gravitational waves in the recently introduced general relativistic polytropic spheres containing a region of trapped null geodesics extended around radius of the stable null circular geodesic that can exist for the polytropic index $N>2.138$ and the relativistic parameter, giving ratio of the central pressure $p_\mathrm{c}$ to the central energy density $ρ_\mathrm{c}$, higher than $σ= 0.677$. In the trapping zones of such polytropes, the effective potential of the axial gravitational wave perturbations resembles those related to the ultracompact uniform density objects, giving thus similar long-lived axial gravitational modes. These long-lived linear perturbations are related to the stable circular null geodesic and due to additional non-linear phenomena could lead to conversion of the trapping zone to a black hole. We give in the eikonal limit examples of the long-lived gravitational modes, their oscillatory frequencies and slow damping rates, for the trapping zones of the polytropes with $N \in (2.138,4)$. However, in the trapping polytropes the long-lived damped modes exist only for very large values of the multipole number $\ell>50$, while for smaller values of $\ell$ the numerical calculations indicate existence of fast growing unstable axial gravitational modes. We demonstrate that for polytropes with $N \geq 3.78$, the trapping region is by many orders smaller than extension of the polytrope, and the mass contained in the trapping zone is about $10^{-3}$ of the total mass of the polytrope. Therefore, the gravitational instability of such trapping zones could serve as a model explaining creation of central supermassive black holes in galactic halos or galaxy clusters.

gr-qc

Polytropic spheres containing regions of trapped null geodesics

We demonstrate that in the framework of standard general relativity polytropic spheres with properly fixed polytropic index $n$ and relativistic parameter $σ$, giving ratio of the central pressure $p_\mathrm{c}$ to the central energy density $ρ_\mathrm{c}$, can contain region of trapped null geodesics. Such trapping polytropes can exist for $n > 2.138$ and they are generally much more extended and massive than the observed neutron stars. We show that in the $n$--$σ$ parameter space the region of allowed trapping increase with polytropic index for interval of physical interest $2.138 < n < 4$. Space extension of the region of trapped null geodesics increases with both increasing $n$ and $σ> 0.677$ from the allowed region. In order to relate the trapping phenomenon to astrophysically relevant situations, we restrict validity of the polytropic configurations to their extension $r_\mathrm{extr}$ corresponding to the gravitational mass $M \sim 2M_{\odot}$ of the most massive observed neutron stars. Then for the central density $ρ_\mathrm{c} \sim 10^{15}$~g\,cm$^{-3}$ the trapped regions are outside $r_\mathrm{extr}$ for all values of $2.138 < n < 4$, for the central density $ρ_\mathrm{c} \sim 5 \times 10^{15}$~g\,cm$^{-3}$ the whole trapped regions are located inside of $r_\mathrm{extr}$ for $2.138 < n < 3.1$, while for $ρ_\mathrm{c} \sim 10^{16}$~g\,cm$^{-3}$ the whole trapped regions are inside of $r_\mathrm{extr}$ for all values of $2.138 < n < 4$, guaranteeing astrophysically plausible trapping for all considered polytropes. The region of trapped null geodesics is located closely to the polytrope centre and could have relevant influence on cooling of such polytropes or for binding of gravitational waves in their interior.

gr-qc

General relativistic polytropes with a repulsive cosmological constant

Spherically symmetric equilibrium configurations of perfect fluid obeying a polytropic equation of state are studied in spacetimes with a repulsive cosmological constant. The configurations are specified in terms of three parameters---the polytropic index $n$, the ratio of central pressure and central energy density of matter $σ$, and the ratio of energy density of vacuum and central density of matter $λ$. The static equilibrium configurations are determined by two coupled first-order nonlinear differential equations that are solved by numerical methods with the exception of polytropes with $n=0$ corresponding to the configurations with a uniform distribution of energy density, when the solution is given in terms of elementary functions. The geometry of the polytropes is conveniently represented by embedding diagrams of both the ordinary space geometry and the optical reference geometry reflecting some dynamical properties of the geodesic motion. The polytropes are represented by radial profiles of energy density, pressure, mass, and metric coefficients. For all tested values of $n>0$, the static equilibrium configurations with fixed parameters $n$, $σ$, are allowed only up to a critical value of the cosmological parameter $λ_{\mathrm{c}}=λ_{\mathrm{c}}(n,σ)$. In the case of $n>3$, the critical value $λ_{\mathrm{c}}$ tends to zero for special values of $σ$. The gravitational potential energy and the binding energy of the polytropes are determined and studied by numerical methods. We discuss in detail the polytropes with an extension comparable to those of the dark matter halos related to galaxies, i.e., with extension $\ell > 100\,\mathrm{kpc}$ and mass $M > 10^{12}\,\mathrm{M}_{\odot}$. ...

gr-qc

A polyphase filter for many-core architectures

In this article we discuss our implementation of a polyphase filter for real-time data processing in radio astronomy. We describe in detail our implementation of the polyphase filter algorithm and its behaviour on three generations of NVIDIA GPU cards, on dual Intel Xeon CPUs and the Intel Xeon Phi (Knights Corner) platforms. All of our implementations aim to exploit the potential for data reuse that the algorithm offers. Our GPU implementations explore two different methods for achieving this, the first makes use of L1/Texture cache, the second uses shared memory. We discuss the usability of each of our implementations along with their behaviours. We measure performance in execution time, which is a critical factor for real-time systems, we also present results in terms of bandwidth (GB/s), compute (GFlop/s) and type conversions (GTc/s). We include a presentation of our results in terms of the sample rate which can be processed in real-time by a chosen platform, which more intuitively describes the expected performance in a signal processing setting. Our findings show that, for the GPUs considered, the performance of our polyphase filter when using lower precision input data is limited by type conversions rather than device bandwidth. We compare these results to an implementation on the Xeon Phi. We show that our Xeon Phi implementation has a performance that is 1.47x to 1.95x greater than our CPU implementation, however is not insufficient to compete with the performance of GPUs. We conclude with a comparison of our best performing code to two other implementations of the polyphase filter, showing that our implementation is faster in nearly all cases. This work forms part of the Astro-Accelerate project, a many-core accelerated real-time data processing library for digital signal processing of time-domain radio astronomy data.

astro-ph.IM

The Implementation of a Real-Time Polyphase Filter

In this article we study the suitability of dierent computational accelerators for the task of real-time data processing. The algorithm used for comparison is the polyphase filter, a standard tool in signal processing and a well established algorithm. We measure performance in FLOPs and execution time, which is a critical factor for real-time systems. For our real-time studies we have chosen a data rate of 6.5GB/s, which is the estimated data rate for a single channel on the SKAs Low Frequency Aperture Array. Our findings how that GPUs are the most likely candidate for real-time data processing. GPUs are better in both performance and power consumption.

cs.DC