SearcharxivSearch

arXiv subjects

S. Plaszczynski

Publications and source records attributed to S. Plaszczynski.

At least 19 recordsLinked to original sources

A stochastic model of discussion

We consider the duration of discussions in face-to-face contacts and propose a stochastic model to describe it. It is based on the points of a Levy flight where the duration of each contact corresponds to the size of the clusters produced during the walk. When confronting it to the data measured from proximity sensors, we show that several datasets obtained in different environments, are precisely reproduced by the model fixing a single parameter, the Levy index, to 1.15. We analyze the dynamics of the cluster formation during the walk and compute analytically the cluster size distribution. We find that discussions are first driven by a maximum-entropy geometric distribution and then by a rich-get-richer mechanism reminiscent of preferential-attachment (the more a discussion lasts, the more it is likely to continue). In this model, conversations may be viewed as an aggregation process with a characteristic scale fixed by the mean interaction time between the two individuals.

physics.soc-ph

Levy geometric graphs

We present a new family of graphs with remarkable properties. They are obtained by connecting the points of a random walk when their distance is smaller than a given scale. Their degree (number of neighbors) does not depend on the graph's size but only on the considered scale. It follows a Gamma distribution and thus presents an exponential decay. Levy flights are particular random walks with some power-law increments of infinite variance. When building the geometric graphs from them, we show from dimensional arguments, that the number of connected components (clusters) follows an inverse power of the scale. The distribution of the size of their components, properly normalized, is scale-invariant, which reflects the self-similar nature of the underlying process. This allows to test if a graph (including non-spatial ones) could possibly result from an underlying Levy process. When the scale increases, these graphs never tend towards a single cluster, the giant component. In other words, while the autocorrelation of the process scales as a power of the distance, they never undergo a phase transition of percolation type. The Levy graphs may find applications in community detection and in the analysis of collective behaviors as in face-to-face interaction networks.

cond-mat.stat-mech

Scaling pair count to next galaxy surveys

Counting pairs of galaxies or stars according to their distance is at the core of real-space correlation analyzes performed in astrophysics and cosmology. Upcoming galaxy surveys (LSST, Euclid) will measure properties of billions of galaxies challenging our ability to perform such counting in a minute-scale time relevant for the usage of simulations. The problem is only limited by efficient access to the data, hence belongs to the big data category. We use the popular Apache Spark framework to address it and design an efficient high-throughput algorithm to deal with hundreds of millions to billions of input data. To optimize it, we revisit the question of nonhierarchical sphere pixelization based on cube symmetries and develop a new one dubbed the "Similar Radius Sphere Pixelization" (SARSPix) with very close to square pixels. It provides the most adapted indexing over the sphere for all distance-related computations. Using LSST-like fast simulations, we compute autocorrelation functions on tomographic bins containing between a hundred million to one billion data points. In each case we achieve the construction of a standard pair-distance histogram in about 2 minutes, using a simple algorithm that is shown to scale, over a moderate number of nodes (16 to 64). This illustrates the potential of this new techniques in the field of astronomy where data access is becoming the main bottleneck. They can be easily adapted to other use-cases as nearest-neighbors search, catalog cross-match or cluster finding. The software is publicly available from https://github.com/astrolabsoftware/SparkCorr.

astro-ph.IM

Analyzing billion-objects catalog interactively: Apache Spark for physicists

Apache Spark is a Big Data framework for working on large distributed datasets. Although widely used in the industry, it remains rather limited in the academic community or often restricted to software engineers. The goal of this paper is to show with practical uses-cases that the technology is mature enough to be used without excessive programming skills by astronomers or cosmologists in order to perform standard analyses over large datasets, as those originating from future galaxy surveys. To demonstrate it, we start from a realistic simulation corresponding to 10 years of LSST data taking (6 billions of galaxies). Then, we design, optimize and benchmark a set of Spark python algorithms in order to perform standard operations as adding photometric redshift errors, measuring the selection function or computing power spectra over tomographic bins. Most of the commands execute on the full 110 GB dataset within tens of seconds and can therefore be performed interactively in order to design full-scale cosmological analyses. A jupyter notebook summarizing the analysis is available at https://github.com/astrolabsoftware/1807.03078.

astro-ph.IM

A direct method to compute the galaxy count angular correlation function including redshift-space distortions

In the near future, cosmology will enter the wide and deep galaxy survey area allowing high-precision studies of the large scale structure of the universe in three dimensions. To test cosmological models and determine their parameters accurately, it is natural to confront data with exact theoretical expectations expressed in the observational parameter space (angles and redshift). The data-driven galaxy number count fluctuations on redshift shells, can be used to build correlation functions $C(θ; z_1, z_2)$ on and between shells which can probe the baryonic acoustic oscillations, the distance-redshift distortions as well as gravitational lensing and other relativistic effects. Transforming the model to the data space usually requires the computation of the angular power spectrum $C_\ell(z_1, z_2)$ but this appears as an artificial and inefficient step plagued by apodization issues. In this article we show that it is not necessary and present a compact expression for $C(θ; z_1, z_2)$ that includes directly the leading density and redshift space distortions terms from the full linear theory. It can be evaluated using a fast integration method based on Clenshaw-Curtis quadrature and Chebyshev polynomial series. This new method to compute the correlation functions without any Limber approximation, allows us to produce and discuss maps of the correlation function directly in the observable space and is a significant step towards disentangling the data from the tested models.

astro-ph.CO

Cosmological constraints on the neutrino mass including systematic uncertainties

When combining cosmological and oscillations results to constrain the neutrino sector, the question of the propagation of systematic uncertainties is often raised. We address this issue in the context of the derivation of an upper bound on the sum of the neutrino masses ($Σm_ν$) with recent cosmological data. This work is performed within the ${\mathrm{Λ{CDM}}}$ model extended to $Σm_ν$, for which we advocate the use of three mass-degenerate neutrinos. We focus on the study of systematic uncertainties linked to the foregrounds modelling in CMB data analysis, and on the impact of the present knowledge of the reionisation optical depth. This is done through the use of different likelihoods built from Planck data. Limits on $Σm_ν$ are derived with various combinations of data, including the latest Baryon Acoustic Oscillations (BAO) and Type Ia Supernovae (SN) results. We also discuss the impact of the preference for current CMB data for amplitudes of the gravitational lensing distortions higher than expected within the ${\mathrm{Λ{CDM}}}$ model, and add the Planck CMB lensing. We then derive a robust upper limit: $Σm_ν< 0.17\hbox{ eV at }95\% \hbox{CL}$, including 0.01 eV of foreground systematics. We also discuss the neutrino mass repartition and show that today's data do not allow one to disentangle normal from inverted hierarchy. The impact on the other cosmological parameters is also reported, for different assumptions on the neutrino mass repartition, and different high and low multipole CMB likelihoods.

astro-ph.CO

Angpow: a software for the fast computation of accurate tomographic power spectra

The statistical distribution of galaxies is a powerful probe to constrain cosmological models and gravity. In particular the matter power spectrum $P(k)$ brings information about the cosmological distance evolution and the galaxy clustering together. However the building of $P(k)$ from galaxy catalogues needs a cosmological model to convert angles on the sky and redshifts into distances, which leads to difficulties when comparing data with predicted $P(k)$ from other cosmological models, and for photometric surveys like LSST. The angular power spectrum $C_\ell(z_1,z_2)$ between two bins located at redshift $z_1$ and $z_2$ contains the same information than the matter power spectrum, is free from any cosmological assumption, but the prediction of $C_\ell(z_1,z_2)$ from $P(k)$ is a costly computation when performed exactly. The Angpow software aims at computing quickly and accurately the auto ($z_1=z_2$) and cross ($z_1 \neq z_2$) angular power spectra between redshift bins. We describe the developed algorithm, based on developments on the Chebyshev polynomial basis and on the Clenshaw-Curtis quadrature method. We validate the results with other codes, and benchmark the performance. Angpow is flexible and can handle any user defined power spectra, transfer functions, and redshift selection windows. The code is fast enough to be embedded inside programs exploring large cosmological parameter spaces through the $C_\ell(z_1,z_2)$ comparison with data. We emphasize that the Limber's approximation, often used to fasten the computation, gives wrong $C_\ell$ values for cross-correlations.

astro-ph.CO

Cosmology with the CMB temperature-polarization correlation

We demonstrate that the cosmic microwave background (CMB) temperature-polarization cross-correlation provides accurate and robust constraints on cosmological parameters. We compare them with the results from temperature or polarization and investigate the impact of foregrounds, cosmic variance, and instrumental noise. This analysis makes use of the Planck high-multipole HiLLiPOP likelihood based on angular power spectra, which takes into account systematics from the instrument and foreground residuals directly modelled using Planck measurements. The temperature-polarization correlation (TE) spectrum is less contaminated by astrophysical emissions than the temperature power spectrum (TT), allowing constraints that are less sensitive to foreground uncertainties to be derived. For ΛCDM parameters, TE gives very competitive results compared to TT. For basic ΛCDM model extensions (such as AL, Σmν, or Neff ), it is still limited by the instrumental noise level in the polarization maps.

astro-ph.CO

Relieving tensions related to the lensing of CMB temperature power spectra

The angular power spectra of the cosmic microwave background (CMB) temperature anisotropies reconstructed from Planck data seem to present too much gravitational lensing distortion. This is quantified by the control parameter $A_L$ that should be compatible with unity for a standard cosmology. With the Class Boltzmann solver and the profile-likelihood method, for this parameter we measure a 2.6$σ$ shift from 1 using the Planck public likelihoods. We show that, owing to strong correlations with the reionization optical depth $τ$ and the primordial perturbation amplitude $A_s$, a $\sim2σ$ tension on $τ$ also appears between the results obtained with the low ($\ell\leq 30$) and high ($30<\ell\lesssim 2500$) multipoles likelihoods. With Hillipop, another high-$\ell$ likelihood built from Planck data, this difference is lowered to $1.3σ$. In this case, the $A_L$ value is still in disagreement with unity by $2.2σ$, suggesting a non-trivial effect of the correlations between cosmological and nuisance parameters. To better constrain the nuisance foregrounds parameters, we include the very high $\ell$ measurements of the Atacama Cosmology Telescope (ACT) and South Pole Telescope (SPT) experiments and obtain $A_L = 1.03 \pm 0.08$. The Hillipop+ACT+SPT likelihood estimate of the optical depth is $τ=0.052\pm{0.035,}$ which is now fully compatible with the low $\ell$ likelihood determination. After showing the robustness of our results with various combinations, we investigate the reasons for this improvement that results from a better determination of the whole set of foregrounds parameters. We finally provide estimates of the $Λ$CDM parameters with our combined CMB data likelihood.

astro-ph.CO

Agnostic cosmology in the CAMEL framework

Cosmological parameter estimation is traditionally performed in the Bayesian context. By adopting an "agnostic" statistical point of view, we show the interest of confronting the Bayesian results to a frequentist approach based on profile-likelihoods. To this purpose, we have developed the Cosmological Analysis with a Minuit Exploration of the Likelihood ("CAMEL") software. Written from scratch in pure C++, emphasis was put in building a clean and carefully-designed project where new data and/or cosmological computations can be easily included. CAMEL incorporates the latest cosmological likelihoods and gives access from the very same input file to several estimation methods: (i) A high quality Maximum Likelihood Estimate (a.k.a "best fit") using MINUIT ; (ii) profile likelihoods, (iii) a new implementation of an Adaptive Metropolis MCMC algorithm that relieves the burden of reconstructing the proposal distribution. We present here those various statistical techniques and roll out a full use-case that can then used as a tutorial. We revisit the $Λ$CDM parameters determination with the latest Planck data and give results with both methodologies. Furthermore, by comparing the Bayesian and frequentist approaches, we discuss a "likelihood volume effect" that affects the optical reionization depth when analyzing the high multipoles part of the Planck data. The software, used in several Planck data analyzes, is available from http://camel.in2p3.fr. Using it does not require advanced C++ skills.

astro-ph.CO

Optimized Large-Scale CMB Likelihood And Quadratic Maximum Likelihood Power Spectrum Estimation

We revisit the problem of exact CMB likelihood and power spectrum estimation with the goal of minimizing computational cost through linear compression. This idea was originally proposed for CMB purposes by Tegmark et al.\ (1997), and here we develop it into a fully working computational framework for large-scale polarization analysis, adopting \WMAP\ as a worked example. We compare five different linear bases (pixel space, harmonic space, noise covariance eigenvectors, signal-to-noise covariance eigenvectors and signal-plus-noise covariance eigenvectors) in terms of compression efficiency, and find that the computationally most efficient basis is the signal-to-noise eigenvector basis, which is closely related to the Karhunen-Loeve and Principal Component transforms, in agreement with previous suggestions. For this basis, the information in 6836 unmasked \WMAP\ sky map pixels can be compressed into a smaller set of 3102 modes, with a maximum error increase of any single multipole of 3.8\% at $\ell\le32$, and a maximum shift in the mean values of a joint distribution of an amplitude--tilt model of 0.006$σ$. This compression reduces the computational cost of a single likelihood evaluation by a factor of 5, from 38 to 7.5 CPU seconds, and it also results in a more robust likelihood by implicitly regularizing nearly degenerate modes. Finally, we use the same compression framework to formulate a numerically stable and computationally efficient variation of the Quadratic Maximum Likelihood implementation that requires less than 3 GB of memory and 2 CPU minutes per iteration for $\ell \le 32$, rendering low-$\ell$ QML CMB power spectrum analysis fully tractable on a standard laptop.

astro-ph.IM

Large-scale CMB temperature and polarization cross-spectra likelihoods

We present a cross-spectra based approach for the analysis of CMB data at large angular scales to constrain the reionization optical depth $τ$, the tensor to scalar ratio $r$ and the amplitude of the primordial scalar perturbations $A_s$. With respect to the pixel-based approach developed so far, using cross-spectra has the unique advantage to eliminate spurious noise bias and to give a better handle over residual systematics, allowing to efficiently combine the cosmological information encoded in cross-frequency or cross-dataset spectra. We present two solutions to deal with the non-Gaussianity of the $\hat{C}_\ell$ estimator distributions at large angular scales: the first one relies on an analytical parametrization of the estimator distribution, while the second one is based on modification of the Hamimache\&Lewis likelihood approximation at large angular scales. The modified HL method (oHL) is powerful and complete. It allows to deal with multipole and mode correlations for a combined temperature and polarization analysis. We validate our likelihoods on numerous simulations that include the realistic noise levels of the \wmap, \planck-LFI and \planck-HFI experiments, demonstrating their validity over a broad range of cross-spectra configurations.

astro-ph.CO

Polarization measurements analysis II. Best estimators of polarization fraction and angle

With the forthcoming release of high precision polarization measurements, such as from the Planck satellite, it becomes critical to evaluate the performance of estimators for the polarization fraction and angle. These two physical quantities suffer from a well-known bias in the presence of measurement noise, as has been described in part I of this series. In this paper, part II of the series, we explore the extent to which various estimators may correct the bias. Traditional frequentist estimators of the polarization fraction are compared with two recent estimators: one inspired by a Bayesian analysis and a second following an asymptotic method. We investigate the sensitivity of these estimators to the asymmetry of the covariance matrix which may vary over large datasets. We present for the first time a comparison among polarization angle estimators, and evaluate the statistical bias on the angle that appears when the covariance matrix exhibits effective ellipticity. We also address the question of the accuracy of the polarization fraction and angle uncertainty estimators. The methods linked to the credible intervals and to the variance estimates are tested against the robust confidence interval method. From this pool of estimators, we build recipes adapted to different use-cases: build a mask, compute large maps, and deal with low S/N data. More generally, we show that the traditional estimators suffer from discontinuous distributions at low S/N, while the asymptotic and Bayesian methods do not. Attention is given to the shape of the output distribution of the estimators, and is compared with a Gaussian. In this regard, the new asymptotic method presents the best performance, while the Bayesian output distribution is shown to be strongly asymmetric with a sharp cut at low S/N.Finally, we present an optimization of the estimator derived from the Bayesian analysis using adapted priors.

astro-ph.IM

Polarization measurements analysis I. Impact of the full covariance matrix on polarization fraction and angle measurements

With the forthcoming release of high precision polarization measurements, such as from the Planck satellite, the metrology of polarization needs to improve. In particular, it is crucial to take into account full knowledge of the noise properties when estimating polarization fraction and angle, which suffer from well-known biases. While strong simplifying assumptions have usually been made in polarization analysis, we present a method for including the full covariance matrix of the Stokes parameters in estimates for the distributions of the polarization fraction and angle. We thereby quantify the impact of the noise properties on the biases in the observational quantities. We derive analytical expressions for the pdf of these quantities, taking into account the full complexity of the covariance matrix, including the Stokes I intensity components. We perform simulations to explore the impact of the noise properties on the statistical variance and bias of the polarization fraction and angle. We show that for low variations of the effective ellipticity between the Q and U components around the symmetrical case the covariance matrix may be simplified as is usually done, with negligible impact on the bias. For S/N on intensity lower than 10 the uncertainty on the total intensity is shown to drastically increase the uncertainty of the polarization fraction but not the relative bias, while a 10\% correlation between the intensity and the polarized components does not significantly affect the bias of the polarization fraction. We compare estimates of the uncertainties affecting polarization measurements, addressing limitations of estimates of the S/N, and we show how to build conservative confidence intervals for polarization fraction and angle simultaneously. This study is the first of a set of papers dedicated to the analysis of polarization measurements.

astro-ph.IM

A novel estimator of the polarization amplitude from normally distributed Stokes parameters

We propose a novel estimator of the polarization amplitude from a single measurement of its normally distributed $(Q,U)$ Stokes components. Based on the properties of the Rice distribution and dubbed 'MAS' (Modified ASymptotic), it meets several desirable criteria:(i) its values lie in the whole positive region; (ii) its distribution is continuous; (iii) it transforms smoothly with the signal-to-noise ratio (SNR) from a Rayleigh-like shape to a Gaussian one; (iv) it is unbiased and reaches its components' variance as soon as the SNR exceeds 2; (v) it is analytic and can therefore be used on large data-sets. We also revisit the construction of its associated confidence intervals and show how the Feldman-Cousins prescription efficiently solves the issue of classical intervals lying entirely in the unphysical negative domain. Such intervals can be used to identify statistically significant polarized regions and conversely build masks for polarization data. We then consider the case of a general $[Q,U]$ covariance matrix and perform a generalization of the estimator that preserves its asymptotic properties. We show that its bias does not depend on the true polarization angle, and provide an analytic estimate of its variance. The estimator value, together with its variance, provide a powerful point-estimate of the true polarization amplitude that follows an unbiased Gaussian distribution for a SNR as low as 2. These results can be applied to the much more general case of transforming any normally distributed random variable from Cartesian to polar coordinates.

astro-ph.CO

A hybrid approach to CMB lensing reconstruction on all-sky intensity maps

Based on realistic simulations, we propose an hybrid method to reconstruct the lensing potential power spectrum, directly on PLANCK-like CMB frequency maps. It implies using a large galactic mask and dealing with a strong inhomogeneous noise. For l < 100, we show that a full-sky inpainting method, already described in a previous work, still allows a minimal variance reconstruction, with a bias that must be accounted for by a Monte-Carlo method, but that does not couple to the deflection field. For l>100 we develop a method based on tiling the cut-sky with local 10x10 degrees overlapping tangent planes (referred to in the following as "patches"). It requires to solve various issues concerning their size/position, non-periodic boundaries and irregularly sampled data after the sphere-to-plane projection. We show how the leading noise term of the quadratic lensing estimator applied onto an apodized patch can still be taken directly from the data. To not loose spatial accuracy, we developed a tool that allows the fast determination of the complex Fourier series coefficients from a bi-dimensional irregularly sampled dataset, without performing an interpolation. We show that the multi-patch approach allows the lensing power spectrum reconstruction with a very small bias, thanks to avoiding the galactic mask and lowering the noise inhomogeneities, while still having almost a minimal variance. The data quality can be assessed at each stage and simple bi-dimensional spectra build, which allows the control of local systematic errors.

astro-ph.CO

Iterative destriping and photometric calibration for Planck-HFI, polarized, multi-detector map-making

We present an iterative scheme designed to recover calibrated I, Q, and U maps from Planck-HFI data using the orbital dipole due to the satellite motion with respect to the Solar System frame. It combines a map reconstruction, based on a destriping technique, juxtaposed with an absolute calibration algorithm. We evaluate systematic and statistical uncertainties incurred during both these steps with the help of realistic, Planck-like simulations containing CMB, foreground components and instrumental noise, and assess the accuracy of the sky map reconstruction by considering the maps of the residuals and their spectra. In particular, we discuss destriping residuals for polarization sensitive detectors similar to those of Planck-HFI under different noise hypotheses and show that these residuals are negligible (for intensity maps) or smaller than the white noise level (for Q and U Stokes maps), for l > 50. We also demonstrate that the combined level of residuals of this scheme remains comparable to those of the destriping-only case except at very low l where residuals from the calibration appear. For all the considered noise hypotheses, the relative calibration precision is on the order of a few 10e-4, with a systematic bias of the same order of magnitude.

astro-ph.CO

Towards a fast, model-independent Cosmic Microwave Background bispectrum estimator

The measurements of the statistical properties of the Cosmic Microwave Background (CMB) fluctuations enable us to probe the physics of the very early Universe especially at the epoch of inflation. A particular interest lays on the detection of the non-Gaussianity of the CMB as it can constrain the current proposed models of inflation and structure formation, or possibly point out new models. The current approach to measure the degree of non-Gaussianity of the CMB is to estimate a single parameter which is highly model-dependent. The bispectrum is a natural and widely studied tool for measuring the non-Gaussianity in a model-independent way. This paper sets the grounds for a full CMB bispectrum estimator based on the decomposition of the sphere onto projected patches. The mean bispectrum estimated this way can be calculated quickly and is model-independent. This approach is very flexible, allowing exclusion of some patches in the processing or consideration of just a specific region of the sphere.

astro-ph.CO