SearcharxivSearch

arXiv subjects

James Jackson

Publications and source records attributed to James Jackson.

13 recordsLinked to original sources

Path-based Design Model for Constructing and Exploring Alternative Visualisations

We present a path-based design model and system for designing and creating visualisations. Our model represents a systematic approach to constructing visual representations of data or concepts following a predefined sequence of steps. The initial step involves outlining the overall appearance of the visualisation by creating a skeleton structure, referred to as a flowpath. Subsequently, we specify objects, visual marks, properties, and appearance, storing them in a gene. Lastly, we map data onto the flowpath, ensuring suitable morphisms. Alternative designs are created by exchanging values in the gene. For example, designs that share similar traits, are created by making small incremental changes to the gene. Our design methodology fosters the generation of diverse creative concepts, space-filling visualisations, and traditional formats like bar charts, circular plots and pie charts. Through our implementation we showcase the model in action. As an example application, we integrate the output visualisations onto a smartwatch and visualisation dashboards. In this article we (1) introduce, define and explain the path model and discuss possibilities for its use, (2) present our implementation, results, and evaluation, and (3) demonstrate and evaluate an application of its use on a mobile watch.

cs.HC

The appeal of the gamma family distribution to protect the confidentiality of contingency tables

Administrative databases, such as the English School Census (ESC), are rich sources of information that are potentially useful for researchers. For such data sources to be made available, however, strict guarantees of privacy would be required. To achieve this, synthetic data methods can be used. Such methods, when protecting the confidentiality of tabular data (contingency tables), often utilise the Poisson or Poisson-mixture distributions, such as the negative binomial (NBI). These distributions, however, are either equidispersed (in the case of the Poisson) or overdispersed (e.g. in the case of the NBI), which results in excessive noise being applied to large low-risk counts. This paper proposes the use of the (discretized) gamma family (GAF) distribution, which allows noise to be applied in a more bespoke fashion. Specifically, it allows less noise to be applied as cell counts become larger, providing an optimal balance in relation to the risk-utility trade-off. We illustrate the suitability of the GAF distribution on an administrative-type data set that is reminiscent of the ESC.

stat.ME

Bernoulli amputation

A novel, stochastic approach to amputation, the process of introducing missing values to a complete dataset, is presented. It allows one to construct a wide variety of missingness patterns by only having to specify distributions of missingness indicators as opposed to specifying each missingness pattern manually. Missingness indicators are modeled in a principled way via copulas and Bernoulli margins, thus allowing one to incorporate dependence in missingness patterns. Besides more classical missingness mechanisms such as missing completely at random, missing at random, and missing not at random, the approach is able to model structured missingness such as block missingness and, via mixtures, monotone missingness, which are patterns of missing data frequently found in real-life datasets. Properties such as joint missingness probabilities or missingness correlation are derived mathematically. The flexibility of the approach in capturing different missingness patterns while only requiring to specify distributional assumptions on missingness indicators is demonstrated with mathematical examples and empirical illustrations in terms of a well-known example dataset of sufficiently small sample size that allows to identify each missing data point visually. Finally, an example application to multivariate financial time series is provided.

stat.AP

Obtaining $(\epsilon,\delta)$-differential privacy guarantees when using a Poisson mechanism to synthesize contingency tables

We show that differential privacy type guarantees can be obtained when using a Poisson synthesis mechanism to protect counts in contingency tables. Specifically, we show how to obtain $(\epsilon, \delta)$-probabilistic differential privacy guarantees via the Poisson distribution's cumulative distribution function. We demonstrate this empirically with the synthesis of an administrative-type confidential database.

cs.CR

A Complete Characterisation of Structured Missingness

Our capacity to process large complex data sources is ever-increasing, providing us with new, important applied research questions to address, such as how to handle missing values in large-scale databases. Mitra et al. (2023) noted the phenomenon of Structured Missingness (SM), which is where missingness has an underlying structure. Existing taxonomies for defining missingness mechanisms typically assume that variables' missingness indicator vectors $M_1$, $M_2$, ..., $M_p$ are independent after conditioning on the relevant portion of the data matrix $\mathbf{X}$. As this is often unsuitable for characterising SM in multivariate settings, we introduce a taxonomy for SM, where each ${M}_j$ can depend on $\mathbf{M}_{-j}$ (i.e., all missingness indicator vectors except ${M}_j$), in addition to $\mathbf{X}$. We embed this new framework within the well-established decomposition of mechanisms into MCAR, MAR, and MNAR (Rubin, 1976), allowing us to recast mechanisms into a broader setting, where we can consider the combined effect of $\mathbf{X}$ and $\mathbf{M}_{-j}$ on ${M}_j$. We also demonstrate, via simulations, the impact of SM on inference and prediction, and consider contextual instances of SM arising in a de-identified nationwide (US-based) clinico-genomic database (CGDB). We hope to stimulate interest in SM, and encourage timely research into this phenomenon.

stat.ME

Science with an ngVLA: The ngVLA Reference Design

The next-generation Very Large Array (ngVLA) is an astronomical observatory planned to operate at centimeter wavelengths (25 to 0.26 centimeters, corresponding to a frequency range extending from 1.2 to 116 GHz). The observatory will be a synthesis radio telescope constituted of approximately 244 reflector antennas each of 18 meters diameter, and 19 reflector antennas each of 6 meters diameter, operating in a phased or interferometric mode. We provide a technical overview of the Reference Design of the ngVLA. This Reference Design forms a baseline for a technical readiness assessment and the construction and operations cost estimate of the ngVLA. The concepts for major system elements such as the antenna, receiving electronics, and central signal processing are presented.

astro-ph.IM

The Radio Ammonia Mid-Plane Survey (RAMPS) Pilot Survey

The Radio Ammonia Mid-Plane Survey (RAMPS) is a molecular line survey that aims to map a portion of the Galactic midplane in the first quadrant of the Galaxy (l = 10 deg - 40 deg, |b| < 0.4 deg) using the Green Bank Telescope. We present results from the pilot survey, which has mapped approximately 6.5 square degrees in fields centered at l = 10 deg, 23 deg, 24 deg, 28 deg, 29 deg, 30 deg, 31 deg, 38 deg, 45 deg, and 47 deg. RAMPS observes the NH3 inversion transitions NH3(1, 1) - (5, 5), the H2O 6(1,6) - 5(2,3) maser line at 22.235 GHz, and several other molecular lines. We present a representative portion of the data from the pilot survey, including NH3(1,1) and NH3(2,2) integrated intensity maps, H2O maser positions, maps of NH3 velocity, NH3 line width, total NH3 column density, and NH3 rotational temperature. These data and the data cubes from which they were produced are publicly available on the RAMPS website (http://sites.bu.edu/ramps/).

astro-ph.GA

The Next Generation Very Large Array: A Technical Overview

The next-generation Very Large Array (ngVLA) is an astronomical observatory planned to operate at centimeter wavelengths (25 to 0.26 centimeters, corresponding to a frequency range extending from 1.2 GHz to 116 GHz). The observatory will be a synthesis radio telescope constituted of approximately 214 reflector antennas each of 18 meters diameter, operating in a phased or interferometric mode. We provide an overview of the current system design of the ngVLA. The concepts for major system elements such as the antenna, receiving electronics, and central signal processing are presented. We also describe the major development activities that are presently underway to advance the design.

astro-ph.IM

The Survey of Water and Ammonia in the Galactic Center (SWAG): Molecular Cloud Evolution in the Central Molecular Zone

The Survey of Water and Ammonia in the Galactic Center (SWAG) covers the Central Molecular Zone (CMZ) of the Milky Way at frequencies between 21.2 and 25.4 GHz obtained at the Australia Telescope Compact Array at $\sim 0.9$ pc spatial and $\sim 2.0$ km s$^{-1}$ spectral resolution. In this paper, we present data on the inner $\sim 250$ pc ($1.4^\circ$) between Sgr C and Sgr B2. We focus on the hyperfine structure of the metastable ammonia inversion lines (J,K) = (1,1) - (6,6) to derive column density, kinematics, opacity and kinetic gas temperature. In the CMZ molecular clouds, we find typical line widths of $8-16$ km s$^{-1}$ and extended regions of optically thick ($\tau > 1$) emission. Two components in kinetic temperature are detected at $25-50$ K and $60-100$ K, both being significantly hotter than dust temperatures throughout the CMZ. We discuss the physical state of the CMZ gas as traced by ammonia in the context of the orbital model by Kruijssen et al. (2015) that interprets the observed distribution as a stream of molecular clouds following an open eccentric orbit. This allows us to statistically investigate the time dependencies of gas temperature, column density and line width. We find heating rates between $\sim 50$ and $\sim 100$ K Myr$^{-1}$ along the stream orbit. No strong signs of time dependence are found for column density or line width. These quantities are likely dominated by cloud-to-cloud variations. Our results qualitatively match the predictions of the current model of tidal triggering of cloud collapse, orbital kinematics and the observation of an evolutionary sequence of increasing star formation activity with orbital phase.

astro-ph.GA

Characterizing the properties of cluster precursors in the MALT90 survey

In the Milky Way there are thousands of stellar clusters each harboring from a hundred to a million stars. Although clusters are common, the initial conditions of cluster formation are still not well understood. To determine the processes involved in the formation and evolution of clusters it is key to determine the global properties of cluster-forming clumps in their earliest stages of evolution. Here, we present the physical properties of 1,244 clumps identified from the MALT90 survey. Using the dust temperature of the clumps as a proxy for evolution we determined how the clump properties change at different evolutionary stages. We find that less-evolved clumps exhibiting dust temperatures lower than 20 K have higher densities and are more gravitationally bound than more-evolved clumps with higher dust temperatures. We also identified a sample of clumps in a very early stage of evolution, thus potential candidates for high-mass star-forming clumps. Only one clump in our sample has physical properties consistent with a young massive cluster progenitor, reinforcing the fact that massive proto-clusters are very rare in the Galaxy.

astro-ph.GA

The Bones of the Milky Way

The very long and thin infrared dark cloud "Nessie" is even longer than had been previously claimed, and an analysis of its Galactic location suggests that it lies directly in the Milky Way's mid-plane, tracing out a highly elongated bone-like feature within the prominent Scutum-Centaurus spiral arm. Re-analysis of mid-infrared imagery from the Spitzer Space Telescope shows that this IRDC is at least 2, and possibly as many as 5 times longer than had originally been claimed by Nessie's discoverers (Jackson et al. 2010); its aspect ratio is therefore at least 300:1, and possibly as large as 800:1. A careful accounting for both the Sun's offset from the Galactic plane ($\sim 25$ pc) and the Galactic center's offset from the $(l^{II},b^{II})=(0,0)$ position shows that the latitude of the true Galactic mid-plane at the 3.1 kpc distance to the Scutum-Centaurus Arm is not $b=0$, but instead closer to $b=-0.4$, which is the latitude of Nessie to within a few pc. An analysis of the radial velocities of low-density (CO) and high-density (${\rm NH}_3$) gas associated with the Nessie dust feature suggests that Nessie runs along the Scutum-Centaurus Arm in position-position-velocity space, which means it likely forms a dense `spine' of the arm in real space as well. The Scutum-Centaurus arm is the closest major spiral arm to the Sun toward the inner Galaxy, and, at the longitude of Nessie, it is almost perpendicular to our line of sight, making Nessie the easiest feature to see as a shadow elongated along the Galactic Plane from our location. Future high-resolution dust mapping and molecular line observations of the harder-to-find Galactic "bones" should allow us to exploit the Sun's position above the plane to gain a (very foreshortened) view "from above" of the Milky Way's structure.

astro-ph.GA

G0.253+0.016: a molecular cloud progenitor of an Arches-like cluster

Young massive clusters (YMCs) with stellar masses of 10^4 - 10^5 Msun and core stellar densities of 10^4 - 10^5 stars per cubic pc are thought to be the `missing link' between open clusters and extreme extragalactic super star clusters and globular clusters. As such, studying the initial conditions of YMCs offers an opportunity to test cluster formation models across the full cluster mass range. G0.253+0.016 is an excellent candidate YMC progenitor. We make use of existing multi-wavelength data including recently available far-IR continuum (Herschel/Hi-GAL) and mm spectral line (HOPS and MALT90) data and present new, deep, multiple-filter, near-IR (VLT/NACO) observations to study G0.253+0.016. These data show G0.253+0.016 is a high mass (1.3x10^5 Msun), low temperature (T_dust~20K), high volume and column density (n ~ 8x10^4 cm^-3; N_{H_2} ~ 4x10^23 cm^-2) molecular clump which is close to virial equilibrium (M_dust ~ M_virial) so is likely to be gravitationally-bound. It is almost devoid of star formation and, thus, has exactly the properties expected for the initial conditions of a clump that may form an Arches-like massive cluster. We compare the properties of G0.253+0.016 to typical Galactic cluster-forming molecular clumps and find it is extreme, and possibly unique in the Galaxy. This uniqueness makes detailed studies of G0.253+0.016 extremely important for testing massive cluster formation models.

astro-ph.GA

The Turbulence Spectrum of Molecular Clouds in the Galactic Ring Survey: A Density-Dependent PCA Calibration

Turbulence plays a major role in the formation and evolution of molecular clouds. The problem is that turbulent velocities are convolved with the density of an observed region. To correct for this convolution, we investigate the relation between the turbulence spectrum of model clouds, and the statistics of their synthetic observations obtained from Principal Component Analysis (PCA). We apply PCA to spectral maps generated from simulated density and velocity fields, obtained from hydrodynamic simulations of supersonic turbulence, and from fractional Brownian motion fields with varying velocity, density spectra, and density dispersion. We examine the dependence of the slope of the PCA structure function, alpha_PCA, on intermittency, on the turbulence velocity (beta_v) and density (beta_n) spectral indexes, and on density dispersion. We find that PCA is insensitive to beta_n and to the log-density dispersion sigma_s, provided sigma_s < 2. For sigma_s>2, alpha_PCA increases with sigma_s due to the intermittent sampling of the velocity field by the density field. The PCA calibration also depends on intermittency. We derive a PCA calibration based on fBms with sigma_s<2 and apply it to 367 CO spectral maps of molecular clouds in the Galactic Ring Survey. The average slope of the PCA structure function, =0.62\pm0.2, is consistent with the hydrodynamic simulations and leads to a turbulence velocity exponent =2.06\pm0.6 for a non-intermittent, low density dispersion flow. Accounting for intermittency and density dispersion, the coincidence between the PCA slope of the GRS clouds and the hydrodynamic simulations suggests beta_v~1.9, consistent with both Burgers and compressible intermittent turbulence.

astro-ph.GA