SearcharxivSearch

arXiv subjects

Matthew Turk

Publications and source records attributed to Matthew Turk.

27 records · Page 2Linked to original sources

Data-Intensive Supercomputing in the Cloud: Global Analytics for Satellite Imagery

We present our experiences using cloud computing to support data-intensive analytics on satellite imagery for commercial applications. Drawing from our background in high-performance computing, we draw parallels between the early days of clustered computing systems and the current state of cloud computing and its potential to disrupt the HPC market. Using our own virtual file system layer on top of cloud remote object storage, we demonstrate aggregate read bandwidth of 230 gigabytes per second using 512 Google Compute Engine (GCE) nodes accessing a USA multi-region standard storage bucket. This figure is comparable to the best HPC storage systems in existence. We also present several of our application results, including the identification of field boundaries in Ukraine, and the generation of a global cloud-free base layer from Landsat imagery.

cs.DC

Large Scale SfM with the Distributed Camera Model

We introduce the distributed camera model, a novel model for Structure-from-Motion (SfM). This model describes image observations in terms of light rays with ray origins and directions rather than pixels. As such, the proposed model is capable of describing a single camera or multiple cameras simultaneously as the collection of all light rays observed. We show how the distributed camera model is a generalization of the standard camera model and describe a general formulation and solution to the absolute camera pose problem that works for standard or distributed cameras. The proposed method computes a solution that is up to 8 times more efficient and robust to rotation singularities in comparison with gDLS. Finally, this method is used in an novel large-scale incremental SfM pipeline where distributed cameras are accurately and robustly merged together. This pipeline is a direct generalization of traditional incremental SfM; however, instead of incrementally adding one camera at a time to grow the reconstruction the reconstruction is grown by adding a distributed camera. Our pipeline produces highly accurate reconstructions efficiently by avoiding the need for many bundle adjustment iterations and is capable of computing a 3D model of Rome from over 15,000 images in just 22 minutes.

cs.CV

Capturing the "Whole Tale" of Computational Research: Reproducibility in Computing Environments

We present an overview of the recently funded "Merging Science and Cyberinfrastructure Pathways: The Whole Tale" project (NSF award #1541450). Our approach has two nested goals: 1) deliver an environment that enables researchers to create a complete narrative of the research process including exposure of the data-to-publication lifecycle, and 2) systematically and persistently link research publications to their associated digital scholarly objects such as the data, code, and workflows. To enable this, Whole Tale will create an environment where researchers can collaborate on data, workspaces, and workflows and then publish them for future adoption or modification. Published data and applications will be consumed either directly by users using the Whole Tale environment or can be integrated into existing or future domain Science Gateways.

cs.DL

One-Class Slab Support Vector Machine

This work introduces the one-class slab SVM (OCSSVM), a one-class classifier that aims at improving the performance of the one-class SVM. The proposed strategy reduces the false positive rate and increases the accuracy of detecting instances from novel classes. To this end, it uses two parallel hyperplanes to learn the normal region of the decision scores of the target class. OCSSVM extends one-class SVM since it can scale and learn non-linear decision functions via kernel methods. The experiments on two publicly available datasets show that OCSSVM can consistently outperform the one-class SVM and perform comparable to or better than other state-of-the-art one-class classifiers.

cs.CV

The Formation of Submillimetre-Bright Galaxies from Gas Infall over a Billion Years

Submillimetre-luminous galaxies at high-redshift are the most luminous, heavily star-forming galaxies in the Universe, and are characterised by prodigious emission in the far-infrared at 850 microns (S850 > 5 mJy). They reside in halos ~ 10^13Msun, have low gas fractions compared to main sequence disks at a comparable redshift, trace complex environments, and are not easily observable at optical wavelengths. Their physical origin remains unclear. Simulations have been able to form galaxies with the requisite luminosities, but have otherwise been unable to simultaneously match the stellar masses, star formation rates, gas fractions and environments. Here we report a cosmological hydrodynamic galaxy formation simulation that is able to form a submillimetre galaxy which simultaneously satisfies the broad range of observed physical constraints. We find that groups of galaxies residing in massive dark matter halos have rising star formation histories that peak at collective rates ~ 500-1000 Msun/yr at z=2-3, by which time the interstellar medium is sufficiently enriched with metals that the region may be observed as a submillimetre-selected system. The intense star formation rates are fueled in part by a reservoir gas supply enabled by stellar feedback at earlier times, not through major mergers. With a duty cycle of nearly a gigayear, our simulations show that the submillimetre-luminous phase of high-z galaxies is a drawn out one that is associated with significant mass buildup in early Universe proto-clusters, and that many submillimetre-luminous galaxies are actually composed of numerous unresolved components (for which there is some observational evidence).

astro-ph.GA

Second Workshop on Sustainable Software for Science: Practice and Experiences (WSSSPE2): Submission, Peer-Review and Sorting Process, and Results

This technical report discusses the submission and peer-review process used by the Second Workshop on Sustainable Software for Science: Practice and Experiences (WSSSPE2) and the results of that process. It is intended to record both the alternative submission and program organization model used by WSSSPE2 as well as the papers associated with the workshop that resulted from that process.

cs.SE

Summary of the First Workshop on Sustainable Software for Science: Practice and Experiences (WSSSPE1)

Challenges related to development, deployment, and maintenance of reusable software for science are becoming a growing concern. Many scientists' research increasingly depends on the quality and availability of software upon which their works are built. To highlight some of these issues and share experiences, the First Workshop on Sustainable Software for Science: Practice and Experiences (WSSSPE1) was held in November 2013 in conjunction with the SC13 Conference. The workshop featured keynote presentations and a large number (54) of solicited extended abstracts that were grouped into three themes and presented via panels. A set of collaborative notes of the presentations and discussion was taken during the workshop. Unique perspectives were captured about issues such as comprehensive documentation, development and deployment practices, software licenses and career paths for developers. Attribution systems that account for evidence of software contribution and impact were also discussed. These include mechanisms such as Digital Object Identifiers, publication of "software papers", and the use of online systems, for example source code repositories like GitHub. This paper summarizes the issues and shared experiences that were discussed, including cross-cutting issues and use cases. It joins a nascent literature seeking to understand what drives software work in science, and how it is impacted by the reward systems of science. These incentives can determine the extent to which developers are motivated to build software for the long-term, for the use of others, and whether to work collaboratively or separately. It also explores community building, leadership, and dynamics in relation to successful scientific software.

cs.SE

Constraints on Hydrodynamical Subgrid Models from Quasar Absorption Line Studies of the Simulated Circumgalactic Medium

Cosmological hydrodynamical simulations of galaxy evolution are increasingly able to produce realistic galaxies, but the largest hurdle remaining is in constructing subgrid models that accurately describe the behavior of stellar feedback. As an alternate way to test and calibrate such models, we propose to focus on the circumgalactic medium. To do so, we generate a suite of adaptive-mesh refinement (AMR) simulations for a Milky-Way-massed galaxy run to z=0, systematically varying the feedback implementation. We then post-process the simulation data to compute the absorbing column density for a wide range of common atomic absorbers throughout the galactic halo, including H I, Mg II, Si II, Si III, Si IV, C IV, N V, O VI, and O VII. The radial profiles of these atomic column densities are compared against several quasar absorption line studies, to determine if one feedback prescription is favored. We find that although our models match some of the observations (specifically those ions with lower ionization strengths), it is particularly difficult to match O VI observations. There is some indication that the models with increased feedback intensity are better matches. We demonstrate that sufficient metals exist in these halos to reproduce the observed column density distribution in principle, but the simulated circumgalactic medium lacks significant multiphase substructure and is generally too hot. Furthermore, we demonstrate the failings of inflow-only models (without energetic feedback) at populating the CGM with adequate metals to match observations even in the presence of multiphase structure. Additionally, we briefly investigate the evolution of the CGM from z=3 to present. Overall, we find that quasar absorption line observations of the gas around galaxies provide a new and important constraint on feedback models.

astro-ph.GA

Face Verification in Polar Frequency Domain: a Biologically Motivated Approach

We present a novel local-based face verification system whose components are analogous to those of biological systems. In the proposed system, after global registration and normalization, three eye regions are converted from the spatial to polar frequency domain by a Fourier-Bessel Transform. The resulting representations are embedded in a dissimilarity space, where each image is represented by its distance to all the other images. In this dissimilarity space a Pseudo-Fisher discriminator is built. ROC and equal error rate verification test results on the FERET database showed that the system performed at least as state-of-the-art methods and better than a system based on polar Fourier features. The local-based system is especially robust to facial expression and age variations, but sensitive to registration errors.

cs.CV