SearcharxivSearch

arXiv subjects

Elizabeth Bradley

Publications and source records attributed to Elizabeth Bradley.

At least 19 recordsLinked to original sources

Introduction to Focus Issue: Topics in Nonlinear Science

Nonlinear science has evolved significantly over the 35 years since the launch of the journal Chaos. This Focus Issue, dedicated to the 80th Birthday of its founding editor-in-chief, David K. Campbell, brings together a selection of contributions on influential topics, many of which were advanced by Campbell's own research program and leadership role. The topics include new phenomena and method development in the realms of network dynamics, machine learning, quantum and material systems, chaos and fractals, localized states, and living systems, with a good balance of literature review, original contributions, and perspectives for future research.

nlin.AO

A Computational Topology-based Spatiotemporal Analysis Technique for Honeybee Aggregation

A primary challenge in understanding collective behavior is characterizing the spatiotemporal dynamics of the group. We employ topological data analysis to explore the structure of honeybee aggregations that form during trophallaxis, which is the direct exchange of food among nestmates. From the positions of individual bees, we build topological summaries called CROCKER matrices to track the morphology of the group as a function of scale and time. Each column of a CROCKER matrix records the number of topological features, such as the number of components or holes, that exist in the data for a range of analysis scales at a given point in time. To detect important changes in the morphology of the group from this information, we first apply dimensionality reduction techniques to these matrices and then use classic clustering and change-point detection algorithms on the resulting scalar data. A test of this methodology on synthetic data from an agent-based model of honeybees and their trophallaxis behavior shows two distinct phases: a dispersed phase that occurs before food is introduced, followed by a food-exchange phase during which aggregations form. We then move to laboratory data, successfully detecting the same two phases across multiple experiments. Interestingly, our method reveals an additional phase change towards the end of the experiments, suggesting the possibility of another dispersed phase that follows the food-exchange phase.

q-bio.QM

Building Resilience to Climate Driven Extreme Events with Computing Innovations: A Convergence Accelerator Report

In 2022, the National Science Foundation (NSF) funded the Computing Research Association (CRA) to conduct a workshop to frame and scope a potential Convergence Accelerator research track on the topic of "Building Resilience to Climate-Driven Extreme Events with Computing Innovations". The CRA's research visioning committee, the Computing Community Consortium (CCC), took on this task, organizing a two-part community workshop series, beginning with a small, in-person brainstorming meeting in Denver, CO on 27-28 October 2022, followed by a virtual event on 10 November 2022. The overall objective was to develop ideas to facilitate convergence research on this critical topic and encourage collaboration among researchers across disciplines. Based on the CCC community white paper entitled Computing Research for the Climate Crisis, we initially focused on five impact areas (i.e. application domains that are both important to society and critically affected by climate change): Energy, Agriculture, Environmental Justice, Transportation, and Physical Infrastructure.

cs.CY

Using scaling-region distributions to select embedding parameters

Reconstructing state-space dynamics from scalar data using time-delay embedding requires choosing values for the delay $τ$ and the dimension $m$. Both parameters are critical to the success of the procedure and neither is easy to formally validate. While embedding theorems do offer formal guidance for these choices, in practice one has to resort to heuristics, such as the average mutual information (AMI) method of Fraser & Swinney for $τ$ or the false near neighbor (FNN) method of Kennel et al. for $m$. Best practice suggests an iterative approach: one of these heuristics is used to make a good first guess for the corresponding free parameter and then an "asymptotic invariant" approach is then used to firm up its value by, e.g., computing the correlation dimension or Lyapunov exponent for a range of values and looking for convergence. This process can be subjective, as these computations often involve finding, and fitting a line to, a scaling region in a plot: a process that is generally done by eye and is not immune to confirmation bias. Moreover, most of these heuristics do not provide confidence intervals, making it difficult to say what "convergence" is. Here, we propose an approach that automates the first step, removing the subjectivity, and formalizes the second, offering a statistical test for convergence. Our approach rests upon a recently developed method for automated scaling-region selection that includes confidence intervals on the results. We demonstrate this methodology by selecting values for the embedding dimension for several real and simulated dynamical systems. We compare these results to those produced by FNN and validate them against known results -- e.g., of the correlation dimension -- where these are available. We note that this method extends to any free parameter in the theory or practice of delay reconstruction.

physics.data-an

Machine Learning Approaches to Solar-Flare Forecasting: Is Complex Better?

Recently, there has been growing interest in the use of machine-learning methods for predicting solar flares. Initial efforts along these lines employed comparatively simple models, correlating features extracted from observations of sunspot active regions with known instances of flaring. Typically, these models have used physics-inspired features that have been carefully chosen by experts in order to capture the salient features of such magnetic field structures. Over time, the sophistication and complexity of the models involved has grown. However, there has been little evolution in the choice of feature sets, nor any systematic study of whether the additional model complexity is truly useful. Our goal is to address these issues. To that end, we compare the relative prediction performance of machine-learning-based, flare-forecasting models with varying degrees of complexity. We also revisit the feature set design, using topological data analysis to extract shape-based features from magnetic field images of the active regions. Using hyperparameter training for fair comparison of different machine-learning models across different feature sets, we show that simpler models with fewer free parameters \textit{generally perform better than more-complicated models}, ie., powerful machinery does not necessarily guarantee better prediction performance. Secondly, we find that \textit{abstract, shape-based features contain just as much useful information}, for the purposes of flare prediction, as the set of hand-crafted features developed by the solar-physics community over the years. Finally, we study the effects of dimensionality reduction, using principal component analysis, to show that streamlined feature sets, overall, perform just as well as the corresponding full-dimensional versions.

astro-ph.SR

Towards automated extraction and characterization of scaling regions in dynamical systems

Scaling regions -- intervals on a graph where the dependent variable depends linearly on the independent variable -- abound in dynamical systems, notably in calculations of invariants like the correlation dimension or a Lyapunov exponent. In these applications, scaling regions are generally selected by hand, a process that is subjective and often challenging due to noise, algorithmic effects, and confirmation bias. In this paper, we propose an automated technique for extracting and characterizing such regions. Starting with a two-dimensional plot -- e.g., the values of the correlation integral, calculated using the Grassberger-Procaccia algorithm over a range of scales -- we create an ensemble of intervals by considering all possible combinations of endpoints, generating a distribution of slopes from least-squares fits weighted by the length of the fitting line and the inverse square of the fit error. The mode of this distribution gives an estimate of the slope of the scaling region (if it exists). The endpoints of the intervals that correspond to the mode provide an estimate for the extent of that region. When there is no scaling region, the distributions will be wide and the resulting error estimates for the slope will be large. We demonstrate this method for computations of dimension and Lyapunov exponent for several dynamical systems, and show that it can be useful in selecting values for the parameters in time-delay reconstructions.

physics.data-an

Computing Research for the Climate Crisis

Climate change is an existential threat to the United States and the world. Inevitably, computing will play a key role in mitigation, adaptation, and resilience in response to this threat. The needs span all areas of computing, from devices and architectures (e.g., low-power sensor systems for wildfire monitoring) to algorithms (e.g., predicting impacts and evaluating mitigation), and robotics (e.g., autonomous UAVs for monitoring and actuation) -- as well as every level of the software stack, from data management systems and energy-aware operating systems to hardware/software co-design. The goal of this white paper is to highlight the role of computing research in addressing climate change-induced challenges. To that end, we outline six key impact areas in which these challenges will arise -- energy, environmental justice, transportation, infrastructure, agriculture, and environmental monitoring and forecasting -- then identify specific ways in which computing research can help address the associated problems. These impact areas will create a driving force behind, and enable, cross-cutting, system-level innovation. We further break down this information into four broad areas of computing research: devices & architectures, software, algorithms/AI/robotics, and sociotechnical computing. Additional contributions by: Ilkay Altintas (San Diego Supercomputer Center), Kyri Baker (University of Colorado Boulder), Sujata Banerjee (VMware), Andrew A. Chien (University of Chicago), Thomas Dietterich (Oregon State University), Ian Foster (Argonne National Labs), Carla P. Gomes (Cornell University), Chandra Krintz (University of California, Santa Barbara), Jessica Seddon (World Resources Institute), and Regan Zane (Utah State University).

cs.CY

Pandemic Informatics: Preparation, Robustness, and Resilience; Vaccine Distribution, Logistics, and Prioritization; and Variants of Concern

Infectious diseases cause more than 13 million deaths a year, worldwide. Globalization, urbanization, climate change, and ecological pressures have significantly increased the risk of a global pandemic. The ongoing COVID-19 pandemic-the first since the H1N1 outbreak more than a decade ago and the worst since the 1918 influenza pandemic-illustrates these matters vividly. More than 47M confirmed infections and 1M deaths have been reported worldwide as of November 4, 2020 and the global markets have lost trillions of dollars. The pandemic will continue to have significant disruptive impacts upon the United States and the world for years; its secondary and tertiary impacts might be felt for more than a decade. An effective strategy to reduce the national and global burden of pandemics must: 1) detect timing and location of occurrence, taking into account the many interdependent driving factors; 2) anticipate public reaction to an outbreak, including panic behaviors that obstruct responders and spread contagion; 3) and develop actionable policies that enable targeted and effective responses.

cs.CY

Shape-based Feature Engineering for Solar Flare Prediction

Solar flares are caused by magnetic eruptions in active regions (ARs) on the surface of the sun. These events can have significant impacts on human activity, many of which can be mitigated with enough advance warning from good forecasts. To date, machine learning-based flare-prediction methods have employed physics-based attributes of the AR images as features; more recently, there has been some work that uses features deduced automatically by deep learning methods (such as convolutional neural networks). We describe a suite of novel shape-based features extracted from magnetogram images of the Sun using the tools of computational topology and computational geometry. We evaluate these features in the context of a multi-layer perceptron (MLP) neural network and compare their performance against the traditional physics-based attributes. We show that these abstract shape-based features outperform the features chosen by the human experts, and that a combination of the two feature sets improves the forecasting capability even further.

astro-ph.SR

An Agenda for Disinformation Research

In the 21st Century information environment, adversarial actors use disinformation to manipulate public opinion. The distribution of false, misleading, or inaccurate information with the intent to deceive is an existential threat to the United States--distortion of information erodes trust in the socio-political institutions that are the fundamental fabric of democracy: legitimate news sources, scientists, experts, and even fellow citizens. As a result, it becomes difficult for society to come together within a shared reality; the common ground needed to function effectively as an economy and a nation. Computing and communication technologies have facilitated the exchange of information at unprecedented speeds and scales. This has had countless benefits to society and the economy, but it has also played a fundamental role in the rising volume, variety, and velocity of disinformation. Technological advances have created new opportunities for manipulation, influence, and deceit. They have effectively lowered the barriers to reaching large audiences, diminishing the role of traditional mass media along with the editorial oversight they provided. The digitization of information exchange, however, also makes the practices of disinformation detectable, the networks of influence discernable, and suspicious content characterizable. New tools and approaches must be developed to leverage these affordances to understand and address this growing challenge.

cs.CY

Detection of Local Mixing in Time-Series Data Using Permutation Entropy

While it is tempting in experimental practice to seek as high a data rate as possible, oversampling can become an issue if one takes measurements too densely. These effects can take many forms, some of which are easy to detect: e.g., when the data sequence contains multiple copies of the same measured value. In other situations, as when there is mixing$-$in the measurement apparatus and/or the system itself$-$oversampling effects can be harder to detect. We propose a novel, model-free technique to detect local mixing in time series using an information-theoretic technique called permutation entropy. By varying the temporal resolution of the calculation and analyzing the patterns in the results, we can determine whether the data are mixed locally, and on what scale. This can be used by practitioners to choose appropriate lower bounds on scales at which to measure or report data. After validating this technique on several synthetic examples, we demonstrate its effectiveness on data from a chemistry experiment, methane records from Mauna Loa, and an Antarctic ice core.

cs.IT

Using Curvature to Select the Time Lag for Delay Reconstruction

We propose a curvature-based approach for choosing good values for the time-delay parameter $τ$ in delay reconstructions. The idea is based on the effects of the delay on the geometry of the reconstructions. If the delay is chosen too small, the reconstructed dynamics are flattened along the main diagonal of the embedding space; too-large delays, on the other hand, can overfold the dynamics. Calculating the curvature of a two-dimensional delay reconstruction is an effective way to identify these extremes and to find a middle ground between them: both the sharp reversals at the ends of an insufficiently unfolded reconstruction and the folds in an overfolded one create spikes in the curvature. We operationalize this observation by computing the mean over the Menger curvature of 2D reconstructions for different time delays. We show that the mean of these values gives an effective heuristic for choosing the time delay. In addition, we show that this curvature-based heuristic is useful even in cases where the customary approach, which uses average mutual information, fails --- e.g., noisy or filtered data.

physics.data-an

Anomaly Detection in Paleoclimate Records using Permutation Entropy

Permutation entropy techniques can be useful in identifying anomalies in paleoclimate data records, including noise, outliers, and post-processing issues. We demonstrate this using weighted and unweighted permutation entropy of water-isotope records in a deep polar ice core. In one region of these isotope records, our previous calculations revealed an abrupt change in the complexity of the traces: specifically, in the amount of new information that appeared at every time step. We conjectured that this effect was due to noise introduced by an older laboratory instrument. In this paper, we validate that conjecture by re-analyzing a section of the ice core using a more-advanced version of the laboratory instrument. The anomalous noise levels are absent from the permutation entropy traces of the new data. In other sections of the core, we show that permutation entropy techniques can be used to identify anomalies in the raw data that are not associated with climatic or glaciological processes, but rather effects occurring during field work, laboratory analysis, or data post-processing. These examples make it clear that permutation entropy is a useful forensic tool for identifying sections of data that require targeted re-analysis---and can even be useful in guiding that analysis.

physics.data-an

Climate entropy production recorded in a deep Antarctic ice core

Paleoclimate records are extremely rich sources of information about the past history of the Earth system. Information theory, the branch of mathematics capable of quantifying the degree to which the present is informed by the past, provides a new means for studying these records. Here, we demonstrate that estimates of the Shannon entropy rate of the water-isotope data from the West Antarctica Ice Sheet (WAIS) Divide ice core, calculated using weighted permutation entropy (WPE), can bring out valuable new information from this record. We find that WPE correlates with accumulation, reveals possible signatures of geothermal heating at the base of the core, and clearly brings out laboratory and data-processing effects that are difficult to see in the raw data. For example, the signatures of Dansgaard-Oeschger events in the information record are small, suggesting that these abrupt warming events may not represent significant changes in the climate system dynamics. While the potential power of information theory in paleoclimatology problems is significant, the associated methods require careful handling and well-dated, high-resolution data. The WAIS Divide ice core is the first such record that can support this kind of analysis. As more high-resolution records become available, information theory will likely become a common forensic tool in climate science.

physics.geo-ph

Computational Topology Techniques for Characterizing Time-Series Data

Topological data analysis (TDA), while abstract, allows a characterization of time-series data obtained from nonlinear and complex dynamical systems. Though it is surprising that such an abstract measure of structure - counting pieces and holes - could be useful for real-world data, TDA lets us compare different systems, and even do membership testing or change-point detection. However, TDA is computationally expensive and involves a number of free parameters. This complexity can be obviated by coarse-graining, using a construct called the witness complex. The parametric dependence gives rise to the concept of persistent homology: how shape changes with scale. Its results allow us to distinguish time-series data from different systems - e.g., the same note played on different musical instruments.

cs.CG

Unix Memory Allocations are Not Poisson

In multitasking operating systems, requests for free memory are traditionally modeled as a stochastic counting process with independent, exponentially-distributed interarrival times because of the analytic simplicity such Poisson models afford. We analyze the distribution of several million unix page commits to show that although this approach could be valid over relatively long timespans, the behavior of the arrival process over shorter periods is decidedly not Poisson. We find that this result holds regardless of the originator of the request: unlike network packets, there is little difference between system- and user-level page-request distributions. We believe this to be due to the bursty nature of page allocations, which tend to occur in either small or extremely large increments. Burstiness and persistent variance have recently been found in self-similar processes in computer networks, but we show that although page commits are both bursty and possess high variance over long timescales, they are probably not self-similar. These results suggest that altogether different models are needed for fine-grained analysis of memory systems, an important consideration not only for understanding behavior but also for the design of online control systems.

cs.PF

Advanced Cyberinfrastructure for Science, Engineering, and Public Policy

Progress in many domains increasingly benefits from our ability to view the systems through a computational lens, i.e., using computational abstractions of the domains; and our ability to acquire, share, integrate, and analyze disparate types of data. These advances would not be possible without the advanced data and computational cyberinfrastructure and tools for data capture, integration, analysis, modeling, and simulation. However, despite, and perhaps because of, advances in "big data" technologies for data acquisition, management and analytics, the other largely manual, and labor-intensive aspects of the decision making process, e.g., formulating questions, designing studies, organizing, curating, connecting, correlating and integrating crossdomain data, drawing inferences and interpreting results, have become the rate-limiting steps to progress. Advancing the capability and capacity for evidence-based improvements in science, engineering, and public policy requires support for (1) computational abstractions of the relevant domains coupled with computational methods and tools for their analysis, synthesis, simulation, visualization, sharing, and integration; (2) cognitive tools that leverage and extend the reach of human intellect, and partner with humans on all aspects of the activity; (3) nimble and trustworthy data cyber-infrastructures that connect, manage a variety of instruments, multiple interrelated data types and associated metadata, data representations, processes, protocols and workflows; and enforce applicable security and data access and use policies; and (4) organizational and social structures and processes for collaborative and coordinated activity across disciplinary and institutional boundaries.

cs.CY

Big Data, Data Science, and Civil Rights

Advances in data analytics bring with them civil rights implications. Data-driven and algorithmic decision making increasingly determine how businesses target advertisements to consumers, how police departments monitor individuals or groups, how banks decide who gets a loan and who does not, how employers hire, how colleges and universities make admissions and financial aid decisions, and much more. As data-driven decisions increasingly affect every corner of our lives, there is an urgent need to ensure they do not become instruments of discrimination, barriers to equality, threats to social justice, and sources of unfairness. In this paper, we argue for a concrete research agenda aimed at addressing these concerns, comprising five areas of emphasis: (i) Determining if models and modeling procedures exhibit objectionable bias; (ii) Building awareness of fairness into machine learning methods; (iii) Improving the transparency and control of data- and model-driven decision making; (iv) Looking beyond the algorithm(s) for sources of bias and unfairness-in the myriad human decisions made during the problem formulation and modeling process; and (v) Supporting the cross-disciplinary scholarship necessary to do all of that well.

cs.CY