SearcharxivSearch

arXiv subjects

Ethan P. White

Publications and source records attributed to Ethan P. White.

10 recordsLinked to original sources

10 quick tips for making your software outlive your job

Loss of key personnel has always been a risk for research software projects. Key members of the team may have to step away due to illness or burnout, to care for a family member, from a loss of financial support, or because their career is going in a new direction. Today, though, political and financial changes are putting large numbers of researchers out of work simultaneously, potentially leaving large amounts of research software abandoned. This article presents ten tips to help researchers ensure that the software they have built will continue to be usable after they have left their present job -- whether in the course of voluntary career moves or researcher mobility, but particularly in cases of involuntary departure due to political or institutional changes.

cs.SE

RandCrowns: A Quantitative Metric for Imprecisely Labeled Tree Crown Delineation

Supervised methods for object delineation in remote sensing require labeled ground-truth data. Gathering sufficient high quality ground-truth data is difficult, especially when targets are of irregular shape or difficult to distinguish from background or neighboring objects. Tree crown delineation provides key information from remote sensing images for forestry, ecology, and management. However, tree crowns in remote sensing imagery are often difficult to label and annotate due to irregular shape, overlapping canopies, shadowing, and indistinct edges. There are also multiple approaches to annotation in this field (e.g., rectangular boxes vs. convex polygons) that further contribute to annotation imprecision. However, current evaluation methods do not account for this uncertainty in annotations, and quantitative metrics for evaluation can vary across multiple annotators. In this paper, we address these limitations by developing an adaptation of the Rand index for weakly-labeled crown delineation that we call RandCrowns. Our new RandCrowns evaluation metric provides a method to appropriately evaluate delineated tree crowns while taking into account imprecision in the ground-truth delineations. The RandCrowns metric reformulates the Rand index by adjusting the areas over which each term of the index is computed to account for uncertain and imprecise object delineation labels. Quantitative comparisons to the commonly used intersection over union method shows a decrease in the variance generated by differences among multiple annotators. Combined with qualitative examples, our results suggest that the RandCrowns metric is more robust for scoring target delineations in the presence of uncertainty and imprecision in annotations that are inherent to tree crown delineation.

cs.CV

Difference Necklaces

An $(a,b)$-difference necklace of length $n$ is a circular arrangement of the integers $0, 1, 2, \ldots , n-1$ such that any two neighbours have absolute difference $a$ or $b$. We prove that, subject to certain conditions on $a$ and $b$, such arrangements exist, and provide recurrence relations for the number of $(a,b)$-difference necklaces for $( a, b ) = ( 1, 2 )$, $( 1, 3 )$, $( 2, 3 )$ and $( 1, 4 )$. Using techniques similar to those employed for enumerating Hamiltonian cycles in certain families of graphs, we obtain these explicit recurrence relations and prove that the number of $(a,b)$-difference necklaces of length $n$ satisfies a linear recurrence relation for all permissible values $a$ and $b$. Our methods generalize to necklaces where an arbitrary number of differences is allowed.

math.CO

On the directions determined by a Cartesian product in an affine Galois plane

We prove that the number of directions contained in a set of the form $A \times B \subset AG(2,p)$, where $p$ is prime, is at least $|A||B| - \min\{|A|,|B|\} + 2$. Here $A$ and $B$ are subsets of $GF(p)$ each with at least two elements and $|A||B| <p$. This bound is tight for an infinite class of examples. Our main tool is the use of the Rédei polynomial with Szőnyi's extension. As an application of our main result, we obtain an upper bound on the clique number of a Paley graph, matching the current best bound obtained recently by Hanson and Petridis.

math.CO

Comparing process-based and constraint-based approaches for modeling macroecological patterns

Ecological patterns arise from the interplay of many different processes, and yet the emergence of consistent phenomena across a diverse range of ecological systems suggests that many patterns may in part be determined by statistical or numerical constraints. Differentiating the extent to which patterns in a given system are determined statistically, and where it requires explicit ecological processes, has been difficult. We tackled this challenge by directly comparing models from a constraint-based theory, the Maximum Entropy Theory of Ecology (METE) and models from a process-based theory, the size-structured neutral theory (SSNT). Models from both theories were capable of characterizing the distribution of individuals among species and the distribution of body size among individuals across 76 forest communities. However, the SSNT models consistently yielded higher overall likelihood, as well as more realistic characterizations of the relationship between species abundance and average body size of conspecific individuals. This suggests that the details of the biological processes contain additional information for understanding community structure that are not fully captured by the METE constraints in these systems. Our approach provides a first step towards differentiating between process- and constraint-based models of ecological systems and a general methodology for comparing ecological models that make predictions for multiple patterns.

q-bio.PE

A process-independent explanation for the general form of Taylor's Law

Taylors Law (TL) describes the scaling relationship between the mean and variance of populations as a power-law. TL is widely observed in ecological systems across space and time with exponents varying largely between 1 and 2. Many ecological explanations have been proposed for TL but it is also commonly observed outside ecology. We propose that TL arises from the constraining influence of two primary variables: the number of individuals and the number of censuses or sites. We show that most possible configurations of individuals among censuses or sites produce the power-law form of TL with exponents between 1 and 2. This feasible set approach suggests that TL is a statistical pattern driven by two constraints, providing an a priori explanation for this ubiquitous pattern. However, the exact form of any specific mean-variance relationship cannot be predicted in this way, i.e., this approach does a poor job of predicting variation in the exponent, suggesting that TL may still contain ecological information.

q-bio.PE

A strong test of the Maximum Entropy Theory of Ecology

The Maximum Entropy Theory of Ecology (METE) is a unified theory of biodiversity that predicts a large number of macroecological patterns using only information on the species richness, total abundance, and total metabolic rate of the community. We evaluated four major predictions of METE simultaneously at an unprecedented scale using data from 60 globally distributed forest communities including over 300,000 individuals and nearly 2000 species. METE successfully captured 96% and 89% of the variation in the species abundance distribution and the individual size distribution, but performed poorly when characterizing the size-density relationship and intraspecific distribution of individual size. Specifically, METE predicted a negative correlation between size and species abundance, which is weak in natural communities. By evaluating multiple predictions with large quantities of data, our study not only identifies a mismatch between abundance and body size in METE, but also demonstrates the importance of conducting strong tests of ecological theories.

q-bio.PE

An empirical evaluation of four variants of a universal species-area relationship

The Maximum Entropy Theory of Ecology (METE) predicts a universal species-area relationship (SAR) that can be fully characterized using only the total abundance (N) and species richness (S) at a single spatial scale. This theory has shown promise for characterizing scale dependence in the SAR. However, there are currently four different approaches to applying METE to predict the SAR and it is unclear which approach should be used due to a lack of empirical evaluation. Specifically, METE can be applied recursively or a non-recursively and can use either a theoretical or observed species-abundance distribution (SAD). We compared the four different combinations of approaches using empirical data from 16 datasets containing over 1000 species and 300,000 individual trees and herbs. In general, METE accurately downscaled the SAR (R^2> 0.94), but the recursive approach consistently under-predicted richness, and METEs accuracy did not depend strongly on using the observed or predicted SAD. This suggests that best approach to scaling diversity using METE is to use a combination of non-recursive scaling and the theoretical abundance distribution, which allows predictions to be made across a broad range of spatial scales with only knowledge of the species richness and total abundance at a single scale.

q-bio.QM

Best Practices for Scientific Computing

Scientists spend an increasing amount of time building and using software. However, most scientists are never taught how to do this efficiently. As a result, many are unaware of tools and practices that would allow them to write more reliable and maintainable code with less effort. We describe a set of best practices for scientific software development that have solid foundations in research and experience, and that improve scientists' productivity and the reliability of their software.

cs.MS