SearcharxivSearch

arXiv subjects

Christopher M. White

Publications and source records attributed to Christopher M. White.

10 recordsLinked to original sources

Scalar Dispersion from Wall-Mounted Cylinders at Large Reynolds Number: Plume Transitions and Regime Classification

This study presents a comprehensive experimental investigation of scalar dispersion from the free end of wall-mounted cylindrical obstacles immersed in a large-Reynolds-number turbulent boundary layer. A key focus is the characterization of transition behavior between distinct dispersion regimes: elevated plumes (EP), ground-level plumes (GLP), and ground-level sources (GLS). Experiments systematically vary the primary and secondary aspect ratios ($AR_1, AR_2$) and the velocity ratio ($ r$) to explore their effects on the evolution of scalar plumes. Plume classification is governed by the non-dimensional parameter $\tilde{h}_s / δ_{cz}$, which quantifies the progressive interaction between the plume and the ground. Here, $\tilde{h}_s$ denotes the effective source height and $δ_{cz}$, the vertical plume half-width. Detailed concentration measurements demonstrate that the EP--GLP--GLS transitions substantially modify both vertical and lateral dispersion characteristics. The measurements reveal systematic departures from classical dispersion-coefficient scaling. To assess the capability of existing models under these conditions, the experimentally determined dispersion coefficients are used to evaluate the Gaussian Dispersion Model (GDM) and a Wall Similarity Model (WSM). The GDM captures general trends but deviates in specific regimes, whereas the WSM offers improved representation under GLS conditions. The resulting dataset, grounded in systematic laboratory measurements, establishes a critical benchmark for validating numerical simulations and informing the development of next-generation predictive models. Finally, leveraging these results, a concise data-informed predictive framework is introduced that captures the EP--GLP--GLS transitions and provides first-order estimates of ground-level concentration across geometric and momentum-ratio parameter space.

physics.flu-dyn

Simple Lifelong Learning Machines

In lifelong learning, data are used to improve performance not only on the present task, but also on past and future (unencountered) tasks. While typical transfer learning algorithms can improve performance on future tasks, their performance on prior tasks degrades upon learning new tasks (called forgetting). Many recent approaches for continual or lifelong learning have attempted to maintain performance on old tasks given new tasks. But striving to avoid forgetting sets the goal unnecessarily low. The goal of lifelong learning should be to use data to improve performance on both future tasks (forward transfer) and past tasks (backward transfer). In this paper, we show that a simple approach -- representation ensembling -- demonstrates both forward and backward transfer in a variety of simulated and benchmark data scenarios, including tabular, vision (CIFAR-100, 5-dataset, Split Mini-Imagenet, and Food1k), and speech (spoken digit), in contrast to various reference algorithms, which typically failed to transfer either forward or backward, or both. Moreover, our proposed approach can flexibly operate with or without a computational budget.

cs.AI

Discovering a change point and piecewise linear structure in a time series of organoid networks via the iso-mirror

Recent advancements have been made in the development of cell-based in-vitro neuronal networks, or organoids. In order to better understand the network structure of these organoids, a super-selective algorithm has been proposed for inferring the effective connectivity networks from multi-electrode array data. In this paper, we apply a novel statistical method called spectral mirror estimation to the time series of inferred effective connectivity organoid networks. This method produces a one-dimensional iso-mirror representation of the dynamics of the time series of the networks which exhibits a piecewise linear structure. A classical change point algorithm is then applied to this representation, which successfully detects a change point coinciding with the neuroscientifically significant time inhibitory neurons start appearing and the percentage of astrocytes increases dramatically. This finding demonstrates the potential utility of applying the iso-mirror dynamic structure discovery method to inferred effective connectivity time series of organoid networks.

stat.AP

When are Deep Networks really better than Decision Forests at small sample sizes, and how?

Deep networks and decision forests (such as random forests and gradient boosted trees) are the leading machine learning methods for structured and tabular data, respectively. Many papers have empirically compared large numbers of classifiers on one or two different domains (e.g., on 100 different tabular data settings). However, a careful conceptual and empirical comparison of these two strategies using the most contemporary best practices has yet to be performed. Conceptually, we illustrate that both can be profitably viewed as "partition and vote" schemes. Specifically, the representation space that they both learn is a partitioning of feature space into a union of convex polytopes. For inference, each decides on the basis of votes from the activated nodes. This formulation allows for a unified basic understanding of the relationship between these methods. Empirically, we compare these two strategies on hundreds of tabular data settings, as well as several vision and auditory settings. Our focus is on datasets with at most 10,000 samples, which represent a large fraction of scientific and biomedical datasets. In general, we found forests to excel at tabular and structured data (vision and audition) with small sample sizes, whereas deep nets performed better on structured data with larger sample sizes. This suggests that further gains in both scenarios may be realized via further combining aspects of forests and networks. We will continue revising this technical report in the coming months with updated results.

cs.LG

Leveraging semantically similar queries for ranking via combining representations

In modern ranking problems, different and disparate representations of the items to be ranked are often available. It is sensible, then, to try to combine these representations to improve ranking. Indeed, learning to rank via combining representations is both principled and practical for learning a ranking function for a particular query. In extremely data-scarce settings, however, the amount of labeled data available for a particular query can lead to a highly variable and ineffective ranking function. One way to mitigate the effect of the small amount of data is to leverage information from semantically similar queries. Indeed, as we demonstrate in simulation settings and real data examples, when semantically similar queries are available it is possible to gainfully use them when ranking with respect to a particular query. We describe and explore this phenomenon in the context of the bias-variance trade off and apply it to the data-scarce settings of a Bing navigational graph and the Drosophila larva connectome.

cs.LG

Workgroup Mapping: Visual Analysis of Collaboration Culture

The digital transformation of work presents new opportunities to understand how informal workgroups organize around the dynamic needs of organizations, potentially in contrast to the formal, static, and idealized hierarchies depicted by org charts. We present a design study that spans multiple enabling capabilities for the visual mapping and analysis of organizational workgroups, including metrics for quantifying two dimensions of collaboration culture: the fluidity of collaborative relationships (measured using network machine learning) and the freedom with which workgroups form across organizational boundaries. These capabilities come together to create a turnkey pipeline that combines the analysis of a target organization, the generation of data graphics and statistics, and their integration in a template-based presentation that enables narrative visualization of results. Our metrics and visuals have supported hundreds of presentations to executives of major US-based and multinational organizations, while our engineering practices have created an ensemble of standalone tools with broad relevance to visualization and visual analytics. We present our work as an example of applied visual analytics research, describing the design iterations that allowed us to move from experimentation to production, as well as the perspectives of the research team and the customer-facing team at each stage in this process.

cs.HC

Design of a Privacy-Preserving Data Platform for Collaboration Against Human Trafficking

Case records on victims of human trafficking are highly sensitive, yet the ability to share such data is critical to evidence-based practice and policy development across government, business, and civil society. We present new methods to anonymize, publish, and explore such data, implemented as a pipeline generating three artifacts: (1) synthetic data mitigating the privacy risk that published attribute combinations might be linked to known individuals or groups; (2) aggregate data mitigating the utility risk that synthetic data might misrepresent statistics needed for official reporting; and (3) visual analytics interfaces to both datasets mitigating the accessibility risk that privacy mechanisms or analysis tools might not be understandable and usable by all stakeholders. We present our work as a design study motivated by the goal of transforming how the world's largest database of identified victims is made available for global collaboration against human trafficking.

cs.HC

A self-sustaining process theory for uniform momentum zones and internal shear layers in high Reynolds number shear flows

Many exact coherent states (ECS) arising in wall-bounded shear flows have an asymptotic structure at extreme Reynolds number Re in which the effective Reynolds number governing the streak and roll dynamics is O(1). Consequently, these viscous ECS are not suitable candidates for quasi-coherent structures away from the wall that necessarily are inviscid in the mean. Specifically, viscous ECS cannot account for the singular nature of the inertial domain, where the flow self-organizes into uniform momentum zones (UMZs) separated by internal shear layers and the instantaneous streamwise velocity develops a staircase-like profile. In this investigation, a large-Re asymptotic analysis is performed to explore the potential for a three-dimensional, short streamwise- and spanwise-wavelength instability of the embedded shear layers to sustain a spatially-distributed array of much larger-scale, effectively inviscid streamwise roll motions. In contrast to other self-sustaining process theories, the rolls are sufficiently strong to differentially homogenize the background shear flow, thereby providing a mechanistic explanation for the formation and maintenance of UMZs and interlaced shear layers that respects the leading-order balance structure of the mean dynamics.

physics.flu-dyn

Likelihood-based semi-supervised model selection with applications to speech processing

In conventional supervised pattern recognition tasks, model selection is typically accomplished by minimizing the classification error rate on a set of so-called development data, subject to ground-truth labeling by human experts or some other means. In the context of speech processing systems and other large-scale practical applications, however, such labeled development data are typically costly and difficult to obtain. This article proposes an alternative semi-supervised framework for likelihood-based model selection that leverages unlabeled data by using trained classifiers representing each model to automatically generate putative labels. The errors that result from this automatic labeling are shown to be amenable to results from robust statistics, which in turn provide for minimax-optimal censored likelihood ratio tests that recover the nonparametric sign test as a limiting case. This approach is then validated experimentally using a state-of-the-art automatic speech recognition system to select between candidate word pronunciations using unlabeled speech data that only potentially contain instances of the words under test. Results provide supporting evidence for the utility of this approach, and suggest that it may also find use in other applications of machine learning.

stat.ML