Searcharxiv⌕ Search

arXiv subjects

Ariana Strandburg-Peshkin

Publications and source records attributed to Ariana Strandburg-Peshkin.

9 recordsLinked to original sources

Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning

Bioacoustic call-type classification relies on costly expert annotation. Active learning can reduce this burden by selecting a small batch of segments for expert annotation and using the labeled segments for training the classifier. The setting is hard: the target calls are extremely sparse and the call-type distribution is long-tailed, so a tight budget must be spent on the few rare, informative segments. We propose BADGE-Greedy-DPP, a deterministic batch selector that greedily adds the segment whose BADGE gradient embedding most enlarges the volume spanned by the batch; because this log-volume objective is submodular, the greedy rule guarantees a batch value at least a (1-1/e) fraction of the optimum of this objective, a guarantee not provided by BADGE's existing k-means++ and MCMC DPP sampling heuristics. There is also a temporal granularity mismatch in the task. The acquisition function scores whole segments, yet the informative frames inside them are few. Uniform averaging therefore washes them out. The BADGE construction naturally addresses this mismatch when applied frame-wise, as prediction residuals weight the aggregated pseudo-gradient, so confidently predicted no-call frames contribute little while a single uncertain rare-call frame can still set the segment's direction. Across 10 runs on a sparse, imbalanced hyena call-type dataset, BADGE-Greedy-DPP achieves the best overall and rare-call-type performance among all compared query strategies, including MFFT, the strongest non-BADGE baseline, and the two vanilla BADGE traversals.

eess.AS↗

animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics

Bioacoustic research, vital for understanding animal behavior, conservation, and ecology, faces a monumental challenge: analyzing vast datasets where animal vocalizations are rare. While deep learning techniques are becoming standard, adapting them to bioacoustics remains difficult. We address this with animal2vec, an interpretable large transformer model, and a self-supervised training scheme tailored for sparse and unbalanced bioacoustic data. It learns from unlabeled audio and then refines its understanding with labeled data. Furthermore, we introduce and publicly release MeerKAT: Meerkat Kalahari Audio Transcripts, a dataset of meerkat (Suricata suricatta) vocalizations with millisecond-resolution annotations, the largest labeled dataset on non-human terrestrial mammals currently available. Our model outperforms existing methods on MeerKAT and the publicly available NIPS4Bplus birdsong dataset. Moreover, animal2vec performs well even with limited labeled data (few-shot learning). animal2vec and MeerKAT provide a new reference point for bioacoustic research, enabling scientists to analyze large amounts of data even with scarce ground truth information.

cs.SD↗

Modeling Animal Communication Using Multivariate Hawkes Processes with Additive Excitation and Multiplicative Inhibition

Animal acoustic communication often exhibits temporal dependence, with calls triggering or suppressing subsequent calls within and across call types, individuals, or species. While Hawkes processes provide a natural framework for modeling excitation, incorporating inhibition in multivariate settings can raise identifiability issues and complicate parameter interpretation. We propose a flexible class of multivariate Hawkes processes that combines additive excitation with multiplicative inhibition. This formulation preserves the branching process interpretation of excitation while reducing confounding between excitation and inhibition, and allows direct quantification of background and excitation contributions to the event rate. Bayesian inference is conducted via Markov chain Monte Carlo, and model adequacy is assessed using the random time change theorem. The proposed methodology is evaluated through simulation and applied to two acoustic communication datasets: group-living meerkats, for which we analyze three selected call types with distinct behavioral roles, and a two-species baleen whale dataset involving humpback and North Atlantic right whales. The meerkat analysis reveals significant within- and cross-type excitation with cross-type inhibition, whereas the whale data show evidence primarily of within-species excitation.

stat.AP↗

Few-shot bioacoustic event detection at the DCASE 2023 challenge

Few-shot bioacoustic event detection consists in detecting sound events of specified types, in varying soundscapes, while having access to only a few examples of the class of interest. This task ran as part of the DCASE challenge for the third time this year with an evaluation set expanded to include new animal species, and a new rule: ensemble models were no longer allowed. The 2023 few shot task received submissions from 6 different teams with F-scores reaching as high as 63% on the evaluation set. Here we describe the task, focusing on describing the elements that differed from previous years. We also take a look back at past editions to describe how the task has evolved. Not only have the F-score results steadily improved (40% to 60% to 63%), but the type of systems proposed have also become more complex. Sound event detection systems are no longer simple variations of the baselines provided: multiple few-shot learning methodologies are still strong contenders for the task.

cs.SD↗

Learning to detect an animal sound from five examples

Automatic detection and classification of animal sounds has many applications in biodiversity monitoring and animal behaviour. In the past twenty years, the volume of digitised wildlife sound available has massively increased, and automatic classification through deep learning now shows strong results. However, bioacoustics is not a single task but a vast range of small-scale tasks (such as individual ID, call type, emotional indication) with wide variety in data characteristics, and most bioacoustic tasks do not come with strongly-labelled training data. The standard paradigm of supervised learning, focussed on a single large-scale dataset and/or a generic pre-trained algorithm, is insufficient. In this work we recast bioacoustic sound event detection within the AI framework of few-shot learning. We adapt this framework to sound event detection, such that a system can be given the annotated start/end times of as few as 5 events, and can then detect events in long-duration audio -- even when the sound category was not known at the time of algorithm training. We introduce a collection of open datasets designed to strongly test a system's ability to perform few-shot sound event detections, and we present the results of a public contest to address the task. We show that prototypical networks are a strong-performing method, when enhanced with adaptations for general characteristics of animal sounds. We demonstrate that widely-varying sound event durations are an important factor in performance, as well as non-stationarity, i.e. gradual changes in conditions throughout the duration of a recording. For fine-grained bioacoustic recognition tasks without massive annotated training data, our results demonstrate that few-shot sound event detection is a powerful new method, strongly outperforming traditional signal-processing detection methods in the fully automated scenario.

cs.SD↗

Link updating strategies influence consensus decisions as a function of the direction of communication

Consensus decision-making in social groups strongly depends on communication links that determine to whom individuals send, and from whom they receive, information. Here, we ask how consensus decisions are affected by strategic updating of links and how this effect varies with the direction of communication. We quantified the co-evolution of link and opinion dynamics in a large population with binary opinions using mean-field numerical simulations of two voter-like models of opinion dynamics: an Incoming model (where individuals choose who to receive opinions from) and an Outgoing model (where individuals choose who to send opinions to). We show that individuals can bias group-level outcomes in their favor by breaking disagreeing links while receiving opinions (Incoming Model) and retaining disagreeing links while sending opinions (Outgoing Model). Importantly, these biases can help the population avoid stalemates and achieve consensus. However, the role of disagreement avoidance is diluted in the presence of strong preferences - highly stubborn individuals can shape decisions to favor their preferences, giving rise to non-consensus outcomes. We conclude that collectively changing communication structures can bias consensus decisions, as a function of the strength of preferences and the direction of communication.

physics.soc-ph↗

Local majority-with-inertia rule can explain global consensus dynamics in a network coordination game

We study how groups reach consensus by varying communication network structure and individual incentives. In 342 networks of seven individuals, single opinionated "leaders" can drive decision outcomes, but do not accelerate consensus formation, whereas conflicting opinions slow consensus. While networks with more links reach consensus faster, this advantage disappears under conflict. Unopinionated individuals make choices consistent with a local majority rule combined with "inertia" favouring their previous choice, while opinionated individuals favour their preferred option but yield under high peer or time pressure. Simulations show these individual rules can account for group patterns, and allow rapid consensus while preventing deadlocks.

physics.soc-ph↗

Coordination Event Detection and Initiator Identification in Time Series Data

Behavior initiation is a form of leadership and is an important aspect of social organization that affects the processes of group formation, dynamics, and decision-making in human societies and other social animal species. In this work, we formalize the "Coordination Initiator Inference Problem" and propose a simple yet powerful framework for extracting periods of coordinated activity and determining individuals who initiated this coordination, based solely on the activity of individuals within a group during those periods. The proposed approach, given arbitrary individual time series, automatically (1) identifies times of coordinated group activity, (2) determines the identities of initiators of those activities, and (3) classifies the likely mechanism by which the group coordination occurred, all of which are novel computational tasks. We demonstrate our framework on both simulated and real-world data: trajectories tracking of animals as well as stock market data. Our method is competitive with existing global leadership inference methods but provides the first approaches for local leadership and coordination mechanism classification. Our results are consistent with ground-truthed biological data and the framework finds many known events in financial data which are not otherwise reflected in the aggregate NASDAQ index. Our method is easily generalizable to any coordinated time-series data from interacting entities.

cs.SI↗

Creation of prompt and thin-sheet splashing by varying surface roughness or increasing air pressure

A liquid drop impacting a solid surface may splash by emitting a thin liquid sheet that subsequently breaks apart or by promptly ejecting droplets from the advancing liquid-solid contact line. Using high-speed imaging, we show that air pressure and surface roughness influence both splash mechanisms. Roughness increases prompt splashing at the advancing contact line but inhibits the formation of the thin sheet. If the air pressure is lowered, droplet ejection is suppressed not only during thin-sheet formation but for prompt splashing as well. The threshold pressure depends on impact velocity, liquid viscosity and surface roughness.

physics.flu-dyn↗