SearcharxivSearch

arXiv subjects

Antonio Celani

Publications and source records attributed to Antonio Celani.

At least 19 recordsLinked to original sources

Olfactory pursuit: catching a moving odor source in complex flows

Locating and intercepting a moving target from possibly delayed, intermittent sensory signals is a paradigmatic problem in decision-making under uncertainty, and a fundamental challenge for, e.g., animals seeking prey or mates and autonomous robotic systems. Odor signals are intermittent, strongly mixed by turbulent-like transport, and typically lag behind the true target position, thereby complicating localization. Here, we formulate olfactory pursuit as a partially observable Markov decision process in which an agent maintains a joint belief over the target's position and velocity. Using a discrete run-and-tumble model, we compute quasi-optimal policies by numerically solving the Bellman equation and benchmark them against well-established information-theoretic strategies such as Infotaxis. We show that purely exploratory policies are near-optimal when the target frequently reorients, but fail dramatically when the target exhibits persistent motion. We thus introduce a computationally efficient hybrid policy that combines the information-gain drive of Infotaxis with a "greedy" value function derived from an associated fully observable control problem. Our heuristic achieves near-optimal performance across all persistence times and substantially outperforms purely exploratory approaches. Moreover, our proposal demonstrates strong robustness even in more complex search scenarios, including continuous run-and-tumble prey motion with moderate persistence time, model mismatch, and more accurate plume dynamics representation. Our results identify predictive inference of target motion as the key ingredient for effective olfactory pursuit and provide a general framework for search in information-poor, dynamically evolving environments.

cs.RO

Stochastic modeling of long-legged ant A. gracilipes locomotion in laboratory experiments

Stochastic modeling of movement behavior provides a valuable way to understand how complex motion can be generated from relatively simple building blocks. Ants demonstrate sophisticated social behavior ranging from foraging to nest relocation; while emphasis is often placed on the communication methods used to synchronize individuals, the movement paradigms of those individuals are of tantamount importance. Here, we apply a stochastic modeling approach to better understand the movement of isolated long-legged ant (A. gracilipes) specimens, informed by extensive laboratory tracking experiments. We find that a combination of active Brownian and run-and-tumble models reproduces the trajectory statistics observed in experiments, both qualitatively and quantitatively. We identify reproducible probability distributions for the turn angles, run times, and waiting times across specimens, and find good agreement between analytical predictions and quantities empirically measured from the trajectories. Having such a model allows for a better understanding and predictions of movement ecology from both simulations and analytics, and even can give insight into the underlying generative mechanisms of motion and the ants' sensory systems.

q-bio.QM

Learning to flock in open space by avoiding collisions and staying together

We investigate the emergence of cohesive flocking in open, boundless space using a multi-agent reinforcement learning framework. Agents integrate positional and orientational information from their closest topological neighbours and learn to balance alignment and attractive interactions by optimizing a local cost function that penalizes both excessive separation and close-range crowding. The resulting Vicsek-like dynamics is robust to algorithmic implementation details and yields cohesive collective motion with high polar order. The optimal policy is dominated by strong aligning interactions when agents are sufficiently close to their neighbours, and a flexible combination of alignment and attraction at larger separations. We further characterize the internal structure and dynamics of the resulting groups using liquid-state metrics and neighbour exchange rates, finding qualitative agreement with empirical observations in starling flocks. These results suggest that flocking may emerge in groups of moving agents as an adaptive response to the biological imperatives of staying together while avoiding collisions.

cond-mat.soft

Information-directed sampling for bandits: a primer

The Multi-Armed Bandit problem provides a fundamental framework for analyzing the tension between exploration and exploitation in sequential learning. This paper explores Information Directed Sampling (IDS) policies, a class of heuristics that balance immediate regret against information gain. We focus on the tractable environment of two-state Bernoulli bandits as a minimal model to rigorously compare heuristic strategies against the optimal policy. We extend the IDS framework to the discounted infinite-horizon setting by introducing a modified information measure and a tuning parameter to modulate the decision-making behavior. We examine two specific problem classes: symmetric bandits and the scenario involving one fair coin. In the symmetric case we show that IDS achieves bounded cumulative regret, whereas in the one-fair-coin scenario the IDS policy yields a regret that scales logarithmically with the horizon, in agreement with classical asymptotic lower bounds. This work serves as a pedagogical synthesis, aiming to bridge concepts from reinforcement learning and information theory for an audience of statistical physicists.

cs.LG

Optimal trajectories for Bayesian olfactory search in turbulent flows: the low information limit and beyond

In turbulent flows, tracking the source of a passive scalar cue requires exploiting the limited information that can be gleaned from rare, stochastic encounters with the cue. When crafting a search policy, the most challenging and important decision is what to do in the absence of an encounter. In this work, we perform high-fidelity direct numerical simulations of a turbulent flow with a stationary source of tracer particles, and obtain quasi-optimal policies (in the sense of minimal average search time) with respect to the empirical encounter statistics. We study the trajectories under such policies and compare the results to those of the infotaxis heuristic. In the presence of a strong mean wind, the optimal motion in the absence of an encounter is zigzagging (akin to the well-known insect behavior "casting") followed by a return to the starting location. The zigzag motion generates characteristic $t^{1/2}$ scaling of the rms displacement envelope. By passing to the limit where the probability of detection vanishes, we connect these results to the classical linear search problem and derive an estimate of the tail of the arrival time pdf as a stretched exponential, which agrees with Monte Carlo results. We also discuss what happens as the wind speed decreases.

physics.flu-dyn

Exploring Bayesian olfactory search in realistic turbulent flows

The problem of tracking the source of a passive scalar in a turbulent flow is relevant to flying insect behavior and several other applications. Extensive previous work has shown that certain Bayesian strategies, such as "infotaxis," can be very effective for this difficult "olfactory search" problem. More recently, a quasi-optimal Bayesian strategy was computed under the assumption that encounters with the scalar are independent. However, the Bayesian approach has not been adequately studied in realistic flows which exhibit spatiotemporal correlations. In this work, we perform direct numerical simulations (DNS) of an incompressible flow at $\mathrm{Re}_λ\simeq150,$ while tracking Lagrangian particles which are emitted by a point source and imposing a uniform mean flow with several magnitudes (including zero). We extract the spatially-dependent statistics of encounters with the particles, which we use to build Bayesian policies, including generalized ("space-aware") infotactic heuristics and quasi-optimal policies. We then assess the relative performance of these policies when they are used to search using scalar cue data from the DNS, and in particular study how this performance depends on correlations between encounters. Among other results, we find that quasi-optimal strategies continue to outperform heuristics in the presence of strong mean flow but fail to do so in the absence of a mean flow. We also explore how to choose optimal search parameters, including the frequency and threshold concentration of observation.

physics.flu-dyn

Harvesting energy from turbulent winds with Reinforcement Learning

Airborne Wind Energy (AWE) is an emerging technology designed to harness the power of high-altitude winds, offering a solution to several limitations of conventional wind turbines. AWE is based on flying devices (usually gliders or kites) that, tethered to a ground station and driven by the wind, convert its mechanical energy into electrical energy by means of a generator. Such systems are usually controlled by manoeuvering the kite so as to follow a predefined path prescribed by optimal control techniques, such as model-predictive control. These methods are strongly dependent on the specific model at use and difficult to generalize, especially in unpredictable conditions such as the turbulent atmospheric boundary layer. Our aim is to explore the possibility of replacing these techniques with an approach based on Reinforcement Learning (RL). Unlike traditional methods, RL does not require a predefined model, making it robust to variability and uncertainty. Our experimental results in complex simulated environments demonstrate that AWE agents trained with RL can effectively extract energy from turbulent flows, relying on minimal local information about the kite orientation and speed relative to the wind.

cs.LG

Olfactory search

The task of olfactory search is ubiquitous in nature and in technology, from animals in the quest of food or of a mating partner, to robots searching for the source of hazardous fumes in a chemical plant. Here, we focus on the algorithmic approach to this task: we systematically review the different olfactory search strategies. Special emphasis is given to the formal description as a Partially Observable Markov Decision Processes, which allows the computation of optimal actions and helps clarifying the relationships between several effective heuristic search strategies.

physics.bio-ph

Optimal policies for Bayesian olfactory search in turbulent flows

In many practical scenarios, a flying insect must search for the source of an emitted cue which is advected by the atmospheric wind. On the macroscopic scales of interest, turbulence tends to mix the cue into patches of relatively high concentration over a background of very low concentration, so that the insect will only detect the cue intermittently and cannot rely on chemotactic strategies which simply climb the concentration gradient. In this work, we cast this search problem in the language of a partially observable Markov decision process (POMDP) and use the Perseus algorithm to compute strategies that are near-optimal with respect to the arrival time. We test the computed strategies on a large two-dimensional grid, present the resulting trajectories and arrival time statistics, and compare these to the corresponding results for several heuristic strategies, including (space-aware) infotaxis, Thompson sampling, and QMDP. We find that the near-optimal policy found by our implementation of Perseus outperforms all heuristics we test by several measures. We use the near-optimal policy to study how the search difficulty depends on the starting location. We discuss additionally the choice of initial belief and the robustness of the policies to changes in the environment. Finally, we present a detailed and pedagogical discussion about the implementation of the Perseus algorithm, including the benefits -- and pitfalls -- of employing a reward shaping function.

physics.flu-dyn

Optimal tracking strategies in a turbulent flow

Pursuing a drifting target in a turbulent flow is an extremely difficult task whenever the searcher has limited propulsion and maneuvering capabilities. Even in the case when the relative distance between pursuer and target stays below the turbulent dissipative scale, the chaotic nature of the trajectory of the target represents a formidable challenge. Here, we show how to successfully apply optimal control theory to find navigation strategies that overcome chaotic dispersion and allow the searcher to reach the target in a minimal time. We contrast the results of optimal control -- which requires perfect observability and full knowledge of the dynamics of the environment -- with heuristic algorithms that are reactive -- relying on local, instantaneous information about the flow. While the latter display significantly worse performances, optimally controlled pursuers can track the target for times much longer than the typical inverse Lyapunov exponent and are considerably more robust.

physics.flu-dyn

Taming Lagrangian Chaos with Multi-Objective Reinforcement Learning

We consider the problem of two active particles in 2D complex flows with the multi-objective goals of minimizing both the dispersion rate and the energy consumption of the pair. We approach the problem by means of Multi Objective Reinforcement Learning (MORL), combining scalarization techniques together with a Q-learning algorithm, for Lagrangian drifters that have variable swimming velocity. We show that MORL is able to find a set of trade-off solutions forming an optimal Pareto frontier. As a benchmark, we show that a set of heuristic strategies are dominated by the MORL solutions. We consider the situation in which the agents cannot update their control variables continuously, but only after a discrete (decision) time, $τ$. We show that there is a range of decision times, in between the Lyapunov time and the continuous updating limit, where Reinforcement Learning finds strategies that significantly improve over heuristics. In particular, we discuss how large decision times require enhanced knowledge of the flow, whereas for smaller $τ$ all a priori heuristic strategies become Pareto optimal.

physics.flu-dyn

Reinforcement learning for pursuit and evasion of microswimmers at low Reynolds number

We consider a model of two competing microswimming agents engaged in a pursue-evasion task within a low-Reynolds-number environment. Agents can only perform simple maneuvers and sense hydrodynamic disturbances, which provide ambiguous (partial) information about the opponent's position and motion. We frame the problem as a zero-sum game: The pursuer has to capture the evader in the shortest time, while the evader aims at deferring capture as long as possible. We show that the agents, trained via adversarial reinforcement learning, are able to overcome partial observability by discovering increasingly complex sequences of moves and countermoves that outperform known heuristic strategies and exploit the hydrodynamic environment.

physics.flu-dyn

Optimal collision avoidance in swarms of active Brownian particles

The effectiveness of collective navigation of biological or artificial agents requires to accommodate for contrasting requirements, such as staying in a group while avoiding close encounters and at the same time limiting the energy expenditure for manoeuvring. Here, we address this problem by considering a system of active Brownian particles in a finite two-dimensional domain and ask what is the control that realizes the optimal tradeoff between collision avoidance and control expenditure. We couch this problem in the language of optimal stochastic control theory and by means of a mean-field game approach we derive an analytic mean-field solution, characterized by a second-order phase transition in the alignment order parameter. We find that a mean-field version of a classical model for collective motion based on alignment interactions (Vicsek model) performs remarkably close to the optimal control. Our results substantiate the view that observed group behaviors may be explained as the result of optimizing multiple objectives and offer a theoretical ground for biomimetic algorithms used for artificial agents.

cond-mat.stat-mech

Collective olfactory search in a turbulent environment

Finding the distant source of an odor dispersed by a turbulent flow is a vital task for many organisms, either for foraging or for mating purposes. At the level of individual search, animals like moths have developed effective strategies to solve this very difficult navigation problem based on the noisy detection of odor concentration and wind velocity alone. When many individuals concurrently perform the same olfactory search task, without any centralized control, sharing information about the decisions made by the members of the group can potentially increase the performance. But how much of this information is actually valuable and exploitable for the collective task ? Here we show that, in a model of a swarm of agents inspired by moth behavior, there is an optimal way to blend the private information about odor and wind detections with the publicly available information about other agents' heading direction. At optimality, the time required for the first agent to reach the source is essentially the shortest flight time from the departure point to the target. Conversely, agents who discard public information are several fold slower and groups that do not put enough weight on private information perform even worse. Our results then suggest an efficient multi-agent olfactory search algorithm that could prove useful in robotics, for instance in the identification of sources of harmful volatile compounds.

physics.bio-ph

Learning to flock through reinforcement

Flocks of birds, schools of fish, insects swarms are examples of coordinated motion of a group that arises spontaneously from the action of many individuals. Here, we study flocking behavior from the viewpoint of multi-agent reinforcement learning. In this setting, a learning agent tries to keep contact with the group using as sensory input the velocity of its neighbors. This goal is pursued by each learning individual by exerting a limited control on its own direction of motion. By means of standard reinforcement learning algorithms we show that: i) a learning agent exposed to a group of teachers, i.e. hard-wired flocking agents, learns to follow them, and ii) that in the absence of teachers, a group of independently learning agents evolves towards a state where each agent knows how to flock. In both scenarios, i) and ii), the emergent policy (or navigation strategy) corresponds to the polar velocity alignment mechanism of the well-known Vicsek model. These results show that a) such a velocity alignment may have naturally evolved as an adaptive behavior that aims at minimizing the rate of neighbor loss, and b) prove that this alignment does not only favor (local) polar order, but it corresponds to best policy/strategy to keep group cohesion when the sensory input is limited to the velocity of neighboring agents. In short, to stay together, steer together.

physics.soc-ph

Generosity, selfishness and exploitation as optimal greedy strategies for resource sharing

Resource sharing outside the kinship bonds is rare. Besides humans, it occurs in chimpanzee, wild dogs and hyenas as well as in vampire bats. Resource sharing is an instance of animal cooperation, where an animal gives away part of the resources that it owns for the benefit of a recipient. Taking inspiration from blood-sharing in vampire bats, here show the emergence of generosity in a Markov game, which couples the resource sharing between two players with the gathering task of that resource. At variance with the classical evolutionary models for cooperation, the optimal strategies of this game can be potentially learned by animals during their life-time. The players act greedily, that is, they try to individually maximize only their personal income. Nonetheless, the analytical solution of the model shows that three non trivial optimal behaviours emerge depending on conditions. Besides the obvious case when players are selfish in their choice of resource division, there are conditions under which both players are generous. Moreover, we also found a range of situations in which one selfish player exploits another generous individual, for the satisfaction of both players. Our results show that resource sharing is favoured by three factors: a long time horizon over which the players try to optimize their own game, the similarity among players in their ability of performing the resource-gathering task, as well as by the availability of resources in the environment. These concurrent requirements lead to identify necessary conditions for the emergence of generosity.

q-bio.PE

Smart Inertial Particles

We performed a numerical study to train smart inertial particles to target specific flow regions with high vorticity through the use of reinforcement learning algorithms. The particles are able to actively change their size to modify their inertia and density. In short, using local measurements of the flow vorticity, the smart particle explores the interplay between its choices of size and its dynamical behaviour in the flow environment. This allows it to accumulate experience and learn approximately optimal strategies of how to modulate its size in order to reach the target high-vorticity regions. We consider flows with different complexities: a two-dimensional stationary Taylor-Green like configuration, a two-dimensional time-dependent flow, and finally a three-dimensional flow given by the stationary Arnold-Beltrami-Childress helical flow. We show that smart particles are able to learn how to reach extremely intense vortical structures in all the tackled cases.

physics.flu-dyn

Chemotaxis emerges as the optimal solution to cooperative search games

Cooperative search games are collective tasks where all agents share the same goal of reaching a target in the shortest time while limiting energy expenditure and avoiding collisions. Here we show that the equations that characterize the optimal strategy are identical to a long-known phenomenological model of chemotaxis, the directed motion of microorganisms guided by chemical cues. Within this analogy, the substance to which searchers respond acts as the memory over which agents share information about the environment. The actions of writing, erasing and forgetting are equivalent to production, consumption and degradation of chemoattractant. The rates at which these biochemical processes take place are tightly related to the parameters that characterize the decision-making problem, such as learning rate, costs for time, control, collisions and their trade-offs, as well as the attitude of agents toward risk. We establish a dictionary that maps notions from decision-making theory to biophysical observables in chemotaxis, and vice versa. Our results offer a fundamental explanation of why search algorithms that mimic microbial chemotaxis can be very effective and suggest how to optimize their performance.

physics.bio-ph