SearcharxivSearch

arXiv subjects

Luis Martí

Publications and source records attributed to Luis Martí.

9 recordsLinked to original sources

Planktonzilla: Multimodal dataset and models for understanding plankton ecosystems

Marine plankton underpin aquatic food webs and play a key role in global CO2 sequestration, making reliable species identification critical for understanding ocean health and climate feedbacks. Existing classification models perform well on individual collections but fail to generalize across instruments and environments due to isolated training datasets and inconsistent labels. To address this, we introduce Planktonzilla-17M, a unified dataset consolidating publicly available plankton image collections spanning thirteen imaging systems. It comprises 17.4 million images with standardized taxonomy and geo-environmental metadata, including 3.74 million plankton images spanning over 602 taxonomic classes, of which 201 are identified at the species level, making it the largest and most comprehensive plankton image dataset to date. Using this large-scale dataset, we perform a controlled comparison between supervised and CLIP-style image--text training on a shared ViT backbone. We find that a supervised classifier matches or exceeds CLIP-style training when trained using taxonomic lineage as text. We further observe that BioCLIP and BioCLIP2 perform poorly on plankton in zero-shot and few-shot settings. Leveraging Planktonzilla-17M improves plankton classification performance, highlighting the limitations of current biological foundation models in marine imaging domains.

cs.CV

From Mean-Field Limits to Semiclassical Concentration: Global Convergence of the Canonical Evolutionary Strategy

We address the issue of global convergence in stochastic continuous optimization. For that purpose, we formulate the Canonical Evolutionary Strategy (CES) as a controlled mathematical framework to analyze global convergence in evolutionary algorithms via the semiclassical limit of a Schr{ö}dinger-type replicator-mutator equation. We provide a rigorous hierarchy from a discrete individual-based dynamics to a deterministic mean-field limit, demonstrating that global convergence is governed by the principal eigenfunction of the underlying operator. This property, defined as Geometric Selection, naturally prioritizes robust, flat optima over narrow local traps, offering a mathematical justification for the ''survival of the flattest'' phenomenon. Moreover, unlike consensus-driven methods that are prone to premature variance collapse when the global minimizer resides outside the initial support, the replicator-mutator dynamics of CES facilitate intrinsic mass transport. High-dimensional benchmarks (d = 30) confirm this advantage, showing that CES achieves lower residual errors in shifted initialization scenarios where standard consensus-driven and gradient-based methods fail to migrate effectively. By shifting the focus from point-wise consensus to spectral concentration, our framework provides a robust theoretical foundation for global convergence in Evolution Strategies (ES) without the need for additional numerical heuristics.

cs.NE

Beyond the Critical Depth: The Metabolic and Physical Drivers of Phytoplankton Persistence in a Changing Ocean

While the classical Critical Depth Hypothesis (CDH) effectively explains the onset of blooms as transient instabilities, it does not fully capture the seasonal decoupling of biological rates and the long-term persistence of phytoplankton communities in fluctuating thermal environments. To address these limitations, we introduce a parsimonious framework that leverages the theory of non-autonomous dynamical systems to diagnose the stability of phytoplankton communities throughout the entire annual cycle. By linearizing the dynamics around the extinction equilibrium, we identify the invasion growth rate -formally the Floquet exponent-and derive the critical nutrient requirement ($γ$crit) as a bifurcation point for uniform persistence. Using end-of-the-century projections from the GFDL-ESM4 model under a high-emission scenario (SSP5-8.5), we identify a global regime shift characterized by a widespread expansion of metabolic-driven regimes, which increasingly displace regions where stability was historically governed by physical mixing. Relevance to Life Sciences. Quantitative analysis of system stability challenges CDH by demonstrating that metabolic constraints increasingly modulates phytoplankton persistence in a changing ocean. Our results, based on high-emission projections, reveal a profound physical-biological decoupling at the poles: while warming reduces the critical nutrient requirement ($γ$crit) facilitating persistence in previously marginal waters, this metabolic expansion is offset at poles. A 1:4 ratio between newly viable niches and ice-free deserts suggests that cryospheric retreat does not guarantee a proportional expansion of life. In addition, we identify the North Atlantic Subpolar Gyre as a ''metabolic refuge'' where mixing dynamics still anchor the ecosystem against global thermalization. By providing a ''radiography'' of the future ocean's complexity, this methodology offers a mechanistic basis to deconstruct how the dynamic balance between environmental energy and metabolic demands may determine the functional integrity of the marine biosphere under extreme anthropogenic forcing. Mathematical Content. The temperature dependence of biological rates is modeled using a thermodynamic equation, coupling population dynamics with seasonal variations in mixed layer depth and temperature. Given the non-autonomous nature of the system under annual forcing, we characterize the stability of the extinction equilibrium through its associated invasion growth rate. This rate is analytically derived as the Floquet exponent $λ$P , which provides a rigorous condition for uniform persistence (Theorem 3.2). The numerical analysis of this exponent, projected onto a global scale, quantifies the relative influence of environmental drivers on the stability threshold $γ$crit. This allows for the definition of the thermal dominance index (DT ), a metric that identifies the geographic transition from mixing-driven to metabolic-driven ecological control.

math.DS

Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America

Large Language Models (LLMs) exhibit inequalities with respect to various cultural contexts. Most prominent open-weights models are trained on Global North data and show prejudicial behavior towards other cultures. Moreover, there is a notable lack of resources to detect biases in non-English languages, especially from Latin America (Latam), a continent containing various cultures, even though they share a common cultural ground. We propose to leverage the content of Wikipedia, the structure of the Wikidata knowledge graph, and expert knowledge from social science in order to create a dataset of question/answer (Q/As) pairs, based on the different popular and social cultures of various Latin American countries. We create the LatamQA database of over 26k questions and associated answers extracted from 26k Wikipedia articles, and transformed into multiple-choice questions (MCQ) in Spanish and Portuguese, in turn translated to English. We use this MCQ to quantify the degree of knowledge of various LLMs and find out (i) a discrepancy in performances between the Latam countries, ones being easier than others for the majority of the models, (ii) that the models perform better in their original language, and (iii) that Iberian Spanish culture is better known than Latam one.

cs.CL

Towards model-free stellar chemical abundances. Potential applications in the search for chemically peculiar stars in large spectroscopic surveys

Chemical abundance determinations from stellar spectra are challenged by observational noise, limitations in stellar models, and departures from simplifying assumptions. While traditional and supervised machine learning methods have made remarkable progress in estimating atmospheric parameters and chemical compositions within existing physical models, these factors still constrain our ability to fully exploit the vast data sets provided by modern spectroscopic surveys. We aim to develop a self-supervised, disentangled representation learning framework that extracts chemically meaningful features directly from spectra, without relying on externally imposed label catalogs. We build a variational autoencoder-based representation learning model with physics-inspired structure: multiple decoders each focus on spectral regions dominated by a particular element, enforcing that each latent dimension maps to a single abundance. To evaluate the potential application of our framework, we trained and validated the model on low-resolution, low signal-to-noise synthetic spectra focusing on $\rm [Fe/H]$, $\rm [C/Fe]$, and $\rm [α/Fe]$. We then demonstrate how the trained model can be used to flag stars as chemically enhanced or depleted in these abundances based on their position within the latent distribution. Our model successfully learns a representation of spectra whose axes correlate tightly with the target abundances ($r=0.92\pm0.01$ for $\rm [Fe/H]$, $r=0.92\pm0.01$ for $\rm [C/Fe]$, $r=0.82\pm0.02$ for $\rm [α/Fe]$). The disentangled representations provide a robust means to distinguish stars based on their chemical properties, offering an efficient and scalable solution for large spectroscopic surveys.

astro-ph.SR

MATATA: Weakly Supervised End-to-End MAthematical Tool-Augmented Reasoning for Tabular Applications

Business documents often contain substantial tabular and textual information with numerical values, requiring mathematical reasoning for effective document understanding. While Small Language Models (SLMs) still struggle at this task, tool-augmented multi-step agents perform better, at the cost of relying on closed-source or larger models, external data, or extensive prompt-engineering. This work introduces MATATA, a novel weakly supervised end-to-end approach to train multi-step reasoning language agents for document tabular applications. MATATA presents an annotation-free paradigm for each agent to enhance 3.8B/8B SLMs. During its two-stage training, MATATA uses the final outcome of the multi-step reasoning chain as weak supervision. This approach avoids having to individually supervise each intermediate agent in the reasoning chain. By employing an adaptive planner and shared tools across different datasets, MATATA shows robust performance. Experiments demonstrate that MATATA achieves state-of-the-art on FinQA, and on TAT-QA among reasoning methods based on open-source SLMs. Although being SLM-based, MATATA closely matches GPT-4-based frameworks on TabMWP. This novel weakly supervised approach enables training an end-to-end multi-step reasoning agent without intermediate supervision, supporting future developments of cost-effective powerful agentic systems.

cs.LG

A baseline on the relation between chemical patterns and birth stellar cluster

The chemical composition of a star's atmosphere reflects the chemical composition of its birth environment. Therefore, it should be feasible to recognize stars born together that have scattered throughout the galaxy, solely based on their chemistry. This concept, known as "strong chemical tagging", is a major objective of spectroscopic studies, but has yet to yield the anticipated results. We assess the existence and the robustness of the relation between chemical abundances and birth place using known member stars of open clusters. We followed a supervised machine learning approach, using chemical abundances obtained from APOGEE DR17, observed open clusters as labels and different data preprocessing techniques. We found that open clusters can be recovered with any classifier and on data whose features are not carefully selected. In the sample with no field stars, we obtain an average accuracy of $75.2\%$ and we find that the prediction accuracy depends mostly on the uncertainties of the chemical abundances. When field stars outnumber the cluster members, the performance degrades. Our results show the difficulty of recovering birth clusters using chemistry alone, even in a supervised scenario. This clearly challenges the feasibility of strong chemical tagging. Nevertheless, including information about ages could potentially enhance the possibility of recovering birth clusters.

astro-ph.GA

Reinforcement-learning robotic sailboats: simulator and preliminary results

This work focuses on the main challenges and problems in developing a virtual oceanic environment reproducing real experiments using Unmanned Surface Vehicles (USV) digital twins. We introduce the key features for building virtual worlds, considering using Reinforcement Learning (RL) agents for autonomous navigation and control. With this in mind, the main problems concern the definition of the simulation equations (physics and mathematics), their effective implementation, and how to include strategies for simulated control and perception (sensors) to be used with RL. We present the modeling, implementation steps, and challenges required to create a functional digital twin based on a real robotic sailing vessel. The application is immediate for developing navigation algorithms based on RL to be applied on real boats.

cs.RO

Towards Optimally Weighted Physics-Informed Neural Networks in Ocean Modelling

The carbon pump of the world's ocean plays a vital role in the biosphere and climate of the earth, urging improved understanding of the functions and influences of the ocean for climate change analyses. State-of-the-art techniques are required to develop models that can capture the complexity of ocean currents and temperature flows. This work explores the benefits of using physics-informed neural networks (PINNs) for solving partial differential equations related to ocean modeling; such as the Burgers, wave, and advection-diffusion equations. We explore the trade-offs of using data vs. physical models in PINNs for solving partial differential equations. PINNs account for the deviation from physical laws in order to improve learning and generalization. We observed how the relative weight between the data and physical model in the loss function influence training results, where small data sets benefit more from the added physics information.

cs.LG