SearcharxivSearch

arXiv subjects

Leonel Aguilar

Publications and source records attributed to Leonel Aguilar.

5 recordsLinked to original sources

Does Your Neural Network Extrapolate? Feature Engineering as Identifiability Bias for OOD Generalization

Successful deep neural networks discover salient features of data. We show when and why they fail to learn out-of-distribution (OOD)-relevant representations from an in-distribution (ID) training window. This requires decoupling feature learning from data-generating-process (DGP) identifiability. From a single training window, OOD extrapolation is non-identifiable: infinitely many DGPs are $\varepsilon$-observationally equivalent on the training data but diverge arbitrarily outside it, and no in-distribution criterion alone reliably breaks the tie. A structural commitment, the feature map, label map, and model class $(φ, ψ, \mathcal{M})$, dictates the assumed DGP and governs OOD generalization while leaving ID performance essentially unchanged. When architecture, pretraining, augmentation, input formats, or domain knowledge implicitly inject the missing commitment, the model succeeds. When it cannot infer OOD-relevant structure from ID evidence, it fails. Changing only the representation can make the same architecture, at the same in-distribution loss, differ by ${\sim}520\times$ out of distribution. When the commitment is correct and identifiable, OOD error vanishes. For example, Fourier coordinates turn periodic extrapolation into interpolation on $\mathbb{S}^1$. The same mechanism predicts outcomes in three natural-science settings (mass-action chemistry; Kepler's-third-law exoplanet prediction, $n=2{,}362$; and cross-species coding-DNA detection) and in a 264-run positional-encoding study across Transformer, Mamba, and S4D. Finally, a controlled study shows: correct features are necessary but not sufficient. The model class must express the target, and the transformed training data must cover the relevant representation space.

cs.LG

IUMENTA: A generic framework for animal digital twins within the Open Digital Twin Platform

IUMENTA (Latin for livestock) is an innovative software framework designed to construct and simulate digital twins of animals. By leveraging the powerful capability of the Open Digital Twin Platform (ODTP) alongside advanced software sensors, IUMENTA offers researchers a user-friendly tool to seamlessly develop adaptive digital replicas of animal-based processes. This framework establishes a dynamic ecosystem that integrates insights from diverse experiments, consequently enhancing our understanding of animal behavioural and physiological responses. Through real-time tracking of an animal's energy balance. IUMENTA provides valuable insights into metabolic rates, nutritional needs, emotional states, and overall well-being of animals. In this article, we explore the application of the IUMENTA framework in developing a digital twin focused on the animal's energy balance. IUMENTA includes the EnergyTag system, a state-of-the-art wearable software sensor, which facilitates real-time monitoring of energy expenditure, allowing for continuous updates and personalisation of the energy balance digital twin.

cs.OH

The Hitchhiker's Guide to Fused Twins: A Review of Access to Digital Twins in situ in Smart Cities

Smart Cities already surround us, and yet they are still incomprehensibly far from directly impacting everyday life. While current Smart Cities are often inaccessible, the experience of everyday citizens may be enhanced with a combination of the emerging technologies Digital Twins (DTs) and Situated Analytics. DTs represent their Physical Twin (PT) in the real world via models, simulations, (remotely) sensed data, context awareness, and interactions. However, interaction requires appropriate interfaces to address the complexity of the city. Ultimately, leveraging the potential of Smart Cities requires going beyond assembling the DT to be comprehensive and accessible. Situated Analytics allows for the anchoring of city information in its spatial context. We advance the concept of embedding the DT into the PT through Situated Analytics to form Fused Twins (FTs). This fusion allows access to data in the location that it is generated in an embodied context that can make the data more understandable. Prototypes of FTs are rapidly emerging from different domains, but Smart Cities represent the context with the most potential for FTs in the future. This paper reviews DTs, Situated Analytics, and Smart Cities as the foundations of FTs. Regarding DTs, we define five components (Physical, Data, Analytical, Virtual, and Connection environments) that we relate to several cognates (i.e., similar but different terms) from existing literature. Regarding Situated Analytics, we review the effects of user embodiment on cognition and cognitive load. Finally, we classify existing partial examples of FTs from the literature and address their construction from Augmented Reality, Geographic Information Systems, Building/City Information Models, and DTs and provide an overview of future direction

cs.CY

Experiments as Code: A Concept for Reproducible, Auditable, Debuggable, Reusable, & Scalable Experiments

A common concern in experimental research is the auditability and reproducibility of experiments. Experiments are usually designed, provisioned, managed, and analyzed by diverse teams of specialists (e.g., researchers, technicians and engineers) and may require many resources (e.g. cloud infrastructure, specialized equipment). Even though researchers strive to document experiments accurately, this process is often lacking, making it hard to reproduce them. Moreover, when it is necessary to create a similar experiment, very often we end up "reinventing the wheel" as it is easier to start from scratch than trying to reuse existing work, thus losing valuable embedded best practices and previous experiences. In behavioral studies this has contributed to the reproducibility crisis. To tackle this challenge, we propose the "Experiments as Code" paradigm, where the whole experiment is not only documented but additionally the automation code to provision, deploy, manage, and analyze it is provided. To this end we define the Experiments as Code concept, provide a taxonomy for the components of a practical implementation, and provide a proof of concept with a simple desktop VR experiment that showcases the benefits of its "as code" representation, i.e., reproducibility, auditability, debuggability, reusability, and scalability.

cs.CY

How intelligence can change the course of evolution

The effect of phenotypic plasticity on evolution, the so-called Baldwin effect, has been studied extensively for more than 100 years. Plasticity is known to influence the speed of evolution towards a specific genetic configuration, but whether it also influences what that genetic configuration is, is still an open question. This question is investigated, in an environment where the distribution of resources follows seasonal cycles, both analytically and experimentally by means of an agent-based model of a foraging task. Individuals can either specialize to foraging only one specific resource type or generalize to foraging all resource types at a low success rate. It is found that the introduction of learning, one instance of phenotypic plasticity, changes what genetic configuration evolves. Specifically, the genome of learning agents evolves a predisposition to adapt quickly to changes in the resource distribution, under the same conditions for which non-learners would evolve a predisposition to maximize the foraging efficiency for a specific resource type. This paper expands the literature at the interface between Biology and Machine Learning by identifying the Baldwin effects in cyclically-changing environments and demonstrating that learning can change the outcome of evolution.

q-bio.PE