SearcharxivSearch

arXiv subjects

Eric Dolores-Cuenca

Publications and source records attributed to Eric Dolores-Cuenca.

10 recordsLinked to original sources

CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where researchers often need to assemble data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting into a reproducible pipeline, even when their primary research contribution concerns only one stage. CheMLFlow provides modular workflow components, ready-to-run reference pipelines, standardized artifacts, and evaluation outputs that reduce orchestration overhead and support benchmarking across methods and datasets. The platform is designed to be extensible, reproducible, and automation friendly, with pluggable representations and models, deterministic splits, explicit run artifacts, batch execution, and report generation. As scientific software increasingly moves toward agent assisted experimentation, CheMLFlow's configuration driven workflows and structured outputs also provide a practical interface for coding agents to help users construct experiments, inspect results, and summarize findings under human supervision. This article describes the system architecture, core workflows, and benchmarks that reach literature performance for quantum mechanical, physicochemical and bioactivity property prediction, and use cases involving time series datasets demonstrating applications beyond molecular chemistry datasets.

cs.LG

PCA-Enhanced Adaptive NVAR Framework for High-Resolution Sea Surface Temperature Forecasting in the East Sea

Accurate forecasting of sea surface temperature (SST) is essential for marine ecosystem monitoring, climate assessment, fisheries management, and operational ocean forecasting. While numerical ocean models provide reliable predictions, they are computationally expensive, and conventional machine learning methods often suffer from high-dimensional inputs and error accumulation during long-term autonomous forecasting. This study extends our previously proposed Adaptive Nonlinear Vector Autoregression (Adaptive NVAR) framework to high-resolution real-world SST prediction by integrating Principal Component Analysis (PCA) through Singular Value Decomposition (SVD). Daily SST fields from the GLORYS12V1 reanalysis dataset covering the East Sea, Yellow Sea, and East China Sea are compressed into a lower-dimensional latent representation that preserves the dominant spatial variability. The proposed reduced-order framework is evaluated using autonomous rolling forecasts up to a 90-day horizon and compared with Standard NVAR (Next Generation Reservoir Computing) and a Persistence baseline. Across all three regions, Adaptive NVAR consistently suppresses long-term error accumulation, achieving up to 96.52% improvement in mean squared error relative to Persistence while maintaining stable predictive performance throughout extended forecasting. Although the adaptive architecture incurs a higher one-time offline optimization cost than Standard NVAR, inference is completed in milliseconds, making the proposed framework an efficient and scalable approach for long-term, high-resolution ocean state forecasting.

cs.LG

Standardization of Post-Publication Code Verification by Journals is Possible with the Support of the Community

Reproducibility remains a challenge in machine learning research. While code and data availability requirements have become increasingly common, post-publication verification in journals is still limited and unformalized. This position paper argues that it is plausible for journals and conference proceedings to implement post-publication verification. We propose a modification to ACM pre-publication verification badges that allows independent researchers to submit post-publication code replications to the journal, leading to visible verification badges included in the article metadata. Each article may earn up to two badges, each linked to verified code in its corresponding public repository. We describe the motivation, related initiatives, a formal framework, the potential impact, possible limitations, and alternative views.

cs.LG

Adaptive Nonlinear Vector Autoregression: Robust Forecasting for Noisy Chaotic Time Series

Nonlinear vector autoregression (NVAR) and reservoir computing (RC) have shown promise in forecasting chaotic dynamical systems, such as the Lorenz-63 model and El Nino-Southern Oscillation. However, their reliance on fixed nonlinear transformations - polynomial expansions in NVAR or random feature maps in RC - limits their adaptability to high noise or complex real-world data. Furthermore, these methods also exhibit poor scalability in high-dimensional settings due to costly matrix inversion during optimization. We propose a data-adaptive NVAR model that combines delay-embedded linear inputs with features generated by a shallow, trainable multilayer perceptron (MLP). Unlike standard NVAR and RC models, the MLP and linear readout are jointly trained using gradient-based optimization, enabling the model to learn data-driven nonlinearities, while preserving a simple readout structure and improving scalability. Initial experiments across multiple chaotic systems, tested under noise-free and synthetically noisy conditions, showed that the adaptive model outperformed in predictive accuracy the standard NVAR, a leaky echo state network (ESN) - the most common RC model - and a hybrid ESN, thereby showing robust forecasting under noisy conditions.

cs.LG

Order Theory in the Context of Machine Learning

The paper ``Tropical Geometry of Deep Neural Networks'' by L. Zhang et al. introduces an equivalence between integer-valued neural networks (IVNN) with $\text{ReLU}_{t}$ and tropical rational functions, which come with a map to polytopes. Here, IVNN refers to a network with integer weights but real biases, and $\text{ReLU}_{t}$ is defined as $\text{ReLU}_{t}(x)=\max(x,t)$ for $t\in\mathbb{R}\cup\{-\infty\}$. For every poset with $n$ points, there exists a corresponding order polytope, i.e., a convex polytope in the unit cube $[0,1]^n$ whose coordinates obey the inequalities of the poset. We study neural networks whose associated polytope is an order polytope. We then explain how posets with four points induce neural networks that can be interpreted as $2\times 2$ convolutional filters. These poset filters can be added to any neural network, not only IVNN. Similarly to maxout, poset pooling filters update the weights of the neural network during backpropagation with more precision than average pooling, max pooling, or mixed pooling, without the need to train extra parameters. We report experiments that support our statements. We also define the structure of algebra over the operad of posets on poset neural networks and tropical polynomials. This formalism allows us to study the composition of poset neural network arquitectures and the effect on their corresponding Newton polytopes, via the introduction of the generalization of two operations on polytopes: the Minkowski sum and the convex envelope.

cs.CV

A poset version of Ramanujan results on Eulerian numbers and zeta values

We explore the operad of finite posets and its algebras. We use order polytopes to investigate the combinatorial properties of zeta values. By generalizing a family of zeta value identities, we demonstrate the applicability of this approach. In addition, we offer new proofs of some of Ramanujan's results on the properties of Eulerian numbers, interpreting his work as dealing with series inheriting the algebraic structure of disjoint unions of points. Finally, we establish a connection between our findings and the linear independence of zeta values.

math.CO

Virtual posets, shuffle algebras and associators

We provide a method to construct new associators out of Drinfel'd's KZ associator. We obtain two analytic families of associators whose coefficients we can describe explicitly by a generalization of multiple zeta values. The two families contain two different paths that deform the Drinfel'd KZ associator into the trivial associator 1. We show that both paths are injective, that is, all of the associators parametrized by them are different. Our construction is based on the observation that one can recover multiple polylogarithms as generating functions of order polynomials of certain formally constructed posets.

math.QA

An algebra over the operad of posets and structural binomial identities

We study generating functions of strict and non-strict order polynomials of series-parallel posets, called order series. These order series are closely related to Ehrhart series and h*-polynomials of the associated order polytopes. We explain how they can be understood as algebras over a certain operad of posets. Our main results are based on the fact that the order series of chains form a basis in the space of order series. This allows to reduce the search space of an algorithm that finds for a given power series f, if possible, a poset P such that f is the generating function of the order polynomial of P. In terms of Ehrhart theory of order polytopes, the coordinates with respect to this basis describe the number of (internal) simplices in the canonical triangulation of the order polytope of P. Furthermore, we derive a new proof of the reciprocity theorem of Stanley. As an application, we find new identities for binomial coefficients and for finite partitions that allow for empty sets, and we describe properties of the negative hypergeometric distribution.

math.CO

Polychrony as Chinampas

In this paper, we study the flow of signals through linear paths with the nonlinear condition that a node emits a signal when it receives external stimuli or when two incoming signals from other nodes arrive coincidentally with a combined amplitude above a fixed threshold. Sets of such nodes form a polychrony group and can sometimes lead to cascades. In the context of this work, cascades are polychrony groups in which the number of nodes activated as a consequence of other nodes is greater than the number of externally activated nodes. The difference between these two numbers is the so-called profit. Given the initial conditions, we predict the conditions for a vertex to activate at a prescribed time and provide an algorithm to efficiently reconstruct a cascade. We develop a dictionary between polychrony groups and graph theory. We call the graph corresponding to a cascade a chinampa. This link leads to a topological classification of chinampas. We enumerate the chinampas of profits zero and one and the description of a family of chinampas isomorphic to a family of partially ordered sets, which implies that the enumeration problem of this family is equivalent to computing the Stanley-order polynomials of those partially ordered sets.

math.CO