SearcharxivSearch

arXiv subjects

Luis Damiano

Publications and source records attributed to Luis Damiano.

4 recordsLinked to original sources

Controller-Augmented Hidden Markov Models: A Computational Framework for Constrained Sequential Inference

Hidden Markov models are foundational for sequential inference, but their Markovian assumption fails under pathwise constraints such as precedence requirements, visitation cardinalities, or monotonic state progression, which induce long-range dependencies that invalidate standard dynamic programming algorithms. To deal with this, we present Controller-Augmented Hidden Markov Models (CHMMs), a framework that compiles each constraint into a finite-state controller tracking the minimal sufficient history, after which standard forward--backward and Viterbi recursions on the augmented chain compute exact constrained posteriors and maximum a posteriori paths in both discrete and continuous time, the latter through uniformization. We establish four theoretical guarantees: exactness of constrained inference, monotone ascent of constrained EM, inference complexity linear in the controller cardinality, and a total-variation bound under constraint misspecification. A catalog of controller encodings covering 11 constraint families across the ordering, visitation, path, and temporal categories operationalizes the framework. Empirically, we evaluate CHMMs against 6 alternative decoders on 3 real-world sequence-labeling tasks of substantively different character: gene-structure decoding in \emph{Drosophila melanogaster}, free-living activity recognition in CASAS smart-home environments, and protocol-defined human activity recognition from wearable sensors. The results reveal a clean local-versus-cumulative dichotomy in which controller augmentation is uniquely able to recover globally feasible trajectories on cumulative-constraint regimes, whilst simpler decoders are matched in validity on locally-dominated regimes. Together, theory and experiment characterize when exact controller augmentation is necessary and when simpler approaches suffice.

stat.ME

Improving the quasi-biennial oscillation via a surrogate-accelerated multi-objective optimization

Simulating the QBO remains a formidable challenge partly due to uncertainties in representing convectively generated gravity waves. We develop an end-to-end uncertainty quantification workflow that calibrates these gravity wave processes in E3SM to yield a more realistic QBO. Central to our approach is a domain knowledge-informed, compressed representation of high-dimensional spatio-temporal wind fields. By employing a parsimonious statistical model that learns the fundamental frequency of the underlying stochastic process from complex observations, we extract a concise set of interpretable and physically meaningful quantities of interest capturing key attributes, such as oscillation amplitude and period. Building on this, we train a probabilistic surrogate model. Leveraging the Karhunen-Loeve decomposition, our surrogate efficiently represents these characteristics as a set of orthogonal features, thereby capturing the cross-correlations among multiple physics quantities evaluated at different stratospheric pressure levels, and enabling rapid surrogate-based inference at a fraction of the computational cost of inference reliant only on full-scale simulations. Finally, we analyze the inverse problem using a multi-objective approach. Our study reveals a tension between amplitude and period that constrains the QBO representation, precluding a single optimal solution that simultaneously satisfies both objectives. To navigate this challenge, we quantify the bi-criteria trade-off and generate a representative set of Pareto optimal physics parameter values that balance the conflicting objectives. This integrated workflow not only improves the fidelity of QBO simulations but also advances toward a practical framework for tuning modes of variability and quasi-periodic phenomena, offering a versatile template for uncertainty quantification in complex geophysical models.

physics.ao-ph

The RITAS algorithm: a constructive yield monitor data processing algorithm

Yield monitor datasets are known to contain a high percentage of unreliable records. The current tool set is mostly limited to observation cleaning procedures based on heuristic or empirically-motivated statistical rules for extreme value identification and removal. We propose a constructive algorithm for handling well-documented yield monitor data artifacts without resorting to data deletion. The four-step Rectangle creation, Intersection assignment and Tessellation, Apportioning, and Smoothing (RITAS) algorithm models sample observations as overlapping, unequally-shaped, irregularly-sized, time-ordered, areal spatial units to better replicate the nature of the destructive sampling process. Positional data is used to create rectangular areal spatial units. Time-ordered intersecting area tessellation and harvested mass apportioning generate regularly-shaped and -sized polygons partitioning the entire harvested area. Finally, smoothing via a Gaussian process is used to provide map users with spatial-trend visualization. The intermediate steps as well as the algorithm output are illustrated in maize and soybean grain yield maps for five years of yield monitor data collected at a research agricultural site located in the US Fish and Wildlife Service Neal Smith National Wildlife Refuge.

stat.ME

Automatic Dynamic Relevance Determination for Gaussian process regression with high-dimensional functional inputs

In the context of Gaussian process regression with functional inputs, it is common to treat the input as a vector. The parameter space becomes prohibitively complex as the number of functional points increases, effectively becoming a hindrance for automatic relevance determination in high-dimensional problems. Generalizing a framework for time-varying inputs, we introduce the asymmetric Laplace functional weight (ALF): a flexible, parametric function that drives predictive relevance over the index space. Automatic dynamic relevance determination (ADRD) is achieved with three unknowns per input variable and enforces smoothness over the index space. Additionally, we discuss a screening technique to assess under complete absence of prior and model information whether ADRD is reasonably consistent with the data. Such tool may serve for exploratory analyses and model diagnostics. ADRD is applied to remote sensing data and predictions are generated in response to atmospheric functional inputs. Fully Bayesian estimation is carried out to identify relevant regions of the functional input space. Validation is performed to benchmark against traditional vector-input model specifications. We find that ADRD outperforms models with input dimension reduction via functional principal component analysis. Furthermore, the predictive power is comparable to high-dimensional models, in terms of both mean prediction and uncertainty, with 10 times fewer tuning parameters. Enforcing smoothness on the predictive relevance profile rules out erratic patterns associated with vector-input models.

stat.ME