SearcharxivSearch

arXiv subjects

Bruno Sansó

Publications and source records attributed to Bruno Sansó.

13 recordsLinked to original sources

Bayesian Quantile Deep Echo State Networks for Nonlinear Time Series

Conditional quantiles are central to asymmetric decision losses, tail-risk assessment, and interval forecasts, but Bayesian quantile regression for nonlinear time series can be difficult when the temporal feature vector is high-dimensional. We develop the quantile deep echo state network (Q-DESN), a Bayesian quantile regression model conditional on fixed features generated by a deep echo state network. Conditional on a specified reservoir construction and feature map, posterior uncertainty is assigned to the regression coefficients, likelihood, and shrinkage parameters. Single-level fits use asymmetric Laplace or quantile-fixed generalized asymmetric Laplace working likelihoods with ridge or regularized-horseshoe priors. Posterior computation uses Markov chain Monte Carlo when computationally practical and a model-specific variational Bayes approximation for analyses requiring repeated fitting. For quantile grids, we compare independent level-wise regressions followed by monotone rearrangement with a joint quantile-vector regression that shrinks adjacent quantile-specific coefficient differences. Synthetic studies, a single-origin retrospective Global Flood Awareness System (GloFAS) streamflow case, and a retrospective PriceFM comparison identify settings where this fixed-feature Bayesian regression improves finite-grid quantile scores.

stat.ME

Mean-Tilted Intervals: Short Tolerance Intervals

Intervals with the same probability content can have different endpoint placements and widths. This matters for tolerance inference, where a reported interval must also satisfy a repeated-sampling content-confidence statement. We develop mean-tilted intervals (MTIs), a fixed-content family indexed by retained-mean balance. The zero-tilt member is the mean-preserving interval (MPI) induced by the residual-product criterion of Pouplin et al.; nonzero tilts move through admissible contiguous windows, including distribution-specific central and shortest intervals. For tolerance inference, we introduce TCSP, a tolerance-calibrated shortest-path action. TCSP chooses the retained order-statistic count by distribution-free scan calibration and reports the shortest closed window at that count. This keeps the certified interval action separate from generalized-posterior endpoint summaries. We also study a calibrated MTI-ECM comparator that profiles fitted content and tilt over a prespecified grid and applies an independent Dirichlet-process content-probability check. In iid simulations at tolerance confidence 0.95, we compare TCSP, MTI-ECM, Young-Mathew interpolation, and Wilks intervals across feasible content-sample-size cells and eight continuous distributions. The study emphasizes skewed distributions, where placement matters most, and excludes cells where the sample range cannot support the requested two-sided distribution-free statement.

stat.ME

Bayesian Quantile-Based Correction and Synthesis of Hydrologic Products

River-flow forecasting requires predictive distributions that remain informative in both routine and extreme conditions. We develop a Bayesian quantile-based correction-and-synthesis framework built on Dynamic Quantile Linear Models (DQLMs). The framework links U.S. Geological Survey (USGS) observations, retrospective products, and ensemble forecast products through a shared latent quantile process, learns dynamic discrepancies for each external source, and combines quantile-specific posterior predictions into a single predictive distribution. We also adapt variational Bayes inference to the extended dynamic quantile linear model using Laplace--Delta approximations for non-conjugate parameters. The methodology is illustrated using daily flow for the San Lorenzo River together with products from the European Centre for Medium-Range Weather Forecasts (ECMWF) Global Flood Awareness System (GloFAS) and the National Oceanic and Atmospheric Administration (NOAA) National Weather Service (NWS), with emphasis on medium-range forecasting and uncertainty quantification across multiple quantile levels.

stat.AP

exdqlm: An R Package for Estimation and Analysis of Flexible Dynamic Quantile Linear Models

We present the R package exdqlm for Bayesian quantile regression, with primary emphasis on dynamic state-space quantile models for time series. The package is built around extended dynamic quantile linear models (exDQLMs), which use the extended asymmetric Laplace (exAL) family, a parametric extension of the asymmetric Laplace (AL) distribution commonly used in quantile regression. The software provides posterior simulation via Markov chain Monte Carlo (MCMC) and fast approximate posterior inference via Laplace-delta variational Bayes (LDVB), supporting posterior uncertainty quantification while also providing a computationally efficient option for longer time series. The same package interface supports static exAL quantile regression with regularized priors, dynamic transfer-function models for nonlinear input effects at a given quantile, post hoc posterior-predictive synthesis across separately fitted quantiles, forecasting, and quantitative and visual diagnostics for model evaluation.

stat.ME

Calibrated Bayesian Nonparametric Tolerance Intervals

Tolerance intervals provide bounds that contain a specified proportion of a population with a given confidence level, yet their construction remains challenging when parametric assumptions fail or sample sizes are small. Traditional nonparametric methods, such as Wilks' intervals, lack flexibility and often require large samples to be valid. We propose a fully nonparametric approach for constructing one-sided and two-sided tolerance intervals using a calibrated Gibbs posterior. Leveraging the connection between tolerance limits and population quantiles, we employ a Gibbs posterior based on the asymmetric Laplace (check) loss function. A key feature of our method is the calibration of the learning rate, which ensures nominal frequentist coverage across diverse distributional shapes. Simulation studies show that the proposed approach often yields shorter intervals than classical nonparametric benchmarks while maintaining reliable coverage. The framework's practical utility is illustrated through applications in ecology, biopharmaceutical manufacturing, and environmental monitoring, demonstrating its flexibility and robustness across diverse applications.

stat.ME

Tolerance Intervals Using Dirichlet Processes

In nonclinical pharmaceutical development, tolerance intervals are critical in ensuring product and process quality. They are statistical intervals designed to contain a specified proportion of the population with a given confidence level. Parametric and non-parametric methods have been developed to obtain tolerance intervals. The former work with small samples but can be affected by distribution misspecification. The latter offer larger flexibility but require large sample sizes. As an alternative, we propose Dirichlet process-based Bayesian nonparametric tolerance intervals to overcome the limitations. We develop a computationally efficient tolerance interval construction algorithm based on the analytically tractable quantile process of the Dirichlet process. Simulation studies show that our new approach is very robust to distributional assumptions and performs as efficiently as existing tolerance interval methods. To illustrate how the model works in practice, we apply our method to the tolerance interval estimation for potency data.

stat.ME

Mixture Modeling for Temporal Point Processes with Memory

We propose a constructive approach to building temporal point processes that incorporate dependence on their history. The dependence is modeled through the conditional density of the duration, i.e., the interval between successive event times, using a mixture of first-order conditional densities for each one of a specific number of lagged durations. Such a formulation for the conditional duration density accommodates high-order dynamics, and it thus enables flexible modeling for point processes with memory. The implied conditional intensity function admits a representation as a local mixture of first-order hazard functions. By specifying appropriate families of distributions for the first-order conditional densities, with different shapes for the associated hazard functions, we can obtain either self-exciting or self-regulating point processes. From the perspective of duration processes, we develop a method to specify a stationary marginal density. The resulting model, interpreted as a dependent renewal process, introduces high-order Markov dependence among identically distributed durations. Furthermore, we provide extensions to cluster point processes. These can describe duration clustering behaviors attributed to different factors, thus expanding the scope of the modeling framework to a wider range of applications. Regarding implementation, we develop a Bayesian approach to inference, model checking, and prediction. We investigate point process model properties analytically, and illustrate the methodology with both synthetic and real data examples.

stat.ME

A Heterogeneous Spatial Model for Soil Carbon Mapping of the Contiguous United States Using VNIR Spectra

The Rapid Carbon Assessment, conducted by the U.S. Department of Agriculture, was implemented in order to obtain a representative sample of soil organic carbon across the contiguous United States. In conjunction with a statistical model, the dataset allows for mapping of soil carbon prediction across the U.S., however there are two primary challenges to such an effort. First, there exists a large degree of heterogeneity in the data, whereby both the first and second moments of the data generating process seem to vary both spatially and for different land-use categories. Second, the majority of the sampled locations do not actually have lab measured values for soil organic carbon. Rather, visible and near-infrared (VNIR) spectra were measured at most locations, which act as a proxy to help predict carbon content. Thus, we develop a heterogeneous model to analyze this data that allows both the mean and the variance to vary as a function of space as well as land-use category, while incorporating VNIR spectra as covariates. After a cross-validation study that establishes the effectiveness of the model, we construct a complete map of soil organic carbon for the contiguous U.S. along with uncertainty quantification.

stat.AP

Nearest-Neighbor Mixture Models for Non-Gaussian Spatial Processes

We develop a class of nearest-neighbor mixture models that provide direct, computationally efficient, probabilistic modeling for non-Gaussian geospatial data. The class is defined over a directed acyclic graph, which implies conditional independence in representing a multivariate distribution through factorization into a product of univariate conditionals, and is extended to a full spatial process. We model each conditional as a mixture of spatially varying transition kernels, with locally adaptive weights, for each one of a given number of nearest neighbors. The modeling framework emphasizes the description of non-Gaussian dependence at the data level, in contrast with approaches that introduce a spatial process for transformed data, or for functionals of the data probability distribution. Thus, it facilitates efficient, full simulation-based inference. We study model construction and properties analytically through specification of bivariate distributions that define the local transition kernels, providing a general strategy for modeling general types of non-Gaussian data. Regarding computation, the framework lays out a new approach to handling spatial data sets, leveraging a mixture model structure to avoid computational issues that arise from large matrix operations. We illustrate the methodology using synthetic data examples and an analysis of Mediterranean Sea surface temperature observations.

stat.ME

Bayesian Geostatistical Modeling for Discrete-Valued Processes

We introduce a flexible and scalable class of Bayesian geostatistical models for discrete data, based on the class of nearest neighbor mixture transition distribution processes (NNMP), referred to as discrete NNMP. The proposed class characterizes spatial variability by a weighted combination of first-order conditional probability mass functions (pmfs) for each one of a given number of neighbors. The approach supports flexible modeling for multivariate dependence through specification of general bivariate discrete distributions that define the conditional pmfs. Moreover, the discrete NNMP allows for construction of models given a pre-specified family of marginal distributions that can vary in space, facilitating covariate inclusion. In particular, we develop a modeling and inferential framework for copula-based NNMPs that can attain flexible dependence structures, motivating the use of bivariate copula families for spatial processes. Compared to the traditional class of spatial generalized linear mixed models, where spatial dependence is introduced through a transformation of response means, our process-based modeling approach provides both computational and inferential advantages. We illustrate the benefits with synthetic data examples and an analysis of North American Breeding Bird Survey data.

stat.ME

On Construction and Estimation of Stationary Mixture Transition Distribution Models

Mixture transition distribution time series models build high-order dependence through a weighted combination of first-order transition densities for each one of a specified number of lags. We present a framework to construct stationary transition mixture distribution models that extend beyond linear, Gaussian dynamics. We study conditions for first-order strict stationarity which allow for different constructions with either continuous or discrete families for the first-order transition densities given a pre-specified family for the marginal density, and with general forms for the resulting conditional expectations. Inference and prediction are developed under the Bayesian framework with particular emphasis on flexible, structured priors for the mixture weights. Model properties are investigated both analytically and through synthetic data examples. Finally, Poisson and Lomax examples are illustrated through real data applications.

stat.ME

Modeling for seasonal marked point processes: An analysis of evolving hurricane occurrences

Seasonal point processes refer to stochastic models for random events which are only observed in a given season. We develop nonparametric Bayesian methodology to study the dynamic evolution of a seasonal marked point process intensity. We assume the point process is a nonhomogeneous Poisson process and propose a nonparametric mixture of beta densities to model dynamically evolving temporal Poisson process intensities. Dependence structure is built through a dependent Dirichlet process prior for the seasonally-varying mixing distributions. We extend the nonparametric model to incorporate time-varying marks, resulting in flexible inference for both the seasonal point process intensity and for the conditional mark distribution. The motivating application involves the analysis of hurricane landfalls with reported damages along the U.S. Gulf and Atlantic coasts from 1900 to 2010. We focus on studying the evolution of the intensity of the process of hurricane landfall occurrences, and the respective maximum wind speed and associated damages. Our results indicate an increase in the number of hurricane landfall occurrences and a decrease in the median maximum wind speed at the peak of the season. Introducing standardized damage as a mark, such that reported damages are comparable both in time and space, we find that there is no significant rising trend in hurricane damages over time.

stat.AP

The 2004 Venezuelan Presidential Recall Referendum: Discrepancies Between Two Exit Polls and Official Results

We present a simulation-based study in which the results of two major exit polls conducted during the recall referendum that took place in Venezuela on August 15, 2004, are compared to the official results of the Venezuelan National Electoral Council "Consejo Nacional Electoral" (CNE). The two exit polls considered here were conducted independently by Súmate, a nongovernmental organization, and Primero Justicia, a political party. We find significant discrepancies between the exit poll data and the official CNE results in about 60% of the voting centers that were sampled in these polls. We show that discrepancies between exit polls and official results are not due to a biased selection of the voting centers or to problems related to the size of the samples taken at each center. We found discrepancies in all the states where the polls were conducted. We do not have enough information on the exit poll data to determine whether the observed discrepancies are the consequence of systematic biases in the selection of the people interviewed by the pollsters around the country. Neither do we have information to study the possibility of a high number of false or nonrespondents. We have limited data suggesting that the discrepancies are not due to a drastic change in the voting patterns that occurred after the exit polls were conducted. We notice that the two exit polls were done independently and had few centers in common, yet their overall results were very similar.

stat.ME