SearcharxivSearch

arXiv subjects

Jorge Castillo-Mateo

Publications and source records attributed to Jorge Castillo-Mateo.

11 recordsLinked to original sources

Double zero-inflated spatio-temporal modeling of daily precipitation under detection thresholds

Explaining precipitation behavior at daily scale is important for fine scale understanding of the mechanisms driving precipitation. However, this effort is challenging because of the frequent incidence of zeros. The challenge is amplified by the acknowledged incidence of two types of zeros -- absence of precipitation as a dry event and absence of measured precipitation due to detection limits. In this work, we propose a multilevel spatio-temporal model which allows us to distinguish and explain the two types of zeros, as well as to model positive precipitation above the detection limit. The methodology combines a point mass at zero with probability modeled through a probit regression, a Gamma regression for latent positive precipitation amounts, and an observation mechanism subject to threshold-induced censoring. To capture spatial dependencies, Gaussian processes are employed in each regression model. Working within a Bayesian framework, we can obtain a rich range of inference with exact uncertainty. In particular, we provide model-based inference tools to compare and quantify differences between the true precipitation process and its observed counterpart across relevant characteristics. We apply our model to the analysis of daily spring observations at 70 sites over 15 years from the Ebro River Basin in northeastern Spain. Our findings indicate that the threshold strongly affects the occurrence of observed precipitation, especially in humid regions. While its impact on total accumulated amounts is small, it can exert a relevant effect on upper quantiles.

stat.ME

Toward a practical handbook for choosing among causal inference methods in non-randomized studies with binary outcomes: A simulation study for applied researchers

Applied researchers in biomedicine and related fields are often interested in estimating the causal effect of a treatment or intervention. Although randomized clinical trials are considered the gold standard for establishing causal effects, they are not always feasible, and real-world data may represent the only available source of evidence. In such settings, causal effects must be estimated using statistical methods applied to observational data. Over the last few decades, modern causal inference methods based on the potential outcomes framework have emerged as useful tools in this field. However, many such techniques exist, and their performance depends on factors such as sample size, the proportion of treated patients, the proportion of patients experiencing the outcome, the magnitude of the treatment effect, the target estimand, and potential violations of the fundamental assumptions of causal inference. Given the wide range of available methods, selecting an appropriate approach can be challenging for applied researchers. This study uses a large-scale simulation experiment to address this issue and provide researchers with a guide in the form of a handbook for a binary treatment and a binary outcome. Particularly, we test four popular statistical techniques: propensity score matching (full matching), inverse of the probability weighting, G-computation, and targeted maximum likelihood estimation. The proposed handbook is applied to two real-world datasets to assess its practical utility: one comprising vulnerable patients with mild COVID-19 (n=534 patients and more than 50% treated), and another of patients undergoing colorectal surgery (n=3635 patients and about 20% treated).

stat.ME

Evaluating the performance of GCM trajectories using Weather Type frequencies for persistence and transitions: the Iberian Peninsula and Lamb classification

This study evaluates the performance of 36 historical CMIP6 GCM trajectories (1979-2005) in reproducing atmospheric circulation over the Iberian Peninsula in the summer months (June-September) using the Lamb Weather Type (WT) classification scheme. Using ERA5 reanalysis as the observational reference, we introduce a methodological framework-applicable to any region worldwide-to evaluate GCM performance. This approach extends traditional daily frequency analysis by evaluating both the daily frequency distribution of WTs and their 24-hour dynamic evolution (i.e., transition probabilities and persistence). Model performance is quantified using the Overlap coefficient. A filtering process is applied where only trajectories that successfully reproduce both daily and conditional distributions with a minimum Overlap threshold $t_{sim}$ across a set number of grid points are retained. The findings show that while several models can adequately reproduce daily WT frequencies (16 out of 36), some struggle to capture day-to-day atmospheric transitions. This leads to a final selection of 12 trajectories over the Iberian Peninsula. Model performance across the region is then evaluated using integrated metrics assessing daily reproduction, conditional reproduction, and transition dynamics. Overall, models from the ec earth3 family-specifically the ec earth3 aerchem trajectory-exhibit the best and most consistent performance across the region. Additionally, the results highlight a geographical performance gap: while models generally represent circulation well in the northwest, they face significant challenges in the central and southern Mediterranean regions of the Peninsula. Ultimately, this study establishes that assessing WT persistence and transitions provides a far more discriminative, objective tool for GCM selection than evaluating daily distributions alone.

stat.ME

A two-stage approach to heat-mortality risk assessment comparing multiple exposure-to-temperature models: the case study in Lazio, Italy

This study investigates how different spatiotemporal temperature models affect the estimation of heat-related mortality in Lazio, Italy (2008--2022). First, we compare three methods to reconstruct daily maximum temperature at the municipality level: 1. a Bayesian quantile regression model with spatial interpolation, 2. a Bayesian Gaussian regression model, 3. the gridded reanalysis data from ERA5-Land. Both Bayesian models are station-based and exhibit higher and more spatially variable temperatures compared to ERA5-Land. Then, using individual mortality data for cardiovascular and respiratory causes, we estimate temperature-mortality associations through Bayesian conditional Poisson models in a case-crossover design. Exposure is defined as the mean maximum temperature over the previous three days. Additional models include heatwave definitions combining different thresholds and durations. All models exhibit a marked increase in relative risk at high temperatures; however, the temperature of minimum risk varies significantly across methods. Stratified analyses reveal higher relative risk increases in females and the elderly (80+). Heatwave effects depend on the definitions used, but all methods capture an increased mortality risk associated with prolonged heat exposure. Results confirm the importance of temperature model choice in epidemiology and provide insights for early warning systems and climate-health adaptation strategies.

stat.AP

Prediction of Maximum Temperature record in Spain 1960-2023 by the means of ERA5 atmospheric geopotentials

The increasing frequency of extreme temperature events, such as daily maximum temperature ($T_x$) records, underscores the need for robust tools to understand their drivers and predict their occurrence. Previous studies have identified increasing and non-stationary trends in $T_x$ records across the Iberian Peninsula, particularly during summer, the literature directly exploring their connection with upper-level atmospheric covariates remains limited. This work develops and applies an innovative methodological framework to model the occurrence of $T_x$ records and their relationship with geopotential height fields. We used daily $T_x$ data from 36 Spanish stations (1960-2023) provided by ECA&D and geopotential height data at 300, 500, and 700 hPa from ERA5. Exploratory analysis revealed a non-stationary trend in records, a higher frequency in the interior of the peninsula, and decreasing spatial co-occurrence with distance. We designed a hierarchical spatio-temporal logistic regression algorithm prioritizing interpretability and high-dimensionality reduction. The approach involves: (1) fitting local models per station; (2) applying a spatial consensus filter based on statistical significance to reduce the initial 1620 covariates to 17 in a base model (M1); and (3) a controlled incorporation of interaction terms. Among the tested models, a global model (M2) that enhances M1 with geodetic interactions was selected for its optimal balance between predictive performance (AUC) and complexity. Model M2 demonstrates high predictive accuracy at interior stations and good performance at coastal stations. It also adequately reproduces key observed properties, including the persistence of record streaks and patterns of spatial co-occurrence. This study provides a novel tool for predicting upcoming record events with high accuracy while maintaining a concise and interpretable structure.

stat.AP

Joint space-time modelling for upper daily maximum and minimum temperature record-breaking

Record-breaking temperature events are now frequently in the news, proffered as evidence of climate change, and often bring significant economic and human impacts. Our previous work undertook the first substantial spatial modelling investigation of temperature record-breaking across years for any given day within the year, employing a dataset consisting of over sixty years of daily maximum temperatures across peninsular Spain. That dataset also supplies daily minimum temperatures (which, in fact, are now available through 2023). Here, the dataset is converted into a daily pair of binary events, indicators, for that day, of whether a yearly record was broken for the daily maximum temperature and/or for the daily minimum temperature. Joint modelling addresses several inference issues: (i) defining/modelling record-breaking with bivariate time series of yearly indicators, (ii) strength of relationship between record-breaking events, (iii) prediction of joint, conditional and marginal record-breaking, (iv) persistence in record-breaking across days, (v) spatial interpolation across peninsular Spain. We substantially expand our previous work to enable investigation of these issues. We observe strong correlation between both processes but a growing trend of climate change that is well differentiated between them both spatially and temporally as well as different strengths of persistence and spatial dependence.

stat.ME

Spatio-temporal modeling for record-breaking temperature events in Spain

Record-breaking temperature events are now very frequently in the news, viewed as evidence of climate change. With this as motivation, we undertake the first substantial spatial modeling investigation of temperature record-breaking across years for any given day within the year. We work with a dataset consisting of over sixty years (1960-2021) of daily maximum temperatures across peninsular Spain. Formal statistical analysis of record-breaking events is an area that has received attention primarily within the probability community, dominated by results for the stationary record-breaking setting with some additional work addressing trends. Such effort is inadequate for analyzing actual record-breaking data. Effective analysis requires rich modeling of the indicator events which define record-breaking sequences. Resulting from novel and detailed exploratory data analysis, we propose hierarchical conditional models for the indicator events. After suitable model selection, we discover explicit trend behavior, necessary autoregression, significance of distance to the coast, useful interactions, helpful spatial random effects, and very strong daily random effects. Illustratively, the model estimates that global warming trends have increased the number of records expected in the past decade almost two-fold, 1.93 (1.89,1.98), but also estimates highly differentiated climate warming rates in space and by season.

stat.ME

Bayesian joint quantile autoregression

Quantile regression continues to increase in usage, providing a useful alternative to customary mean regression. Primary implementation takes the form of so-called multiple quantile regression, creating a separate regression for each quantile of interest. However, recently, advances have been made in joint quantile regression, supplying a quantile function which avoids crossing of the regression across quantiles. Here, we turn to quantile autoregression (QAR), offering a fully Bayesian version. We extend the initial quantile regression work of Koenker and Xiao (2006) in the spirit of Tokdar and Kadane (2012). We offer a directly interpretable parametric model specification for QAR. Further, we offer a p-th order QAR(p) version, a multivariate QAR(1) version, and a spatial QAR(1) version. We illustrate with simulation as well as a temperature dataset collected in Aragón, Spain.

stat.ME

Model-based tools for assessing space and time change in daily maximum temperature: an application to the Ebro basin in Spain

There is continuing interest in the investigation of change in temperature over space and time. We offer a set of tools to illuminate such change temporally, at desired temporal resolution, and spatially, according to region of interest, using data generated from suitable space-time models. These tools include predictive spatial probability surfaces and spatial extents for an event. Working with exceedance events around the center of the temperature distribution, the probability surfaces capture the spatial variation in the risk of an exceedance event, while the spatial extents capture the expected proportion of incidence of a given exceedance event for a region of interest. Importantly, the proposed tools can be used with the output from any suitable model fitted to any set of spatially referenced time series data. As an illustration, we employ a dataset from 1956 to 2015 collected at 18 stations over Aragón in Spain, and a collection of daily maximum temperature series obtained from posterior predictive simulation of a Bayesian hierarchical daily temperature model. The results for the summer period show that although there is an increasing risk in all the events used to quantify the effects of climate change, it is not spatially homogeneous, with the largest increase arising in the center of Ebro valley and Eastern Pyrenees area. The risk of an increase of the average temperature between 1966-1975 and 2006-2015 higher than $1^\circ$C is higher than 0.5 all over the region, and close to 1 in the previous areas. The extent of daily temperature higher than the reference mean has increased 3.5% per decade. The mean of the extent indicates that 95% of the area under study has suffered a positive increment of the average temperature, and almost 70% higher than $1^{\circ}$C.

stat.ME

Distribution-free changepoint detection tests based on the breaking of records

The analysis of record-breaking events is of interest in fields such as climatology, hydrology or anthropology. In connection with the record occurrence, we propose three distribution-free statistics for the changepoint detection problem. They are CUSUM-type statistics based on the upper and/or lower record indicators observed in a series. Using a version of the functional central limit theorem, we show that the CUSUM-type statistics are asymptotically Kolmogorov distributed. The main results under the null hypothesis are based on series of independent and identically distributed random variables, but a statistic to deal with series with seasonal component and serial correlation is also proposed. A Monte Carlo study of size, power and changepoint estimate has been performed. Finally, the methods are illustrated by analyzing the time series of temperatures at Madrid, Spain. The R package $\texttt{RecordTest}$ publicly available on CRAN implements the proposed methods.

stat.ME

Spatial modeling of day-within-year temperature time series: an examination of daily maximum temperatures in Aragón, Spain

Acknowledging a considerable literature on modeling daily temperature data, we propose a multi-level spatio-temporal model which introduces several innovations in order to explain the daily maximum temperature in the summer period over 60 years in a region containing Aragón, Spain. The model operates over continuous space but adopts two discrete temporal scales, year and day within year. It captures temporal dependence through autoregression on days within year and also on years. Spatial dependence is captured through spatial process modeling of intercepts, slope coefficients, variances, and autocorrelations. The model is expressed in a form which separates fixed effects from random effects and also separates space, years, and days for each type of effect. Motivated by exploratory data analysis, fixed effects to capture the influence of elevation, seasonality and a linear trend are employed. Pure errors are introduced for years, for locations within years, and for locations at days within years. The performance of the model is checked using a leave-one-out cross-validation. Applications of the model are presented including prediction of the daily temperature series at unobserved or partially observed sites and inference to investigate climate change comparison.

stat.ME