SearcharxivSearch

arXiv subjects

Nancy L. Garcia

Publications and source records attributed to Nancy L. Garcia.

At least 19 recordsLinked to original sources

A Bayesian Spatial-Temporal Functional Model for Data with Block Structure and Repeated Measures

The analysis of spatio-temporal data has been the object of research in several areas of knowledge. One of the main objectives of such research is the need to evaluate the behavior of climate effects in certain regions across a period of time. When certain climate patterns appear for several days or even weeks, causing the areas affected by them to have the same kind of weather for an extended period of time, the use of blocks for these phenomena may be a good strategy. Additionally, having repeated measures for observations within blocks helps to control for differences between observations, thus gaining more statistical power. In view of these perspectives, this study presents a spatio-temporal regression model with block structure with repeated measures incorporating as predictors functional variables of fixed and random nature. To accommodate complex spatial, temporal and block structures, functional components based on random effects were considered in addition to the class Matérn covariance structure, which was responsible to account for spatial covariance. This work is motivated by a precipitation dataset collected monthly from various meteorological stations in Goiás State, Brazil, covering the years 1980 to 2001 (21 years). In this framework, spatial effects are represented by individual meteorological stations, temporal effects by months, block effects by climate patterns, and repeated measures by the years within those patterns. The proposed model demonstrated promising results in simulation studies and effectively estimated precipitation using the available data.

stat.ME

Predicting Dengue Outbreaks: A Dynamic Approach with Variable Length Markov Chains and Exogenous Factors

Variable Length Markov Chains with Exogenous Covariates (VLMCX) are stochastic models that use Generalized Linear Models to compute transition probabilities, taking into account both the state history and time-dependent exogenous covariates. The beta-context algorithm selects a relevant finite suffix (context) for predicting the next symbol. This algorithm estimates flexible tree-structured models by aggregating irrelevant states in the process history and enables the model to incorporate exogenous covariates over time. This research uses data from multiple sources to extend the beta-context algorithm to incorporate time-dependent and time-invariant exogenous covariates. Within this approach, we have a distinct Markov chain for every data source, allowing for a comprehensive understanding of the process behavior across multiple situations, such as different geographic locations. Despite using data from different sources, we assume that all sources are independent and share identical parameters - we explore contexts within each data source and combine them to compute transition probabilities, deriving a unified tree. This approach eliminates the necessity for spatial-dependent structural considerations within the model. Furthermore, we incorporate modifications in the estimation procedure to address contexts that appear with low frequency. Our motivation was to investigate the impact of previous dengue rates, weather conditions, and socioeconomic factors on subsequent dengue rates across various municipalities in Brazil, providing insights into dengue transmission dynamics.

stat.ME

Unsupervised Bayesian classification for models with scalar and functional covariates

We consider unsupervised classification by means of a latent multinomial variable which categorizes a scalar response into one of L components of a mixture model. This process can be thought as a hierarchical model with first level modelling a scalar response according to a mixture of parametric distributions, the second level models the mixture probabilities by means of a generalised linear model with functional and scalar covariates. The traditional approach of treating functional covariates as vectors not only suffers from the curse of dimensionality since functional covariates can be measured at very small intervals leading to a highly parametrised model but also does not take into account the nature of the data. We use basis expansion to reduce the dimensionality and a Bayesian approach to estimate the parameters while providing predictions of the latent classification vector. By means of a simulation study we investigate the behaviour of our approach considering normal mixture model and zero inflated mixture of Poisson distributions. We also compare the performance of the classical Gibbs sampling approach with Variational Bayes Inference.

stat.ME

Graphical Construction of Spatial Gibbs Random Graphs

We consider a Random Graph Model on $\mathbb{Z}^{d}$ that incorporates the interplay between the statistics of the graph and the underlying space where the vertices are located. Based on a graphical construction of the model as the invariant measure of a birth and death process, we prove the existence and uniqueness of a measure defined on graphs with vertices in $\mathbb{Z}^{d}$ which coincides with the limit along the measures over graphs with finite vertex set. As a consequence, theoretical properties such as exponential mixing of the infinite volume measure and central limit theorem for averages of a real-valued function of the graph are obtained. Moreover, a perfect simulation algorithm based on the clan of ancestors is described in order to sample a finite window of the equilibrium measure defined on $\mathbb{Z}^{d}$.

math.ST

Aggregated functional data model applied on clustering and disaggregation of UK electrical load profiles

Understanding electrical energy demand at the consumer level plays an important role in planning the distribution of electrical networks and offering of off-peak tariffs, but observing individual consumption patterns is still expensive. On the other hand, aggregated load curves are normally available at the substation level. The proposed methodology separates substation aggregated loads into estimated mean consumption curves, called typical curves, including information given by explanatory variables. In addition, a model-based clustering approach for substations is proposed based on the similarity of their consumers typical curves and covariance structures. The methodology is applied to a real substation load monitoring dataset from the United Kingdom and tested in eight simulated scenarios.

stat.AP

Discrete one-dimensional coverage process on a renewal process

We consider the {following} coverage model on $\mathbb{N}$. For each site $i\in \mathbb{N}$ we associate a pair $(ξ_i, R_i)$ where $\{ξ_0, ξ_1, \ldots \}$ is a 1-dimensional {undelayed} discrete renewal point process and $\{R_0,R_1,\ldots\}$ is an i.i.d. sequence of $\mathbb{N}$-valued random variables. At each site where $ξ_i=1$ we start an interval of length $R_i$. Coverage occurs if every site of $\mathbb{N}$ is covered by some interval. We obtain sharp conditions for both, positive and null probability of coverage. As corollaries, we extend results of the literature of rumor processes and discrete one-dimensional Boolean percolation.

math.PR

Continuity properties of a factor of Markov chains

Starting from a Markov chain with a finite alphabet, we consider the chain obtained when all but one symbol are undistinguishable for the practitioner. We study necessary and sufficient conditions for this chain to have continuous transition probabilities with respect to the past.

math.PR

Rumor processes on $\N$ and discrete renewal processes

We study two rumor processes on $\N$, the dynamics of which are related to an SI epidemic model with long range transmission. Both models start with one spreader at site $0$ and ignorants at all the other sites of $\N$, but differ by the transmission mechanism. In one model, the spreaders transmit the information within a random distance on their right, and in the other the ignorants take the information from a spreader within a random distance on their left. We obtain the probability of survival, information on the distribution of the range of the rumor and limit theorems for the proportion of spreaders. The key step of our proofs is to show that, in each model, the position of the spreaders on $\N$ can be related to a suitably chosen discrete renewal process.

math.PR

Aggregated functional data model for Near-Infrared Spectroscopy calibration and prediction

Calibration and prediction for NIR spectroscopy data are performed based on a functional interpretation of the Beer-Lambert formula. Considering that, for each chemical sample, the resulting spectrum is a continuous curve obtained as the summation of overlapped absorption spectra from each analyte plus a Gaussian error, we assume that each individual spectrum can be expanded as a linear combination of B-splines basis. Calibration is then performed using two procedures for estimating the individual analytes curves: basis smoothing and smoothing splines. Prediction is done by minimizing the square error of prediction. To assess the variance of the predicted values, we use a leave-one-out jackknife technique. Departures from the standard error models are discussed through a simulation study, in particular, how correlated errors impact on the calibration step and consequently on the analytes' concentration prediction. Finally, the performance of our methodology is demonstrated through the analysis of two publicly available datasets.

stat.ME

Stochastically Perturbed Chains of Variable Memory

In this paper, we study inference for chains of variable order under two distinct contamination regimes. Consider we have a chain of variable memory on a finite alphabet containing zero. At each instant of time an independent coin is flipped and if it turns head a contamination occurs. In the first regime a zero is read independent of the value of the chain. In the second regime, the value of another chain of variable memory is observed instead of the original one. Our results state that the difference between the transition probabilities of the original process and the corresponding ones of the contaminated process may be bounded above uniformly. Moreover, if the contamination probability is small enough, using a version of the Context algorithm we are able to recover the context tree of the original process through a contaminated sample.

math.PR

Perfect simulation for locally continuous chains of infinite order

We establish sufficient conditions for perfect simulation of chains of infinite order on a countable alphabet. The new assumption, localized continuity, is formalized with the help of the notion of context trees, and includes the traditional continuous case, probabilistic context trees and discontinuous kernels. Since our assumptions are more refined than uniform continuity, our algorithms perfectly simulate continuous chains faster than the existing algorithms of the literature. We provide several illustrative examples.

math.PR

Context tree selection and linguistic rhythm retrieval from written texts

The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet they are conjectured, on linguistic grounds, to implement different rhythms. We show that this linguistic question can be formulated as a problem of model selection in the class of variable length Markov chains. To carry on this approach, we compare texts from European and Brazilian Portuguese. These texts are previously encoded according to some basic rhythmic features of the sentences which can be automatically retrieved. This is an entirely new approach from the linguistic point of view. Our statistical contribution is the introduction of the smallest maximizer criterion which is a constant free procedure for model selection. As a by-product, this provides a solution for the problem of optimal choice of the penalty constant when using the BIC to select a variable length Markov chain. Besides proving the consistency of the smallest maximizer criterion when the sample size diverges, we also make a simulation study comparing our approach with both the standard BIC selection and the Peres-Shields order estimation. Applied to the linguistic sample constituted for our case study, the smallest maximizer criterion assigns different context-tree models to the two dialects of Portuguese. The features of the selected models are compatible with current conjectures discussed in the linguistic literature.

stat.ML

A Hierarchical Model for Aggregated Functional Data

In many areas of science one aims to estimate latent sub-population mean curves based only on observations of aggregated population curves. By aggregated curves we mean linear combination of functional data that cannot be observed individually. We assume that several aggregated curves with linear independent coefficients are available. More specifically, we assume each aggregated curve is an independent partial realization of a Gaussian process with mean modeled through a weighted linear combination of the disaggregated curves. We model the mean of the Gaussian processes as a smooth function approximated by a function belonging to a finite dimensional space ${\cal H}_K$ which is spanned by $K$ B-splines basis functions. We explore two different specifications of the covariance function of the Gaussian process: one that assumes a constant variance across the domain of the process, and a more general variance structure which is itself modelled as a smooth function, providing a nonstationary covariance function. Inference procedure is performed following the Bayesian paradigm allowing experts' opinion to be considered when estimating the disaggregated curves. Moreover, it naturally provides the uncertainty associated with the parameters estimates and fitted values. Our model is suitable for a wide range of applications. We concentrate on two different real examples: calibration problem for NIR spectroscopy data and an analysis of distribution of energy among different type of consumers.

stat.ME

Perfect simulation of a coupling achieving the $\bar{d}$-distance between ordered pairs of binary chains of infinite order

We explicitly construct a coupling attaining Ornstein's $\bar{d}$-distance between ordered pairs of binary chains of infinite order. Our main tool is a representation of the transition probabilities of the coupled bivariate chain of infinite order as a countable mixture of Markov transition probabilities of increasing order. Under suitable conditions on the loss of memory of the chains, this representation implies that the coupled chain can be represented as a concatenation of iid sequence of bivariate finite random strings of symbols. The perfect simulation algorithm is based on the fact that we can identify the first regeneration point to the left of the origin almost surely.

math.PR

Perfect simulation for stochastic chains of infinite memory: relaxing the continuity assumption

This paper is composed of two main results concerning chains of infinite order which are not necessarily continuous. The first one is a decomposition of the transition probability kernel as a countable mixture of unbounded probabilistic context trees. This decomposition is used to design a simulation algorithm which works as a combination of the algorithms given by Comets et al. (2002) and Gallo (2009). The second main result gives sufficient conditions on the kernel for this algorithm to stop after an almost surely finite number of steps. Direct consequences of this last result are existence and uniqueness of the stationary chain compatible with the kernel.

math.PR

Spatial birth and death processes as solutions of stochastic equations

Spatial birth and death processes are obtained as solutions of a system of stochastic equations. The processes are required to be locally finite, but may involve an infinite population over the full (noncompact) type space. Conditions are given for existence and uniqueness of such solutions, and for temporal and spatial ergodicity. For birth and death processes with constant death rate, a sub-criticality condition on the birth rate implies that the process is ergodic and converges exponentially fast to the stationary distribution.

math.PR

On mixing times for stratified walks on the d-cube

Using the electric and coupling approaches, we derive a series of results concerning the mixing times for the stratified random walk on the d-cube, inspired in the results of Chung and Graham (1997) Stratified random walks on the n-cube.

physics.data-an