SearcharxivSearch

arXiv subjects

Yuichi Goto

Publications and source records attributed to Yuichi Goto.

11 recordsLinked to original sources

Parametric generalized spectrum for heavy-tailed time series

Recently, several spectra have emerged, designed to encapsulate the distributional characteristics of non-Gaussian stationary processes. This article introduces parametric families of generalized spectra based on the characteristic function, alongside inference procedures enabling $\sqrt{n}$-consistent estimation of the unknown parameters in a broad class of parametric models. These spectra capture non-linear dependencies without requiring that the underlying stochastic processes satisfy any moment assumptions. Crucially, this approach facilitates frequency domain analysis for heavy-tailed time series, including possibly non-causal Cauchy autoregressive models and discrete-stable integer-valued autoregressive models. To the best of our knowledge, the latter models have not been studied theoretically in the literature. By estimating parameters across both causal and non-causal parameter spaces, our method automatically identifies the causal or non-causal structure of Cauchy autoregressive models. Furthermore, our estimator does not depend on smoothing parameters since it is based on the integrated periodogram associated with the generalized spectrum. As applications, we develop goodness-of-fit tests, moving average unit-root tests, and tests for non-invertibility. We study the finite-sample performance of the proposed estimators and tests via Monte Carlo simulations, and apply the methodology to estimation and forecasting of a measles count dataset. We evaluate finite-sample performance using Monte Carlo simulations and illustrate the practical value of the procedure with an application to measles case-count estimation and forecasting.

math.ST

Mixed difference integer-valued GARCH model for $ \mathbb{Z}$-valued time series

In this paper, we introduce flexible observation-driven $\mathbb{Z}$-valued time series models constructed from mixtures of negative and non-negative components. Compared to models based on the standard Skellam distribution or on a difference of two integer-valued variables, our specification offers greater versatility. For example, it easily allows for skewness and bimodality. Furthermore, the observation of one component of the mixture makes interpretation and statistical analysis easier. We establish conditions for stationarity and mixing, and develop a mixed Poisson quasi-maximum likelihood estimator with proven asymptotic properties. A portmanteau test is proposed to diagnose residual serial dependence. The finite-sample performance of the methodology is assessed via simulation, and an empirical application on tick prices demonstrates its practical usefulness.

math.ST

On spectral clustering under non-isotropic Gaussian mixture models

We evaluate the misclustering probability of a spectral clustering algorithm under a Gaussian mixture model with a general covariance structure. The algorithm partitions the data into two groups based on the sign of the first principal component score. As a corollary of the main result, the clustering procedure is shown to be consistent in a high-dimensional regime.

math.ST

Robust $M$-Estimation of Scatter Matrices via Precision Structure Shrinkage

Maronna's and Tyler's $M$-estimators are among the most widely used robust estimators for scatter matrices. However, when the dimension of observations is relatively high, their performance can substantially deteriorate in certain situations, particularly in the presence of clustered outliers. To address this issue, we propose an estimator that shrinks the estimated precision matrix toward the identity matrix. We derive a sufficient condition for its existence, discuss its statistical interpretation, and establish upper and lower bounds for its additive finite sample breakdown point. Numerical experiments confirm the robustness of the proposed method.

stat.ME

ANOVATS: A subsampling-based test to detect differences among short time series in marine studies

Assessing marine ecosystems is important for understanding the impacts of climate change and human activity, as well as for maintaining healthy oceans and ecosystems. In marine science, it is common for biologists and geologists to identify regional differences based on expert knowledge, frequently through data visualization. However, time series data collected through surveys in marine studies typically span only a few decades, limiting the applicability of classical time series methods. Additionally, without expert knowledge, detecting significant differences becomes challenging. To address these issues, we introduce ANOVATS (ANOVA for small-sample time series data), a subsampling-based method to detect regional differences in small-sample time series data with a fixed number of groups. This method bypasses the need for spectral density estimation, which requires a large number of time points in the data. Furthermore, after detecting differences in homogeneity across all areas using the ANOVATS procedure, we devised a simple ANOVATS post hoc procedure to group the areas. Finally, we demonstrate the effectiveness of our method by analyzing zooplankton biomass data collected in different strata of the North Sea, showing its ability to quantify differences in species between geographical areas without relying on prior biological or geographical knowledge.

stat.ME

A model-free test of the time-reversibility of climate change processes

Time-reversibility is a crucial feature of many time series models, while time-irreversibility is the rule rather than the exception in real-life data. Testing the null hypothesis of time-reversibilty, therefore, should be an important step preliminary to the identification and estimation of most traditional time-series models. Existing procedures, however, mostly consist of testing necessary but not sufficient conditions, leading to under-rejection, or sufficient but non-necessary ones, which leads to over-rejection. Moreover, they generally are model-besed. In contrast, the copula spectrum studied by Goto et al. ($\textit{Ann. Statist.}$ 2022, $\textbf{50}$: 3563--3591) allows for a model-free necessary and sufficient time-reversibility condition. A test based on this copula-spectrum-based characterization has been proposed by authors. This paper illustrates the performance of this test, with an illustration in the analysis of climatic data.

stat.ME

Residual spectrum: Brain functional connectivity detection beyond coherence

Coherence is a widely used measure to assess linear relationships between time series. However, it fails to capture nonlinear dependencies. To overcome this limitation, this paper introduces the notion of residual spectral density as a higher-order extension of the squared coherence. The method is based on an orthogonal decomposition of time series regression models. We propose a test for the existence of the residual spectrum and derive its fundamental properties. A numerical study illustrates finite sample performance of the proposed method. An application of the method shows that the residual spectrum can effectively detect brain connectivity. Our study reveals a noteworthy contrast in connectivity patterns between schizophrenia patients and healthy individuals. Specifically, we observed that non-linear connectivity in schizophrenia patients surpasses that of healthy individuals, which stands in stark contrast to the established understanding that linear connectivity tends to be higher in healthy individuals. This finding sheds new light on the intricate dynamics of brain connectivity in schizophrenia.

math.ST

Spectral clustering algorithm for the allometric extension model

The spectral clustering algorithm is often used as a binary clustering method for unclassified data by applying the principal component analysis. To study theoretical properties of the algorithm, the assumption of conditional homoscedasticity is often supposed in existing studies. However, this assumption is restrictive and often unrealistic in practice. Therefore, in this paper, we consider the allometric extension model, that is, the directions of the first eigenvectors of two covariance matrices and the direction of the difference of two mean vectors coincide, and we provide a non-asymptotic bound of the error probability of the spectral clustering algorithm for the allometric extension model. As a byproduct of the result, we obtain the consistency of the clustering method in high-dimensional settings.

math.ST

A test for counting sequences of integer-valued autoregressive models

The integer autoregressive (INAR) model is one of the most commonly used models in nonnegative integer-valued time series analysis and is a counterpart to the traditional autoregressive model for continuous-valued time series. To guarantee the integer-valued nature, the binomial thinning operator or more generally the generalized Steutel and van Harn operator is used to define the INAR model. However, the distributions of the counting sequences used in the operators have been determined by the preference of analyst without statistical verification so far. In this paper, we propose a test based on the mean and variance relationships for distributions of counting sequences and a disturbance process to check if the operator is reasonable. We show that our proposed test has asymptotically correct size and is consistent. Numerical simulation is carried out to evaluate the finite sample performance of our test. As a real data application, we apply our test to the monthly number of anorexia cases in animals submitted to animal health laboratories in New Zealand and we conclude that binomial thinning operator is not appropriate.

math.ST

The integrated copula spectrum

Frequency domain methods form a ubiquitous part of the statistical toolbox for time series analysis. In recent years, considerable interest has been given to the development of new spectral methodology and tools capturing dynamics in the entire joint distributions and thus avoiding the limitations of classical, $L^2$-based spectral methods. Most of the spectral concepts proposed in that literature suffer from one major drawback, though: their estimation requires the choice of a smoothing parameter, which has a considerable impact on estimation quality and poses challenges for statistical inference. In this paper, associated with the concept of copula-based spectrum, we introduce the notion of copula spectral distribution function or integrated copula spectrum. This integrated copula spectrum retains the advantages of copula-based spectra but can be estimated without the need for smoothing parameters. We provide such estimators, along with a thorough theoretical analysis, based on a functional central limit theorem, of their asymptotic properties. We leverage these results to test various hypotheses that cannot be addressed by classical spectral methods, such as the lack of time-reversibility or asymmetry in tail dynamics.

math.ST

Sparse principal component analysis for high-dimensional stationary time series

We consider the sparse principal component analysis for high-dimensional stationary processes. The standard principal component analysis performs poorly when the dimension of the process is large. We establish the oracle inequalities for penalized principal component estimators for the processes including heavy-tailed time series. The rate of convergence of the estimators is established. We also elucidate the theoretical rate for choosing the tuning parameter in penalized estimators. The performance of the sparse principal component analysis is demonstrated by numerical simulations. The utility of the sparse principal component analysis for time series data is exemplified by the application to average temperature data.

math.ST