SearcharxivSearch

arXiv subjects

Ivor Cribben

Publications and source records attributed to Ivor Cribben.

10 recordsLinked to original sources

Tensor time series change-point detection in cryptocurrency network data

Financial fraud has been growing exponentially in recent years. The rise of cryptocurrencies as an investment asset has simultaneously seen a parallel growth in cryptocurrency scams. To detect possible cryptocurrency fraud, and in particular market manipulation, previous research focused on the detection of changes in the network of trades; however, market manipulators are now trading across multiple cryptocurrency platforms, making their detection more difficult. Hence, it is important to consider the identification of changes across several trading networks or a `network of networks' over time. To this end, in this article, we propose a new change-point detection method in the network structure of tensor-variate data. This new method, labeled TenSeg, first employs a tensor decomposition, and second detects multiple change-points in the second-order (cross-covariance or network) structure of the decomposed data. It allows for change-point detection in the presence of frequent changes of possibly small magnitudes and is computationally fast. We apply our method to several simulated datasets and to a cryptocurrency dataset, which consists of network tensor-variate data from the Ethereum blockchain. We demonstrate that our approach substantially outperforms other state-of-the-art change-point techniques, and the detected change-points in the Ethereum data set coincide with changes across several trading networks or a `network of networks' over time. Finally, all the relevant \textsf{R} code implementing the method in the article are available on https://github.com/Anastasiou-Andreas/TenSeg.

stat.ME

Bayesian Time-Varying Tensor Vector Autoregressive Models for Dynamic Effective Connectivity

In contemporary neuroscience, a key area of interest is dynamic effective connectivity, which is crucial for understanding the dynamic interactions and causal relationships between different brain regions. Dynamic effective connectivity can provide insights into how brain network interactions are altered in neurological disorders such as dyslexia. Time-varying vector autoregressive (TV-VAR) models have been employed to draw inferences for this purpose. However, their significant computational requirements pose challenges, since the number of parameters to be estimated increases quadratically with the number of time series. In this paper, we propose a computationally efficient Bayesian time-varying VAR approach. For dealing with large-dimensional time series, the proposed framework employs a tensor decomposition for the VAR coefficient matrices at different lags. Dynamically varying connectivity patterns are captured by assuming that at any given time only a subset of components in the tensor decomposition is active. Latent binary time series select the active components at each time via an innovative and parsimonious Ising model in the time-domain. Furthermore, we propose parsity-inducing priors to achieve global-local shrinkage of the VAR coefficients, determine automatically the rank of the tensor decomposition and guide the selection of the lags of the auto-regression. We show the performances of our model formulation via simulation studies and data from a real fMRI study involving a book reading experiment.

stat.ME

Do NHL goalies get hot in the playoffs? A multilevel logistic regression analysis

The hot-hand theory posits that an athlete who has performed well in the recent past performs better in the present. We use multilevel logistic regression to test this theory for National Hockey League playoff goaltenders, controlling for a variety of shot-related and game-related characteristics. Our data consists of 48,431 shots for 93 goaltenders in the 2008-2016 playoffs. Using a wide range of shot-based windows to quantify recent save performance, we find no evidence for the hot-hand theory, and some evidence that good recent save performance negatively impacts the next-shot save probability. We use a permutation test to rule out a regression to the mean explanation for our findings.

stat.AP

Factorized Binary Search: change point detection in the network structure of multivariate high-dimensional time series

Functional magnetic resonance imaging (fMRI) time series data presents a unique opportunity to understand the behavior of temporal brain connectivity, and models that uncover the complex dynamic workings of this organ are of keen interest in neuroscience. We are motivated to develop accurate change point detection and network estimation techniques for high-dimensional whole-brain fMRI data. To this end, we introduce factorized binary search (FaBiSearch), a novel change point detection method in the network structure of multivariate high-dimensional time series in order to understand the large-scale characterizations and dynamics of the brain. FaBiSearch employs non-negative matrix factorization, an unsupervised dimension reduction technique, and a new binary search algorithm to identify multiple change points. In addition, we propose a new method for network estimation for data between change points. We seek to understand the dynamic mechanism of the brain, particularly for two fMRI data sets. The first is a resting-state fMRI experiment, where subjects are scanned over three visits. The second is a task-based fMRI experiment, where subjects read Chapter 9 of Harry Potter and the Sorcerer's Stone. For the resting-state data set, we examine the test-retest behavior of dynamic functional connectivity, while for the task-based data set, we explore network dynamics during the reading and whether change points across subjects coincide with key plot twists in the story. Further, we identify hub nodes in the brain network and examine their dynamic behavior. Finally, we make all the methods discussed available in the R package fabisearch on CRAN.

stat.ME

The state of play of reproducibility in Statistics: an empirical analysis

Reproducibility, the ability to reproduce the results of published papers or studies using their computer code and data, is a cornerstone of reliable scientific methodology. Studies where results cannot be reproduced by the scientific community should be treated with caution. Over the past decade, the importance of reproducible research has been frequently stressed in a wide range of scientific journals such as \textit{Nature} and \textit{Science} and international magazines such as \textit{The Economist}. However, multiple studies have demonstrated that scientific results are often not reproducible across research areas such as psychology and medicine. Statistics, the science concerned with developing and studying methods for collecting, analyzing, interpreting and presenting empirical data, prides itself on its openness when it comes to sharing both computer code and data. In this paper, we examine reproducibility in the field of statistics by attempting to reproduce the results in 93 published papers in prominent journals utilizing functional magnetic resonance imaging (fMRI) data during the 2010-2021 period. Overall, from both the computer code and the data perspective, among all the 93 examined papers, we could only reproduce the results in 14 (15.1%) papers, that is, the papers provide both executable computer code (or software) with the real fMRI data, and our results matched the results in the paper. Finally, we conclude with some author-specific and journal-specific recommendations to improve the research reproducibility in statistics.

stat.AP

fabisearch: A Package for Change Point Detection in and Visualization of the Network Structure of Multivariate High-Dimensional Time Series in R

Change point detection is a commonly used technique in time series analysis, capturing the dynamic nature in which many real-world processes function. With the ever increasing troves of multivariate high-dimensional time series data, especially in neuroimaging and finance, there is a clear need for scalable and data-driven change point detection methods. Currently, change point detection methods for multivariate high-dimensional data are scarce, with even less available in high-level, easily accessible software packages. To this end, we introduce the R package fabisearch, available on the Comprehensive R Archive Network (CRAN), which implements the factorized binary search (FaBiSearch) methodology. FaBiSearch is a novel statistical method for detecting change points in the network structure of multivariate high-dimensional time series which employs non-negative matrix factorization (NMF), an unsupervised dimension reduction and clustering technique. Given the high computational cost of NMF, we implement the method in C++ code and use parallelization to reduce computation time. Further, we also utilize a new binary search algorithm to efficiently identify multiple change points and provide a new method for network estimation for data between change points. We show the functionality of the package and the practicality of the method by applying it to a neuroimaging and a finance data set. Lastly, we provide an interactive, 3-dimensional, brain-specific network visualization capability in a flexible, stand-alone function. This function can be conveniently used with any node coordinate atlas, and nodes can be color coded according to community membership (if applicable). The output is an elegantly displayed network laid over a cortical surface, which can be rotated in the 3-dimensional space.

stat.CO

Extremal Dependence in Australian Electricity Markets

Electricity markets are significantly more volatile than other comparable financial or commodity markets. Extreme price outcomes and their transmission between regions pose significant risks for market participants. We examine the dependence between extreme spot price outcomes in the Australian National Electricity Market (NEM). We investigate extremal dependence both in a univariate and multivariate setting, applying the extremogram developed by Davis and Mikosch (2009) and Davis et al. (2011, 2012). We measure the persistence of extreme prices within individual regional markets and the transmission of extreme prices across different regions. With both 5-minute and 30-minute price data, we find that extreme prices are more persistent in the market with a higher share of intermittent renewable energy. We also find that persistence of extreme prices is more prevalent in more concentrated markets. We also show significant extremal price dependence between different regions, which is typically stronger between physically interconnected markets. The dependence structure of extreme prices shows asymmetric and time-dependent patterns. Applying the extremograms, we further show the effectiveness of the Australian Energy Market Commission's 2016 rebidding rule with respect to reducing the share of isolated price spikes that are often considered as an indication of strategic bidding. Our results provide important information for hedging decisions of market participants and for policy makers who aim to reduce market volatility and extreme price outcomes through effective regulations which guide the trading behaviour of market participants as well as improved network interconnections.

q-fin.RM

Generalized reliability based on distances

The intraclass correlation coefficient (ICC) is a classical index of measurement reliability. With the advent of new and complex types of data for which the ICC is not defined, there is a need for new ways to assess reliability. To meet this need, we propose a new distance-based intraclass correlation coefficient (dbICC), defined in terms of arbitrary distances among observations. We introduce a bias correction to improve the coverage of bootstrap confidence intervals for the dbICC, and demonstrate its efficacy via simulation. We illustrate the proposed method by analyzing the test-retest reliability of brain connectivity matrices derived from a set of repeated functional magnetic resonance imaging scans. The Spearman-Brown formula, which shows how more intensive measurement increases reliability, is extended to encompass the dbICC.

stat.ME

Estimating whole brain dynamics using spectral clustering

The estimation of time-varying networks for functional Magnetic Resonance Imaging (fMRI) data sets is of increasing importance and interest. In this work, we formulate the problem in a high-dimensional time series framework and introduce a data-driven method, namely Network Change Points Detection (NCPD), which detects change points in the network structure of a multivariate time series, with each component of the time series represented by a node in the network. NCPD is applied to various simulated data and a resting-state fMRI data set. This new methodology also allows us to identify common functional states within and across subjects. Finally, NCPD promises to offer a deep insight into the large-scale characterisations and dynamics of the brain

stat.AP

Estimating Extremal Dependence in Univariate and Multivariate Time Series via the Extremogram

Davis and Mikosch [7] introduced the extremogram as a flexible quantitative tool for measuring various types of extremal dependence in a stationary time series. There we showed some standard statistical properties of the sample extremogram. A major difficulty was the construction of credible confidence bands for the extremogram. In this paper, we employ the stationary bootstrap to overcome this problem. Moreover, we introduce the cross extremogram as a measure of extremal serial dependence between two or more time series. We also study the extremogram for return times between extremal events. The use of the stationary bootstrap for the extremogram and the resulting interpretations are illustrated in several univariate and multivariate financial time series examples.

stat.ME