SearcharxivSearch

arXiv subjects

Georgios Fellouris

Publications and source records attributed to Georgios Fellouris.

At least 19 recordsLinked to original sources

Active Sequential Signal Detection with Asynchronous Decisions

This work considers the problem of detecting signals from multiple sequentially observed data streams, where only one stream can be observed at every time instant. The goal is to detect signals as quickly as possible while controlling the global probabilities of false alarm and missed detection. In this active sampling setup, it is impossible to minimize the expected detection time simultaneously for every signal, so we formulate a novel set of performance criteria that aim to minimize the expectations of the order statistics of the detection times. A novel procedure is proposed, which incorporates an exploration mechanism to a "follow-the-leader" procedure, and is shown to optimize all the criteria asymptotically as the global error probabilities go to zero. Its finite-sample performance is compared with existing and oracle procedures in simulation studies.

stat.ME

Ordering sampling rules for sequential anomaly identification under sampling constraints

We consider the problem of sequential anomaly identification over multiple independent data streams, under the presence of a sampling constraint. The goal is to quickly identify those that exhibit anomalous statistical behavior, when it is not possible to sample every source at each time instant. Thus, in addition to a stopping rule that determines when to stop sampling, and a decision rule that indicates which sources to identify as anomalous upon stopping, one needs to specify a sampling rule that determines which sources to sample at each time instant. We focus on the family of ordering sampling rules that select the sources to be sampled at each time instant based not only on the currently estimated subset of anomalous sources as the probabilistic sampling rules \cite{Tsopela_2022}, but also on the ordering of the sources' test-statistics. We show that under an appropriate design specified explicitly, an ordering sampling rule leads to the optimal expected time for stopping among all policies that satisfy the same sampling and error constraints to a first-order asymptotic approximation as the false positive and false negative error thresholds go to zero. This is the first asymptotic optimality result for ordering sampling rules, when more than one sources can be sampled per time instant, and it is established under a general setup where the number of anomalous sources is not required to be known. A novel proof technique is introduced that encompasses all different cases of the problem concerning sources' homogeneity, and prior information on the number of anomalies. Simulations show that ordering sampling rules have better performance in finite regime compared to probabilistic sampling rules.

math.ST

Efficient Importance Sampling for Wrong Exit Probabilities over Combinatorially Many Rare Regions

We consider importance sampling for estimating the probability that a light-tailed $d$-dimensional random walk exits through one of many disjoint rare-event regions before reaching an anticipated target. This problem arises in sequential multiple hypothesis testing, where the number of such regions may grow combinatorially and in some cases exponentially with the dimension. While mixtures over all associated exponential tilts are asymptotically efficient, they become computationally infeasible even for moderate values of $d$. We develop a method for constructing asymptotically efficient mixtures with substantially fewer components by combining optimal tilts for a small number of regions with additional proposals that control variance across a large collection of regions. The approach is applied to the estimation of three probabilities that arise in sequential multiple testing, including a multidimensional extension of Siegmund's classical exit problem, and is supported by both theoretical analysis and numerical experiments.

math.PR

Sequential anomaly identification with observation control under generalized error metrics

The problem of sequential anomaly detection and identification is considered, where multiple data sources are simultaneously monitored and the goal is to identify in real time those, if any, that exhibit ``anomalous" statistical behavior. An upper bound is postulated on the number of data sources that can be sampled at each sampling instant, but the decision maker selects which ones to sample based on the already collected data. Thus, in this context, a policy consists not only of a stopping rule and a decision rule that determine when sampling should be terminated and which sources to identify as anomalous upon stopping, but also of a sampling rule that determines which sources to sample at each time instant subject to the sampling constraint. Two distinct formulations are considered, which require control of different, ``generalized" error metrics. The first one tolerates a certain user-specified number of errors, of any kind, whereas the second tolerates distinct, user-specified numbers of false positives and false negatives. For each of them, a universal asymptotic lower bound on the expected time for stopping is established as the error probabilities go to 0, and it is shown to be attained by a policy that combines the stopping and decision rules proposed in the full-sampling case with a probabilistic sampling rule that achieves a specific long-run sampling frequency for each source. Moreover, the optimal to a first order asymptotic approximation expected time for stopping is compared in simulation studies with the corresponding factor in a finite regime, and the impact of the sampling constraint and tolerance to errors is assessed.

math.ST

Change Acceleration and Detection

A novel sequential change detection problem is proposed, in which the goal is to not only detect but also accelerate the change. Specifically, it is assumed that the sequentially collected observations are responses to treatments selected in real time. The assigned treatments determine the pre-change and post-change distributions of the responses and also influence when the change happens. The goal is to find a treatment assignment rule and a stopping rule that minimize the expected total number of observations subject to a user-specified bound on the false alarm probability. The optimal solution is obtained under a general Markovian change-point model. Moreover, an alternative procedure is proposed, whose applicability is not restricted to Markovian change-point models and whose design requires minimal computation. For a large class of change-point models, the proposed procedure is shown to achieve the optimal performance in an asymptotic sense. Finally, its performance is found in simulation studies to be comparable to the optimal, uniformly with respect to the error probability.

math.ST

Round Robin Active Sequential Change Detection for Dependent Multi-Channel Data

This paper considers the problem of sequentially detecting a change in the joint distribution of multiple data sources under a sampling constraint. Specifically, the channels or sources generate observations that are independent over time, but not necessarily independent at any given time instant. The sources follow an initial joint distribution, and at an unknown time instant, the joint distribution of an unknown subset of sources changes. Importantly, there is a hard constraint that only a fixed number of sources are allowed to be sampled at each time instant. The goal is to sequentially observe the sources according to the constraint, and stop sampling as quickly as possible after the change while controlling the false alarm rate below a user-specified level. The sources can be selected dynamically based on the already collected data, and thus, a policy for this problem consists of a joint sampling and change-detection rule. A non-randomized policy is studied, and an upper bound is established on its worst-case conditional expected detection delay with respect to both the change point and the observations from the affected sources before the change.

stat.ME

Quickest Change Detection with Controlled Sensing

In the problem of quickest change detection, a change occurs at some unknown time in the distribution of a sequence of random vectors that are monitored in real time, and the goal is to detect this change as quickly as possible subject to a certain false alarm constraint. In this work we consider this problem in the presence of parametric uncertainty in the post-change regime and controlled sensing. That is, the post-change distribution contains an unknown parameter, and the distribution of each observation, before and after the change, is affected by a control action. In this context, in addition to a stopping rule that determines the time at which it is declared that the change has occurred, one also needs to determine a sequential control policy, which chooses the control action at each time based on the already collected observations. We formulate this problem mathematically using Lorden's minimax criterion, and assuming that there are finitely many possible actions and post-change parameter values. We then propose a specific procedure for this problem that employs an adaptive CuSum statistic in which (i) the estimate of the parameter is based on a fixed number of the more recent observations, and (ii) each action is selected to maximize the Kullback-Leibler divergence of the next observation based on the current parameter estimate, apart from a small number of exploration times. We show that this procedure, which we call the Windowed Chernoff-CuSum (WCC), is first-order asymptotically optimal under Lorden's minimax criterion, for every possible possible value of the unknown post-change parameter, as the mean time to false alarm goes to infinity. We also provide simulation results to illustrate the performance of the WCC procedure.

cs.IT

Worst-Case Misidentification Control in Sequential Change Diagnosis using the min-CuSum

The problem of sequential change diagnosis is considered, where a sequence of independent random elements is accessed sequentially, there is an abrupt change in its distribution at some unknown time, and there are two main operational goals: to quickly detect the change and, upon stopping, to accurately identify the post-change distribution among a finite set of alternatives. The algorithm that raises an alarm as soon as the CuSum statistic that corresponds to one of the post-change alternatives exceeds a certain threshold is studied. When the data are generated over independent channels and the change can occur in only one of them, its worst-case with respect to the change point conditional probability of misidentification, given that there was not a false alarm, is shown to decay exponentially fast in the threshold. As a corollary, in this setup, this algorithm is shown to asymptotically minimize Lorden's detection delay criterion, simultaneously for every possible post-change distribution, within the class of schemes that satisfy prescribed bounds on the false alarm rate and the worst-case conditional probability of misidentification, as the former goes to zero sufficiently faster than the latter. Finally, these theoretical results are also illustrated in simulation studies.

math.ST

Asymptotically Optimal Sequential Multiple Testing with Asynchronous Decisions

The problem of simultaneously testing the marginal distributions of sequentially monitored, independent data streams is considered. The decisions for the various testing problems can be made at different times, using data from all streams, which can be monitored until all decisions have been made. Moreover, arbitrary a priori bounds are assumed on the number of signals, i.e., data streams in which the alternative hypothesis is correct. A novel sequential multiple testing procedure is proposed and it is shown to achieve the minimum expected decision time, simultaneously in every data stream and under every signal configuration, asymptotically as certain metrics of global error rates go to zero. This optimality property is established under general parametric composite hypotheses, various error metrics, and weak distributional assumptions that allow for temporal dependence. Furthermore, the limit of the factor by which the expected decision time in a data stream increases when one is limited to synchronous or decentralized procedures is evaluated. Finally, two existing sequential multiple testing procedures in the literature are compared with the proposed one in various simulation studies.

stat.ME

Sequential Change Diagnosis Revisited and the Adaptive Matrix CuSum

The problem of sequential change diagnosis is considered, where observations are obtained on-line, an abrupt change occurs in their distribution, and the goal is to quickly detect the change and accurately identify the post-change distribution, while controlling the false alarm rate. A finite set of alternatives is postulated for the post-change regime, but no prior information is assumed for the unknown change-point. A drawback of many algorithms that have been proposed for this problem is the implicit use of pre-change data for determining the post-change distribution. This can lead to very large conditional probabilities of misidentification, given that there was no false alarm, unless the change occurs soon after monitoring begins. A novel, recursive algorithm is proposed and shown to resolve this issue without the use of additional tuning parameters and without sacrificing control of the worst-case delay in Lorden's sense. A theoretical analysis is conducted for a general family of sequential change diagnosis procedures, which supports the proposed algorithm and revises certain state-of-the-art results. Additionally, a novel, comprehensive method is proposed for the design and evaluation of sequential change diagnosis algorithms. This method is illustrated with simulation studies, where existing procedures are compared to the proposed.

math.ST

Signal Recovery With Multistage Tests And Without Sparsity Constraints

A signal recovery problem is considered, where the same binary testing problem is posed over multiple, independent data streams. The goal is to identify all signals, i.e., streams where the alternative hypothesis is correct, and noises, i.e., streams where the null hypothesis is correct, subject to prescribed bounds on the classical or generalized familywise error probabilities. It is not required that the exact number of signals be a priori known, only upper bounds on the number of signals and noises are assumed instead. A decentralized formulation is adopted, according to which the sample size and the decision for each testing problem must be based only on observations from the corresponding data stream. A novel multistage testing procedure is proposed for this problem and is shown to enjoy a high-dimensional asymptotic optimality property. Specifically, it achieves the optimal, average over all streams, expected sample size, uniformly in the true number of signals, as the maximum possible numbers of signals and noises go to infinity at arbitrary rates, in the class of all sequential tests with the same global error control. In contrast, existing multistage tests in the literature are shown to achieve this high-dimensional asymptotic optimality property only under additional sparsity or symmetry conditions. These results are based on an asymptotic analysis for the fundamental binary testing problem as the two error probabilities go to zero. For this problem, unlike existing multistage tests in the literature, the proposed test achieves the optimal expected sample size under both hypotheses, in the class of all sequential tests with the same error control, as the two error probabilities go to zero at arbitrary rates. These results are further supported by simulation studies and extended to problems with non-iid data and composite hypotheses.

stat.ME

Joint Sequential Detection and Isolation for Dependent Data Streams

The problem of joint sequential detection and isolation is considered in the context of multiple, not necessarily independent, data streams. A multiple testing framework is proposed, where each hypothesis corresponds to a different subset of data streams, the sample size is a stopping time of the observations, and the probabilities of four kinds of error are controlled below distinct, user-specified levels. Two of these errors reflect the detection component of the formulation, whereas the other two the isolation component. The optimal expected sample size is characterized to a first-order asymptotic approximation as the error probabilities go to 0. Different asymptotic regimes, expressing different prioritizations of the detection and isolation tasks, are considered. A novel, versatile family of testing procedures is proposed, in which two distinct, in general, statistics are computed for each hypothesis, one addressing the detection task and the other the isolation task. Tests in this family, of various computational complexities, are shown to be asymptotically optimal under different setups. The general theory is applied to the detection and isolation of anomalous, not necessarily independent, data streams, as well as to the detection and isolation of an unknown dependence structure.

math.ST

3-stage and 4-stage tests with deterministic stage sizes and non-iid data

Given a fixed-sample-size test that controls the error probabilities under two specific, but arbitrary, distributions, a 3-stage and two 4-stage tests are proposed and analyzed. For each of them, a novel, concrete, non-asymptotic, non-conservative design is specified, which guarantees the same error control as the given fixed-sample-size test. Moreover, first-order asymptotic approximation are established on their expected sample sizes under the two prescribed distributions as the error probabilities go to zero. As a corollary, it is shown that the proposed multistage tests can achieve, in this asymptotic sense, the optimal expected sample size under these two distributions in the class of all sequential tests with the same error control. Furthermore, they are shown to be much more robust than Wald's SPRT when applied to one-sided testing problems and the error probabilities under control are small enough. These general results are applied to testing problems in the iid setup and beyond, such as testing the correlation coefficient of a first-order autoregression, or the transition matrix of a finite-state Markov chain, and are illustrated in various numerical studies.

math.ST

Sequential anomaly detection with sampling constraints

The problem of sequential anomaly detection is considered, where multiple data sources are monitored in real time and the goal is to identify the "anomalous" ones among them, when it is not possible to sample all sources at all times. A detection scheme in this context requires specifying not only when to stop sampling and which sources to identify as anomalous upon stopping, but also which sources to sample at each time instance until stopping. A novel formulation for this problem is proposed, in which the number of anomalous sources is not necessarily known in advance and the number of sampled sources per time instance is not necessarily fixed. Instead, an arbitrary lower bound and an arbitrary upper bound are assumed on the number of anomalous sources, and the fraction of the expected number of samples over the expected time until stopping is required to not exceed an arbitrary, user-specified level. In addition to this sampling constraint, the probabilities of at least one false alarm and at least one missed detection are controlled below user-specified tolerance levels. A general criterion is established for a policy to achieve the minimum expected time until stopping to a first-order asymptotic approximation as both familywise error rates go to zero. This criterion is used to prove the asymptotic optimality of a family of policies that sample each source at each time instance with a probability that depends on the past observations only through the current estimate of the subset of anomalous sources. In particular, the asymptotic optimality is established of a policy that requires minimal computation under any setup of the problem.

math.ST

Quickest Change Detection under Transient Dynamics: Theory and Asymptotic Analysis

The problem of quickest change detection (QCD) under transient dynamics is studied, where the change from the initial distribution to the final persistent distribution does not happen instantaneously, but after a series of transient phases. The observations within the different phases are generated by different distributions. The objective is to detect the change as quickly as possible, while controlling the average run length (ARL) to false alarm, when the durations of the transient phases are completely unknown. Two algorithms are considered, the dynamic Cumulative Sum (CuSum) algorithm, proposed in earlier work, and a newly constructed weighted dynamic CuSum algorithm. Both algorithms admit recursions that facilitate their practical implementation, and they are adaptive to the unknown transient durations. Specifically, their asymptotic optimality is established with respect to both Lorden's and Pollak's criteria as the ARL to false alarm and the durations of the transient phases go to infinity at any relative rate. Numerical results are provided to demonstrate the adaptivity of the proposed algorithms, and to validate the theoretical results.

math.ST

Sequential multiple testing with generalized error control: an asymptotic optimality theory

The sequential multiple testing problem is considered under two generalized error metrics. Under the first one, the probability of at least $k$ mistakes, of any kind, is controlled. Under the second, the probabilities of at least $k_1$ false positives and at least $k_2$ false negatives are simultaneously controlled. For each formulation, the optimal expected sample size is characterized, to a first-order asymptotic approximation as the error probabilities go to 0, and a novel multiple testing procedure is proposed and shown to be asymptotically efficient under every signal configuration. These results are established when the data streams for the various hypotheses are independent and each local log-likelihood ratio statistic satisfies a certain Strong Law of Large Numbers. In the special case of i.i.d. observations in each stream, the gains of the proposed sequential procedures over fixed-sample size schemes are quantified.

math.ST

Efficient Byzantine Sequential Change Detection

In the multisensor sequential change detection problem, a disruption occurs in an environment monitored by multiple sensors. This disruption induces a change in the observations of an unknown subset of sensors. In the Byzantine version of this problem, which is the focus of this work, it is further assumed that the postulated change-point model may be misspecified for an unknown subset of sensors. The problem then is to detect the change quickly and reliably, for any possible subset of affected sensors, even if the misspecified sensors are controlled by an adversary. Given a user-specified upper bound on the number of compromised sensors, we propose and study three families of sequential change-detection rules for this problem. These are designed and evaluated under a generalization of Lorden's criterion, where conditional expected detection delay and expected time to false alarm are both computed in the worst-case scenario for the compromised sensors. The first-order asymptotic performance of these procedures is characterized as the worst-case false alarm rate goes to 0. The insights from these theoretical results are corroborated by a simulation study.

math.ST

Asymptotically optimal, sequential, multiple testing procedures with prior information on the number of signals

Assuming that data are collected sequentially from independent streams, we consider the simultaneous testing of multiple binary hypotheses under two general setups; when the number of signals (correct alternatives) is known in advance, and when we only have a lower and an upper bound for it. In each of these setups, we propose feasible procedures that control, without any distributional assumptions, the familywise error probabilities of both type I and type II below given, user-specified levels. Then, in the case of i.i.d. observations in each stream, we show that the proposed procedures achieve the optimal expected sample size, under every possible signal configuration, asymptotically as the two error probabilities vanish at arbitrary rates. A simulation study is presented in a completely symmetric case and supports insights obtained from our asymptotic results, such as the fact that knowledge of the exact number of signals roughly halves the expected number of observations compared to the case of no prior information.

math.ST