SearcharxivSearch

arXiv subjects

Pouria Saidi

Publications and source records attributed to Pouria Saidi.

4 recordsLinked to original sources

Unraveling overoptimism and publication bias in ML-driven science

Machine Learning (ML) is increasingly used across many disciplines with impressive reported results. However, recent studies suggest published performance of ML models are often overoptimistic. Validity concerns are underscored by findings of an inverse relationship between sample size and reported accuracy in published ML models, contrasting with the theory of learning curves where accuracy should improve or remain stable with increasing sample size. This paper investigates factors contributing to overoptimism in ML-driven science, focusing on overfitting and publication bias. We introduce a novel stochastic model for observed accuracy, integrating parametric learning curves and the aforementioned biases. We construct an estimator that corrects for these biases in observed data. Theoretical and empirical results show that our framework can estimate the underlying learning curve, providing realistic performance assessments from published results. Applying the model to meta-analyses of classifications of neurological conditions, we estimate the inherent limits of ML-based prediction in each domain.

cs.LG

Active Sequential Two-Sample Testing

A two-sample hypothesis test is a statistical procedure used to determine whether the distributions generating two samples are identical. We consider the two-sample testing problem in a new scenario where the sample measurements (or sample features) are inexpensive to access, but their group memberships (or labels) are costly. To address the problem, we devise the first \emph{active sequential two-sample testing framework} that not only sequentially but also \emph{actively queries}. Our test statistic is a likelihood ratio where one likelihood is found by maximization over all class priors, and the other is provided by a probabilistic classification model. The classification model is adaptively updated and used to predict where the (unlabelled) features have a high dependency on labels; labeling the ``high-dependency'' features leads to the increased power of the proposed testing framework. In theory, we provide the proof that our framework produces an \emph{anytime-valid} $p$-value. In addition, we characterize the proposed framework's gain in testing power by analyzing the mutual information between the feature and label variables in asymptotic and finite-sample scenarios. In practice, we introduce an instantiation of our framework and evaluate it using several experiments; the experiments on the synthetic, MNIST, and application-specific datasets demonstrate that the testing power of the instantiated active sequential test significantly increases while the Type I error is under control.

cs.LG

Support Recovery of Periodic Mixtures with Nested Periodic Dictionaries

Periodic signals composed of periodic mixtures admit sparse representations in nested periodic dictionaries (NPDs). Therefore, their underlying hidden periods can be estimated by recovering the exact support of said representations. In this paper, support recovery guarantees of such signals are derived both in noise-free and noisy settings. While exact recovery conditions have been studied in the theory of compressive sensing, existing conditions fall short of yielding meaningful achievability regions in the context of periodic signals with sparse representations in NPDs, in part since existing bounds do not capture structures intrinsic to these dictionaries. We leverage known properties of NPDs to derive several conditions for exact sparse recovery of periodic mixtures in the noise-free setting. These conditions rest on newly introduced notions of nested periodic coherence and restricted coherence, which can be efficiently computed. In the presence of noise, we obtain improved conditions for recovering the exact support set of the sparse representation of the periodic mixture via orthogonal matching pursuit based on the introduced notions of coherence. The theoretical findings are corroborated using numerical experiments for different families of NPDs. Our results show significant improvement over generic recovery bounds as the conditions hold over a larger range of sparsity levels.

cs.IT

Detection of Brain Stimuli Using Ramanujan Periodicity Transforms

The ability to efficiently match the frequency of the brain's response to repetitive visual stimuli in real time is the basis for reliable SSVEP-based Brain-Computer-Interfacing (BCI). The detection of different stimuli is posed as a composite hypothesis test, where SSVEPs are assumed to admit a sparse representation in a Ramanujan Periodicity Transform (RPT) dictionary. For the binary case, we develop and analyze the performance of an RPT detector based on a derived generalized likelihood ratio test. Our approach is extended to multi-hypothesis multi-electrode settings, where we capture the spatial correlation between the electrodes using pre-stimulus data. We also introduce a new metric for evaluating SSVEP detection schemes based on their achievable efficiency and discrimination rate tradeoff for given system resources. We obtain exact distributions of the test statistic in terms of confluent hypergeometric functions. Results based on extensive simulations with both synthesized and real data indicate that the RPT detector substantially outperforms spectral-based methods. Its performance also surpasses the state-of-the-art Canonical Correlation Analysis (CCA) methods with respect to accuracy and sample complexity in short data lengths regimes crucial for real-time applications. The proposed approach is asymptotically optimal as it closes the gap to a perfect measurement bound as the data length increases. In contrast to existing supervised methods which are highly data-dependent, the RPT detector only uses pre-stimulus data to estimate the per-subject spatial correlation, thereby dispensing with considerable overhead associated with data collection for a large number of subjects and stimuli. Our work advances the theory and practice of emerging real-time BCI and affords a new framework for comparing SSVEP detection schemes across a wider spectrum of operating regimes.

eess.SP