SearcharxivSearch

arXiv subjects

Kevin Coakley

Publications and source records attributed to Kevin Coakley.

5 recordsLinked to original sources

Learning to be Reproducible: Custom Loss Design for Robust Neural Networks

To enhance the reproducibility and reliability of deep learning models, we address a critical gap in current training methodologies: the lack of mechanisms that ensure consistent and robust performance across runs. Our empirical analysis reveals that even under controlled initialization and training conditions, the accuracy of the model can exhibit significant variability. To address this issue, we propose a Custom Loss Function (CLF) that reduces the sensitivity of training outcomes to stochastic factors such as weight initialization and data shuffling. By fine-tuning its parameters, CLF explicitly balances predictive accuracy with training stability, leading to more consistent and reliable model performance. Extensive experiments across diverse architectures for both image classification and time series forecasting demonstrate that our approach significantly improves training robustness without sacrificing predictive performance. These results establish CLF as an effective and efficient strategy for developing more stable, reliable and trustworthy neural networks.

cs.LG

Examining the Effect of Implementation Factors on Deep Learning Reproducibility

Reproducing published deep learning papers to validate their conclusions can be difficult due to sources of irreproducibility. We investigate the impact that implementation factors have on the results and how they affect reproducibility of deep learning studies. Three deep learning experiments were ran five times each on 13 different hardware environments and four different software environments. The analysis of the 780 combined results showed that there was a greater than 6% accuracy range on the same deterministic examples introduced from hardware or software environment variations alone. To account for these implementation factors, researchers should run their experiments multiple times in different hardware and software environments to verify their conclusions are not affected.

cs.AI

Sources of Irreproducibility in Machine Learning: A Review

Background: Many published machine learning studies are irreproducible. Issues with methodology and not properly accounting for variation introduced by the algorithm themselves or their implementations are attributed as the main contributors to the irreproducibility.Problem: There exist no theoretical framework that relates experiment design choices to potential effects on the conclusions. Without such a framework, it is much harder for practitioners and researchers to evaluate experiment results and describe the limitations of experiments. The lack of such a framework also makes it harder for independent researchers to systematically attribute the causes of failed reproducibility experiments. Objective: The objective of this paper is to develop a framework that enable applied data science practitioners and researchers to understand which experiment design choices can lead to false findings and how and by this help in analyzing the conclusions of reproducibility experiments. Method: We have compiled an extensive list of factors reported in the literature that can lead to machine learning studies being irreproducible. These factors are organized and categorized in a reproducibility framework motivated by the stages of the scientific method. The factors are analyzed for how they can affect the conclusions drawn from experiments. A model comparison study is used as an example. Conclusion: We provide a framework that describes machine learning methodology from experimental design decisions to the conclusions inferred from them.

cs.LG

Performance of Test Supermartingale Confidence Intervals for the Success Probability of Bernoulli Trials

Given a composite null hypothesis H, test supermartingales are non-negative supermartingales with respect to H with initial value 1. Large values of test supermartingales provide evidence against H. As a result, test supermartingales are an effective tool for rejecting H, particularly when the p-values obtained are very small and serve as certificates against the null hypothesis. Examples include the rejection of local realism as an explanation of Bell test experiments in the foundations of physics and the certification of entanglement in quantum information science. Test supermartingales have the advantage of being adaptable during an experiment and allowing for arbitrary stopping rules. By inversion of acceptance regions, they can also be used to determine confidence sets. We use an example to compare the performance of test supermartingales for computing p-values and confidence intervals to Chernoff-Hoeffding bounds and the "exact" p-value. The example is the problem of inferring the probability of success in a sequence of Bernoulli trials. There is a cost in using a technique that has no restriction on stopping rules, and for a particular test supermartingale, our study quantifies this cost.

math.ST

Bell Inequalities for Continuously Emitting Sources

A common experimental strategy for demonstrating non-classical correlations is to show violation of a Bell inequality by measuring a continuously emitted stream of entangled photon pairs. The measurements involve the detection of photons by two spatially separated parties. The detection times are recorded and compared to quantify the violation. The violation critically depends on determining which detections are coincident. Because the recorded detection times have "jitter", coincidences cannot be inferred perfectly. In the presence of settings-dependent timing errors, this can allow a local-realistic system to show apparent violation--the so-called "coincidence loophole". Here we introduce a family of Bell inequalities based on signed, directed distances between the parties' sequences of recorded timetags. Given that the timetags are recorded for synchronized, fixed observation periods and that the settings choices are random and independent of the source, violation of these inequalities unambiguously shows non-classical correlations violating local realism. Distance-based Bell inequalities are generally useful for two-party configurations where the effective size of the measurement outcome space is large or infinite. We show how to systematically modify the underlying Bell functions to improve the signal to noise ratio and to quantify the significance of the violation.

quant-ph