Searcharxiv⌕ Search

arXiv subjects

Stanley E. Lazic

Publications and source records attributed to Stanley E. Lazic.

10 recordsLinked to original sources

Relative plausibility versus probabilism: A level-of-analysis error in juridical proof

Debates about juridical proof are often framed as a conflict between probabilistic approaches and relative plausibility theory (RPT). This paper argues that this opposition rests on a level-of-analysis error. Drawing on Marr's distinction between levels of analysis, we show that RPT and probabilistic approaches operate at different conceptual levels and are therefore compatible rather than competing theories. RPT provides a computational-level description of juridical proof, characterizing the task of comparing explanations in light of the evidence and assessing whether a standard of proof has been met. Probabilistic approaches supply algorithmic-level accounts that specify how such comparative assessments can be represented and computed. When plausibility judgments satisfy minimal coherence conditions, relative plausibility corresponds to posterior odds. Recognizing this distinction clarifies longstanding disputes and highlights the complementary roles of explanation and probability in legal reasoning.

stat.AP↗

A weighted-likelihood framework for class imbalance in Bayesian prediction models

Class imbalance is a pervasive problem in predictive toxicology, where the number of non-toxic compounds often exceeds the number of toxic ones. Models trained on such data often perform well on the majority class but poorly on the minority class, which is most relevant for safety assessment. We propose a simple and general Bayesian framework that addresses class imbalance by modifying the likelihood function. Each observation's likelihood is raised to a power inversely proportional to its class proportion, with the weights normalized to preserve the overall information content. This weighted-likelihood (or power-likelihood) approach embeds cost-sensitive learning directly into Bayesian updating. The method is demonstrated using simulated binary data and an ordered logistic model for drug-induced liver injury (DILI). Weighting alters parameter estimates and decision boundaries, improving balanced accuracy and sensitivity for the minority (toxic) class. The approach can be implemented with minimal changes in standard probabilistic programming languages such as Stan, PyMC, and Turing.jl. This framework provides an easily extensible foundation for developing Bayesian prediction models that better reflect the asymmetric costs of safety-critical decisions.

stat.AP↗

Internal replication as a tool for evaluating reproducibility in preclinical experiments

Reproducibility is central to the credibility of scientific findings, yet complete replication studies are costly and infrequent. However, many biological experiments contain internal replication, which is defined as repetition across batches, runs, days, litters, or sites that can be used to estimate reproducibility without requiring additional experiments. This internal replication is analogous to internal validation in prediction or machine learning models, but is often treated as a nuisance and removed by normalisation, missing an opportunity to assess the stability of results. Here, six types of internal replication are defined based on independence and timing. Using mice data from an experiment conducted at three independent sites, we demonstrate how to quantify and test for internal reproducibility. This approach provides a framework for quantifying reproducibility from existing data and reporting more robust statistical inferences in preclinical research.

stat.AP↗

The ultimate issue error in scientific inference: mistaking parameters for hypotheses

Statistical inference often conflates the probability of a parameter with the probability of a hypothesis, a critical misunderstanding termed the ultimate issue error. This error is pervasive across the social, biological, and medical sciences, where null hypothesis significance testing (NHST) is mistakenly understood to be testing hypotheses rather than evaluating parameter estimates. Here, we advocate for using the Weight of Evidence (WoE) approach, which integrates quantitative data with qualitative background information for more accurate and transparent inference. Through a detailed example involving the relationship between vitamin D (25-hydroxy vitamin D) levels and COVID-19 risk, we demonstrate how WoE quantifies support for hypotheses while accounting for study design biases, power, and confounding factors. These findings emphasise the necessity of combining statistical metrics with contextual evaluation. This offers a structured framework to enhance reproducibility, reduce false interpretations, and foster robust scientific conclusions across disciplines.

stat.ME↗

Why multiple hypothesis test corrections provide poor control of false positives in the real world

Most scientific disciplines use significance testing to draw conclusions about experimental or observational data. This classical approach provides a theoretical guarantee for controlling the number of false positives across a set of hypothesis tests, making it an appealing framework for scientists seeking to limit the number of false effects or associations that they claim to observe. Unfortunately, this theoretical guarantee applies to few experiments, and the true false positive rate (FPR) is much higher. Scientists have plenty of freedom to choose the error rate to control, the tests to include in the adjustment, and the method of correction, making strong error control difficult to attain. In addition, hypotheses are often tested after finding unexpected relationships or patterns, the data are analysed in several ways, and analyses may be run repeatedly as data accumulate. As a result, adjusted p-values are too small, incorrect conclusions are often reached, and results are harder to reproduce. In the following, I argue why the FPR is rarely controlled meaningfully and why shrinking parameter estimates is preferable to p-value adjustments.

stat.AP↗

Quantifying sources of uncertainty in drug discovery predictions with probabilistic models

Knowing the uncertainty in a prediction is critical when making expensive investment decisions and when patient safety is paramount, but machine learning (ML) models in drug discovery typically provide only a single best estimate and ignore all sources of uncertainty. Predictions from these models may therefore be over-confident, which can put patients at risk and waste resources when compounds that are destined to fail are further developed. Probabilistic predictive models (PPMs) can incorporate uncertainty in both the data and model, and return a distribution of predicted values that represents the uncertainty in the prediction. PPMs not only let users know when predictions are uncertain, but the intuitive output from these models makes communicating risk easier and decision making better. Many popular machine learning methods have a PPM or Bayesian analogue, making PPMs easy to fit into current workflows. We use toxicity prediction as a running example, but the same principles apply for all prediction models used in drug discovery. The consequences of ignoring uncertainty and how PPMs account for uncertainty are also described. We aim to make the discussion accessible to a broad non-mathematical audience. Equations are provided to make ideas concrete for mathematical readers (but can be skipped without loss of understanding) and code is available for computational researchers (https://github.com/stanlazic/ML_uncertainty_quantification).

cs.LG↗

Quantifying the behavioural relevance of hippocampal neurogenesis

Few studies that examine the neurogenesis--behaviour relationship formally establish covariation between neurogenesis and behaviour or rule out competing explanations. The behavioural relevance of neurogenesis might therefore be overestimated if other mechanisms account for some, or even all, of the experimental effects. A systematic review of the literature was conducted and the data reanalysed using causal mediation analysis, which can estimate the behavioural contribution of new hippocampal neurons separately from other mechanisms that might be operating. Results from eleven eligible individual studies were then combined in a meta-analysis to increase precision (representing data from 215 animals) and showed that neurogenesis made a negligible contribution to behaviour (standarised effect = 0.15; 95% CI = -0.04 to 0.34; p = 0.128); other mechanisms accounted for the majority of experimental effects (standardised effect = 1.06; 95% CI = 0.74 to 1.38; p = 1.7 $\times 10^{-11}$).

q-bio.NC↗

Improving basic and translational science by accounting for litter-to-litter variation in animal models

Background: Animals from the same litter are often more alike compared with animals from different litters. This litter-to-litter variation, or "litter effects", can influence the results in addition to the experimental factors of interest. Furthermore, an experimental treatment can be applied to whole litters rather than to individual offspring. For example, in the valproic acid (VPA) model of autism, VPA is administered to pregnant females thereby inducing the disease phenotype in the offspring. With this type of experiment the sample size is the number of litters and not the total number of offspring. If such experiments are not appropriately designed and analysed, the results can be severely biased as well as extremely underpowered. Results: A review of the VPA literature showed that only 9% (3/34) of studies correctly determined that the experimental unit (n) was the litter and therefore made valid statistical inferences. In addition, litter effects accounted for up to 61% (p <0.001) of the variation in behavioural outcomes, which was larger than the treatment effects. In addition, few studies reported using randomisation (12%) or blinding (18%), and none indicated that a sample size calculation or power analysis had been conducted. Conclusions: Litter effects are common, large, and ignoring them can make replication of findings difficult and can contribute to the low rate of translating preclinical in vivo studies into successful therapies. Only a minority of studies reported using rigorous experimental methods, which is consistent with much of the preclinical in vivo literature.

q-bio.QM↗

Using causal models to distinguish between neurogenesis-dependent and -independent effects on behaviour

There has been a substantial amount of research on the relationship between hippocampal neurogenesis and behaviour over the past fifteen years, but the causal role that new neurons have on cognitive and affective behavioural tasks is still far from clear. This is partly due to the difficulty of manipulating levels of neurogenesis without inducing off-target effects, which might also influence behaviour. In addition, the analytical methods typically used do not directly test whether neurogenesis mediates the effect of an intervention on behaviour. Previous studies may have incorrectly attributed changes in behavioural performance to neurogenesis because the role of known (or unknown) neurogenesis-independent mechanisms were not formally taken into consideration during the analysis. Causal models can tease apart complex causal relationships and were used to demonstrate that the effect of exercise on pattern separation is via neurogenesis-independent mechanisms. Many studies in the neurogenesis literature would benefit from the use of statistical methods that can separate neurogenesis-dependent from neurogenesis-independent effects on behaviour.

q-bio.NC↗

Modelling hippocampal neurogenesis across the lifespan in seven species

The aim of this study was to estimate the number of new cells and neurons added to the dentate gyrus across the lifespan, and to compare the rate of age-associated decline in neurogenesis across species. Data from mice (Mus musculus), rats (Rattus norvegicus), lesser hedgehog tenrecs (Echinops telfairi), macaques (Macaca mulatta), marmosets (Callithrix jacchus), tree shrews (Tupaia belangeri), and humans (Homo sapiens) were extracted from twenty one data sets published in fourteen different papers. ANOVA, exponential, Weibull, and power models were fit to the data to determine which best described the relationship between age and neurogenesis. Exponential models provided a suitable fit and were used to estimate the relevant parameters. The rate of decrease of neurogenesis correlated with species longevity r = 0.769, p = 0.043), but not body mass or basal metabolic rate. Of all the cells added postnatally to the mouse dentate gyrus, only 8.5% (95% CI = 1.0% to 14.7%) of these will be added after middle age. In addition, only 5.7% (95% CI = 0.7% to 9.9%) of the existing cell population turns over from middle age onwards. Thus, relatively few new cells are added for much of an animal's life, and only a proportion of these will mature into functional neurons.

q-bio.NC↗