SearcharxivSearch

arXiv subjects

Michael Wood

Publications and source records attributed to Michael Wood.

10 recordsLinked to original sources

A Beamdump Facility at Jefferson Lab

This White Paper is exploring the potential of intense secondary muon, neutrino, and (hypothetical) light dark matter beams produced in interactions of high-intensity electron beams with beam dumps. Light dark matter searches with the approved Beam Dump eXperiment (BDX) are driving the realization of a new underground vault at Jefferson Lab that could be extended to a Beamdump Facility with minimal additional installations. The paper summarizes contributions and discussions from the International Workshop on Secondary Beams at Jefferson Lab (BDX & Beyond). Several possible muon physics applications and neutrino detector technologies for Jefferson Lab are highlighted. The potential of a secondary neutron beam will be addressed in a future edition.

physics.acc-ph

MindBenchAI: An Actionable Platform to Evaluate the Profile and Performance of Large Language Models in a Mental Healthcare Context

Individuals are increasingly utilizing large language model (LLM)based tools for mental health guidance and crisis support in place of human experts. While AI technology has great potential to improve health outcomes, insufficient empirical evidence exists to suggest that AI technology can be deployed as a clinical replacement; thus, there is an urgent need to assess and regulate such tools. Regulatory efforts have been made and multiple evaluation frameworks have been proposed, however,field-wide assessment metrics have yet to be formally integrated. In this paper, we introduce a comprehensive online platform that aggregates evaluation approaches and serves as a dynamic online resource to simplify LLM and LLM-based tool assessment: MindBenchAI. At its core, MindBenchAI is designed to provide easily accessible/interpretable information for diverse stakeholders (patients, clinicians, developers, regulators, etc.). To create MindBenchAI, we built off our work developing MINDapps.org to support informed decision-making around smartphone app use for mental health, and expanded the technical MINDapps.org framework to encompass novel large language model (LLM) functionalities through benchmarking approaches. The MindBenchAI platform is designed as a partnership with the National Alliance on Mental Illness (NAMI) to provide assessment tools that systematically evaluate LLMs and LLM-based tools with objective and transparent criteria from a healthcare standpoint, assessing both profile (i.e. technical features, privacy protections, and conversational style) and performance characteristics (i.e. clinical reasoning skills).

cs.HC

On Bayesian Search for the Feasible Space Under Computationally Expensive Constraints

We are often interested in identifying the feasible subset of a decision space under multiple constraints to permit effective design exploration. If determining feasibility required computationally expensive simulations, the cost of exploration would be prohibitive. Bayesian search is data-efficient for such problems: starting from a small dataset, the central concept is to use Bayesian models of constraints with an acquisition function to locate promising solutions that may improve predictions of feasibility when the dataset is augmented. At the end of this sequential active learning approach with a limited number of expensive evaluations, the models can accurately predict the feasibility of any solution obviating the need for full simulations. In this paper, we propose a novel acquisition function that combines the probability that a solution lies at the boundary between feasible and infeasible spaces (representing exploitation) and the entropy in predictions (representing exploration). Experiments confirmed the efficacy of the proposed function.

cs.LG

Simple Methods for Estimating Confidence Levels, or Tentative Probabilities, for Hypotheses Instead of P Values

In many fields of research null hypothesis significance tests and p values are the accepted way of assessing the degree of certainty with which research results can be extrapolated beyond the sample studied. However, there are very serious concerns about the suitability of p values for this purpose. An alternative approach is to cite confidence intervals for a statistic of interest, but this does not directly tell readers how certain a hypothesis is. Here, I suggest how the framework used for confidence intervals could easily be extended to derive confidence levels, or "tentative probabilities", for hypotheses. I also outline four quick methods for estimating these. This allows researchers to state their confidence in a hypothesis as a direct probability, instead of circuitously by p values referring to an unstated, hypothetical null hypothesis. The inevitable difficulties of statistical inference mean that these probabilities can only be tentative, but probabilities are the natural way to express uncertainties, so, arguably, researchers using statistical methods have an obligation to estimate how probable their hypotheses are by the best available method. Otherwise misinterpretations will fill the void. Key words: Confidence, Null hypothesis significance test, p value, Statistical inference

stat.ME

How sure are we? Two approaches to statistical inference

Suppose you are told that taking a statin will reduce your risk of a heart attack or stroke by 3% in the next ten years, or that women have better emotional intelligence than men. You may wonder how accurate the 3% is, or how confident we should be about the assertion about women's emotional intelligence, bearing in mind that these conclusions are only based on samples of data? My aim here is to present two statistical approaches to questions like these. Approach 1 is often called null hypothesis testing but I prefer the phrase "baseline hypothesis": this is the standard approach in many areas of inquiry but is fraught with problems. Approach 2 can be viewed as a generalisation of the idea of confidence intervals, or as the application of Bayes' theorem. Unlike Approach 1, Approach 2 provides a tentative estimate of the probability of hypotheses of interest. For both approaches, I explain, from first principles, building only on "common sense" statistical concepts like averages and randomness, both how to derive answers, and the rationale behind the answers. This is achieved by using computer simulation methods (resampling and bootstrapping using a spreadsheet available on the web) which avoid the use of probability distributions (t, normal, etc). Such a minimalist, but reasonably rigorous, analysis is particularly useful in a discipline like statistics which is widely used by people who are not specialists. My intended audience includes both statisticians, and users of statistical methods who are not statistical experts.

stat.OT

Beyond p values: practical methods for analyzing uncertainty in research

This article explains, and discusses the merits of, three approaches for analyzing the certainty with which statistical results can be extrapolated beyond the data gathered. Sometimes it may be possible to use more than one of these approaches. (1) If there is an exact null hypothesis which is credible and interesting (usually not the case), researchers should cite a p value (significance level), although jargon is best avoided. (2) If the research result is a numerical value, researchers should cite a confidence interval. (3) If there are one or more hypotheses of interest, it may be possible to adapt the methods used for confidence intervals to derive an "estimated probability" for each. Under certain circumstances these could be interpreted as Bayesian posterior probabilities. These estimated probabilities can easily be worked out from the p values and confidence intervals produced by packages such as SPSS. Estimating probabilities for hypotheses means researchers can give a direct answer to the question "How certain can we be that this hypothesis is right?".

stat.ME

P values, confidence intervals, or confidence levels for hypotheses?

Null hypothesis significance tests and p values are widely used despite very strong arguments against their use in many contexts. Confidence intervals are often recommended as an alternative, but these do not achieve the objective of assessing the credibility of a hypothesis, and the distinction between confidence and probability is an unnecessary confusion. This paper proposes a more straightforward (probabilistic) definition of confidence, and suggests how the idea can be applied to whatever hypotheses are of interest to researchers. The relative merits of the different approaches are discussed using a series of illustrative examples: usually confidence based approaches seem more transparent and useful, but there are some contexts in which p values may be appropriate. I also suggest some methods for converting results from one format to another. (The attractiveness of the idea of confidence is demonstrated by the widespread persistence of the completely incorrect idea that p=5% is equivalent to 95% confidence in the alternative hypothesis. In this paper I show how p values can be used to derive meaningful confidence statements, and the assumptions underlying the derivation.) Key words: Confidence interval, Confidence level, Hypothesis testing, Null hypothesis significance tests, P value, User friendliness.

stat.ME

Beyond journals and peer review: towards a more flexible ecosystem for scholarly communication

This article challenges the assumption that journals and peer review are essential for developing,evaluating and disseminating scientific and other academic knowledge. It suggests a more flexible ecosystem, and examines some of the possibilities this might facilitate. The market for academic outputs should be opened up by encouraging the separation of the dissemination service from the evaluation service. Publishing research in subject-specific journals encourages compartmentalising research into rigid categories. The dissemination of knowledge would be better served by an open access, web-based repository system encompassing all disciplines. There would then be a role for organisations to assess the items in this repository to help users find relevant, high-quality work. There could be a variety of such organisations which could enable reviews from peers to be supplemented with evaluation by non-peers from a variety of different perspectives: user reviews, statistical reviews, reviews from the perspective of different disciplines, and so on. This should reduce the inevitably conservative influence of relying on two or three peers, and make the evaluation system more critical, multi-dimensional and responsive to the requirements of different audience groups, changing circumstances, and new ideas. Non-peer review might make it easier to challenge dominant paradigms, and expanding the potential audience beyond a narrow group of peers might encourage the criterion of simplicity to be taken more seriously - which is essential if human knowledge is to continue to progress. Keywords: Academic journals, Growth of knowledge, Non-peer review, Paradigm change, Peer review,Scholarly communication, Science communication, Simplicity.

cs.DL

Making statistical methods in management research more useful: some suggestions from a case study

I present a critique of the methods used in a typical paper. This leads to three broad conclusions about the conventional use of statistical methods. First, results are often reported in an unnecessarily obscure manner. Second, the null hypothesis testing paradigm is deeply flawed: estimating the size of effects and citing confidence intervals or levels is usually better. Third, there are several issues, independent of the particular statistical concepts employed, which limit the value of any statistical approach: e.g. difficulties of generalizing to different contexts, and the weakness of some research in terms of the size of the effects found. The first two of these are easily remedied: I illustrate some of the possibilities by re-analyzing the data from the case study article. The third means that in some contexts a statistical approach may not be worthwhile. My case study is a management paper, but similar problems arise in other social sciences. Keywords: Confidence, Hypothesis testing, Null hypothesis significance tests, Philosophy of statistics, Statistical methods, User-friendliness.

stat.AP

Bootstrapping Confidence Levels for Hypotheses about Quadratic (U-Shaped) Regression Models

Bootstrapping can produce confidence levels for hypotheses about quadratic regression models - such as whether the U-shape is inverted, and the location of optima. The method has several advantages over conventional methods: it provides more, and clearer, information, and is flexible - it could easily be applied to a wide variety of different types of models. The utility of the method can be enhanced by formulating models with interpretable coefficients, such as the location and value of the optimum. Keywords: Bootstrap resampling; Confidence level; Quadratic model; Regression, U-shape.

stat.ME