SearcharxivSearch

arXiv subjects

William Becker

Publications and source records attributed to William Becker.

5 recordsLinked to original sources

ODD: A Benchmark Dataset for the Natural Language Processing based Opioid Related Aberrant Behavior Detection

Opioid related aberrant behaviors (ORABs) present novel risk factors for opioid overdose. This paper introduces a novel biomedical natural language processing benchmark dataset named ODD, for ORAB Detection Dataset. ODD is an expert-annotated dataset designed to identify ORABs from patients' EHR notes and classify them into nine categories; 1) Confirmed Aberrant Behavior, 2) Suggested Aberrant Behavior, 3) Opioids, 4) Indication, 5) Diagnosed opioid dependency, 6) Benzodiazepines, 7) Medication Changes, 8) Central Nervous System-related, and 9) Social Determinants of Health. We explored two state-of-the-art natural language processing models (fine-tuning and prompt-tuning approaches) to identify ORAB. Experimental results show that the prompt-tuning models outperformed the fine-tuning models in most categories and the gains were especially higher among uncommon categories (Suggested Aberrant Behavior, Confirmed Aberrant Behaviors, Diagnosed Opioid Dependence, and Medication Change). Although the best model achieved the highest 88.17% on macro average area under precision recall curve, uncommon classes still have a large room for performance improvement. ODD is publicly available.

cs.CL

A comprehensive comparison of total-order estimators for global sensitivity analysis

Sensitivity analysis helps identify which model inputs convey the most uncertainty to the model output. One of the most authoritative measures in global sensitivity analysis is the Sobol' total-order index, which can be computed with several different estimators. Although previous comparisons exist, it is hard to know which estimator performs best since the results are contingent on the benchmark setting defined by the analyst (the sampling method, the distribution of the model inputs, the number of model runs, the test function or model and its dimensionality, the weight of higher order effects or the performance measure selected). Here we compare several total-order estimators in an eight-dimension hypercube where these benchmark parameters are treated as random parameters. This arrangement significantly relaxes the dependency of the results on the benchmark design. We observe that the most accurate estimators are Razavi and Gupta's, Jansen's or Janon/Monod's for factor prioritization, and Jansen's, Janon/Monod's or Azzini and Rosati's for approaching the "true" total-order indices. The rest lag considerably behind. Our work helps analysts navigate the myriad of total-order formulae by reducing the uncertainty in the selection of the most appropriate estimator.

stat.AP

Computational Framework for Behind-The-Meter DER Techno-Economic Modeling and Optimization -- REopt Lite

The global energy system is undergoing a major transformation. Renewable energy generation is growing and is projected to accelerate further with the global emphasis on decarbonization. Furthermore, distributed generation is projected to play a significant role in the new energy system, and energy models are playing a key role in understanding how distributed generation can be integrated reliably and economically. The deployment of massive amounts of distributed generation requires understanding the interface of technology, economics, and policy in the energy modeling process. In this work, we present an end-to-end computational framework for distributed energy resource (DER) modeling, REopt Lite which addresses this need effectively. We describe the problem space, the building blocks of the model, the scaling capabilities of the design, the optimization formulation, and the accessibility of the model. We present a framework for accelerating the techno-economic analysis of behind-the-meter distributed energy resources to enable rapid planning and decision-making, thereby significantly boosting the rate the renewable energy deployment. Lastly, but equally importantly, this computation framework is open-sourced to facilitate transparency, flexibility, and wider collaboration opportunities within the worldwide energy modeling community.

cs.OH

Why So Many Published Sensitivity Analyses Are False. A Systematic Review of Sensitivity Analysis Practices

Sensitivity analysis (SA) has much to offer for a very large class of applications, such as model selection, calibration, optimization, quality assurance and many others. Sensitivity analysis offers crucial contextual information regarding a prediction by answering the question "Which uncertain input factors are responsible for the uncertainty in the prediction?" SA is distinct from uncertainty analysis (UA), which instead addresses the question "How uncertain is the prediction?" As we discuss in the present paper much confusion exists in the use of these terms. A proper uncertainty analysis of the output of a mathematical model needs to map what the model does when the input factors are left free to vary over their range of existence. A fortiori, this is true of a sensitivity analysis. Despite this, most UA and SA still explore the input space; moving along mono-dimensional corridors which leave the space of variation of the input factors mostly unscathed. We use results from a bibliometric analysis to show that many published SA fail the elementary requirement to properly explore the space of the input factors. The results, while discipline-dependent, point to a worrying lack of standards and of recognized good practices. The misuse of sensitivity analysis in mathematical modelling is at least as serious as the misuse of the p-test in statistical modelling. Mature methods have existed for about two decades to produce a defensible sensitivity analysis. We end by offering a rough guide for proper use of the methods.

stat.AP

Exploring Hoover and Perez's experimental designs using global sensitivity analysis

This paper investigates variable-selection procedures in regression that make use of global sensitivity analysis. The approach is combined with existing algorithms and it is applied to the time series regression designs proposed by Hoover and Perez. A comparison of an algorithm employing global sensitivity analysis and the (optimized) algorithm of Hoover and Perez shows that the former significantly improves the recovery rates of original specifications.

stat.CO