SearcharxivSearch

arXiv subjects

Monnie McGee

Publications and source records attributed to Monnie McGee.

10 recordsLinked to original sources

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis. The framework is applied to explanations generated by 15 current-generation LLMs responding to 90 questions drawn from four statistics examinations spanning high school, undergraduate, and graduate levels. Accuracy varied substantially across models, ranging from 55\% to 78\%. In contrast, structural topic modeling revealed a common conceptual organization of statistical reasoning across all models, while lexical similarity analysis identified modest but consistent vendor-specific differences in explanatory style. Models developed by the same vendor (e.g. Anthropic, OpenAI) produced explanations that were slightly more similar than models from different vendors. These findings demonstrate that statistical reasoning in contemporary LLMs cannot be characterized by accuracy alone and illustrate how complementary analyses of response behavior and model-generated explanations provide a more comprehensive evaluation of statistical reasoning in generative AI.

cs.CL

Tree Estimation and Saddlepoint-Based Diagnostics for the Nested Dirichlet Distribution: Application to Compositional Behavioral Data

The Nested Dirichlet Distribution (NDD) provides a flexible alternative to the Dirichlet distribution for modeling compositional data, relaxing constraints on component variances and correlations through a hierarchical tree structure. While theoretically appealing, the NDD is underused in practice due to two main limitations: the need to predefine the tree structure and the lack of diagnostics for evaluating model fit. This paper addresses both issues. First, we introduce a data-driven, greedy tree-finding algorithm that identifies plausible NDD tree structures from observed data. Second, we propose novel diagnostic tools, including pseudo-residuals based on a saddlepoint approximation to the marginal distributions and a likelihood displacement measure to detect influential observations. These tools provide accurate and computationally tractable assessments of model fit, even when marginal distributions are analytically intractable. We demonstrate our approach through simulation studies and apply it to data from a Morris water maze experiment, where the goal is to detect differences in spatial learning strategies among cognitively impaired and unimpaired mice. Our methods yield interpretable structures and improved model evaluation in a realistic compositional setting. An accompanying R package is provided to support reproducibility and application to new datasets.

stat.ME

Layered Dirichlet Modeling to Assess the Changing Contributions of MLB Players as they Age

The productive career of a professional athlete is limited compared to the normal human lifespan. Most professional athletes have retired by age 40. The early retirement age is due to a combination of age-related performance and life considerations. While younger players typically are stronger and faster than their older teammates, older teammates add value to a team due to their experience and perspective. Indeed, the highest--paid major league baseball players are those over the age of 35. These players contribute intangibly to a team through mentorship of younger players; however, their peak athletic performance has likely passed. Given this, it is of interest to learn how more mature players contribute to a team in measurable ways. We examine the distribution of plate appearance outcomes from three different age groups as compositional data, using Layered Dirichlet Modeling (LDM). We develop a hypothesis testing framework to compare the average proportions of outcomes for each component among 3 of more groups. LDM can not only determine evidence for differences among populations, but also pinpoint within which component the largest changes are likely to occur. This framework can determine where players can be of most use as they age.

stat.ME

Generative AI Takes a Statistics Exam: A Comparison of Performance between ChatGPT3.5, ChatGPT4, and ChatGPT4o-mini

Many believe that use of generative AI as a private tutor has the potential to shrink access and achievement gaps between students and schools with abundant resources versus those with fewer resources. Shrinking the gap is possible only if paid and free versions of the platforms perform with the same accuracy. In this experiment, we investigate the performance of GPT versions 3.5, 4.0, and 4o-mini on the same 16-question statistics exam given to a class of first-year graduate students. While we do not advocate using any generative AI platform to complete an exam, the use of exam questions allows us to explore aspects of ChatGPT's responses to typical questions that students might encounter in a statistics course. Results on accuracy indicate that GPT 3.5 would fail the exam, GPT4 would perform well, and GPT4o-mini would perform somewhere in between. While we acknowledge the existence of other Generative AI/LLMs, our discussion concerns only ChatGPT because it is the most widely used platform on college campuses at this time. We further investigate differences among the AI platforms in the answers for each problem using methods developed for text analytics, such as reading level evaluation and topic modeling. Results indicate that GPT3.5 and 4o-mini have characteristics that are more similar than either of them have with GPT4.

stat.OT

Maximizing Diver Score by Examining Discrepancies in Diver Competency and Judges' Marks

Central to diving competitions is the diver's ``dive list'', which is the list of dives an athlete will perform during a competition. Creating a dive list that contains enough difficulty to be competitive yet not beyond the capability of the diver is an important consideration in diving. In this work, we examine the discrepancy between a diver's ability and judges' scores in springboard diving meets with the purpose of discovering biases in scoring that might aid a diver in completing a dive list. As a measure of the ability of a diver, we calculate a mean score for all dives and all meets in which the diver has participated. We call this mean score a diver's competency score. We use the difference between judges' scores within a given meet and the diver's competency to define a discrepancy: the difference between a judge's estimation of a diver's ability and their true ability. The notions of competency and discrepancy are applied to a data set, gathered from divemeets.com for high-school one meter diving competitions in the US from 2017 to 2022.

stat.AP

GRB Redshift Classifier to Follow-up High-Redshift GRBs Using Supervised Machine Learning

Gamma-ray bursts (GRBs) are intense, short-lived bursts of gamma-ray radiation observed up to a high redshift ($z \sim 10$) due to their luminosities. Thus, they can serve as cosmological tools to probe the early Universe. However, we need a large sample of high$-z$ GRBs, currently limited due to the difficulty in securing time at the large aperture Telescopes. Thus, it is painstaking to determine quickly whether a GRB is high$z$ or low$-z$, which hampers the possibility of performing rapid follow-up observations. Previous efforts to distinguish between high$-$ and low$-z$ GRBs using GRB properties and machine learning (ML) have resulted in limited sensitivity. In this study, we aim to improve this classification by employing an ensemble ML method on 251 GRBs with measured redshifts and plateaus observed by the Neil Gehrels Swift Observatory. Incorporating the plateau phase with the prompt emission, we have employed an ensemble of classification methods to enhance the sensitivity unprecedentedly. Additionally, we investigate the effectiveness of various classification methods using different redshift thresholds, $z_{threshold}$=$z_t$ at $z_{t}=$ 2.0, 2.5, 3.0, and 3.5. We achieve a sensitivity of 87\% and 89\% with a balanced sampling for both $z_{t}=3.0$ and $z_{t}=3.5$, respectively, representing a 9\% and 11\% increase in the sensitivity over Random Forest used alone. Overall, the best results are at $z_{t} = 3.5$, where the difference between the sensitivity of the training set and the test set is the smallest. This enhancement of the proposed method paves the way for new and intriguing follow-up observations of high$-z$ GRBs.

astro-ph.HE

Equity in the Use of ChatGPT for the Classroom: A Comparison of the Accuracy and Precision of ChatGPT 3.5 vs. ChatGPT4 with Respect to Statistics and Data Science Exams

A college education historically has been seen as method of moving upward with regards to income brackets and social status. Indeed, many colleges recognize this connection and seek to enroll talented low income students. While these students might have their education, books, room, and board paid; there are other items that they might be expected to use that are not part of most college scholarship packages. One of those items that has recently surfaced is access to generative AI platforms. The most popular of these platforms is ChatGPT, and it has a paid version (ChatGPT4) and a free version (ChatGPT3.5). We seek to explore differences in the free and paid versions in the context of homework questions and data analyses as might be seen in a typical introductory statistics course. We determine the extent to which students who cannot afford newer and faster versions of generative AI programs would be disadvantaged in terms of writing such projects and learning these methods.

stat.OT

Analysis of Compositional Data with Positive Correlations among Components using a Nested Dirichlet Distribution with Application to a Morris Water Maze Experiment

In a typical Morris water maze experiment, a mouse is placed in a circular water tank and allowed to swim freely until it finds a platform, triggering a route of escape from the tank. For reference purposes, the tank is divided into four quadrants: the target quadrant where the trigger to escape resides, the opposite quadrant to the target, and two adjacent quadrants. Several response variables can be measured: the amount of time that a mouse spends in different quadrants of the water tank, the number of times the mouse crosses from one quadrant to another, or how quickly a mouse triggers an escape from the tank. When considering time within each quadrant, it is hypothesized that normal mice will spend smaller amounts of time in quadrants that do not contain the escape route, while mice with an acquired or induced mental deficiency will spend equal time in all quadrants of the tank. Clearly, proportion of time in the quadrants must sum to one and are therefore statistically dependent; however, most analyses of data from this experiment treat time in quadrants as statistically independent. A recent paper introduced a hypothesis testing method that involves fitting such data to a Dirichlet distribution. While an improvement over studies that ignore the compositional structure of the data, we show that methodology is flawed. We introduce a two-sample test to detect differences in proportion of components for two independent groups where both groups are from either a Dirichlet or nested Dirichlet distribution. This new test is used to reanalyze the data from a previous study and come to a different conclusion.

stat.ME

A Simple Correction Procedure for High-Dimensional Generalized Linear Models with Measurement Error

We consider high-dimensional generalized linear models when the covariates are contaminated by measurement error. Estimates from errors-in-variables regression models are well-known to be biased in traditional low-dimensional settings if the error is unincorporated. Such models have recently become of interest when regularizing penalties are added to the estimation procedure. Unfortunately, correcting for the mismeasurements can add undue computational difficulties onto the optimization, which a new tool set for practitioners to successfully use the models. We investigate a general procedure that utilizes the recently proposed Imputation-Regularized Optimization algorithm for high-dimensional errors-in-variables models, which we implement for continuous, binary, and count response type. Crucially, our method allows for off-the-shelf linear regression methods to be employed in the presence of contaminated covariates. We apply our correction to gene microarray data, and illustrate that it results in a great reduction in the number of false positives whilst still retaining most true positives.

stat.CO

Bayesian Regularization of Gaussian Graphical Models with Measurement Error

We consider a framework for determining and estimating the conditional pairwise relationships of variables when the observed samples are contaminated with measurement error in high dimensional settings. Assuming the true underlying variables follow a multivariate Gaussian distribution, if no measurement error is present, this problem is often solved by estimating the precision matrix under sparsity constraints. However, when measurement error is present, not correcting for it leads to inconsistent estimates of the precision matrix and poor identification of relationships. We propose a new Bayesian methodology to correct for the measurement error from the observed samples. This Bayesian procedure utilizes a recent variant of the spike-and-slab Lasso to obtain a point estimate of the precision matrix, and corrects for the contamination via the recently proposed Imputation-Regularization Optimization procedure designed for missing data. Our method is shown to perform better than the naive method that ignores measurement error in both identification and estimation accuracy. To show the utility of the method, we apply the new method to establish a conditional gene network from a microarray dataset.

stat.ME