Searcharxiv⌕ Search

arXiv subjects

Rebeckah K. Fussell

Publications and source records attributed to Rebeckah K. Fussell.

4 recordsLinked to original sources

A Framework for Deductive Semantic Content Analysis at Scale in Science Education Using Text Embeddings

Qualitative content analysis of open-ended survey responses is a commonly used research method in science education. However, traditional coding approaches are often time-consuming and prone to inconsistency, especially when applied to large datasets. Existing solutions from Natural Language Processing such as supervised classifiers, topic modeling techniques, and generative large language models have limited applicability in analysis of open-ended survey responses, since they demand extensive labeled data, disrupt established qualitative workflows, and/or yield variable results. In this paper, we introduce a text embedding-based classification framework called Deductive Semantic Content Analysis (DeSCA) that requires only a handful of examples per category to run, is transparent and replicable, and fits well with standard qualitative workflows. When benchmarked against human analysis of a physics education survey consisting of 2899 open-ended responses, the method described by our framework achieves high agreement with expert human coders across ten embeddings models on a simulated exhaustive coding task, using approximately 1-2% of the total dataset for training. The method achieves lower agreement on a complete selective coding task; this performance, however, improves with fine-tuning of the text embedding model, which can be done with a small amount of additional data. We unpack these results in terms of the theoretical assumptions of text embeddings, and further demonstrate how embeddings can be used to audit previously-analyzed datasets for coding consistency. These findings demonstrate that text embedding-assisted coding can flexibly scale to thousands of responses without sacrificing interpretability, opening avenues for deductive qualitative analysis at scale.

cs.CL↗

Comparing large language models for supervised analysis of students' lab notes

Recent advancements in large language models (LLMs) hold significant promise in improving physics education research that uses machine learning. In this study, we compare the application of various models to perform large-scale analysis of written text grounded in a physics education research classification problem: identifying skills in students' typed lab notes through sentence-level labeling. Specifically, we use training data to fine-tune two different LLMs, BERT and LLaMA, and compare the performance of these models to both a traditional bag of words approach and a few-shot LLM (without fine-tuning).} We evaluate the models based on their resource use, performance metrics, and research outcomes when identifying skills in lab notes. We find that higher-resource models often, but not necessarily, perform better than lower-resource models. We also find that all models estimate similar trends in research outcomes, although the absolute values of the estimated measurements are not always within uncertainties of each other. We use the results to discuss relevant considerations for education researchers seeking to select a model type to use as a classifier.

physics.ed-ph↗

A method to assess trustworthiness of machine coding at scale

Physics education researchers are interested in using the tools of machine learning and natural language processing to make quantitative claims from natural language and text data, such as open-ended responses to survey questions. The aspiration is that this form of machine coding may be more efficient and consistent than human coding, allowing much larger and broader data sets to be analyzed than is practical with human coders. Existing work that uses these tools, however, does not investigate norms that allow for trustworthy quantitative claims without full reliance on cross-checking with human coding, which defeats the purpose of using these automated tools. Here we propose a four-part method for making such claims with supervised natural language processing: evaluating a trained model, calculating statistical uncertainty, calculating systematic uncertainty from the trained algorithm, and calculating systematic uncertainty from novel data sources. We provide evidence for this method using data from two distinct short response survey questions with two distinct coding schemes. We also provide a real-world example of using these practices to machine code a data set unseen by human coders. We offer recommendations to guide physics education researchers who may use machine-coding methods in the future.

physics.ed-ph↗

Instructing nontraditional physics labs: Toward responsiveness to student epistemic framing

Research on nontraditional laboratory (lab) activities in physics shows that students often expect to verify predetermined results, as takes place in traditional activities. This understanding of what is taking place, or epistemic framing, may impact their behaviors in the lab, either productively or unproductively. In this paper, we present an analysis of student epistemic framing in a nontraditional lab to understand how instructional context, specifically instructor behaviors, may shape student framing. We present video data from a lab section taught by an experienced teaching assistant (TA), with 19 students working in seven groups. We argue that student framing in this lab is evidenced by whether or not students articulate experimental predictions and by the extent to which they take up opportunities to construct knowledge (epistemic agency). We show that the TA's attempts to shift student frames generally succeed with respect to experimental predictions but are less successful with respect to epistemic agency. In part, we suggest, the success of the TA's attempts reflects whether and how they are responsive to students' current framing. This work offers evidence that instructors can shift students' frames in nontraditional labs, while also illuminating the complexities of both student framing and the role of the instructor in shifting that framing in this context.

physics.ed-ph↗