SearcharxivSearch

arXiv subjects

Sara Colando

Publications and source records attributed to Sara Colando.

3 recordsLinked to original sources

Analyzing Students' Statistics Writing Before and After the Emergence of Large Language Models

The ability to communicate statistical results to domain experts and stakeholders is an important goal of the undergraduate statistics and data science curriculum. However, as large language models (LLMs) have become more accessible, a major concern is that students are offloading important cognitive tasks to generative AI. Using a corpus of over 1,600 undergraduate students' data analysis reports from 2021 to 2025, we show how students' writing style and verb usage have become more similar to that of LLMs. This shift is most pronounced in the first and fifth quintiles of students' reports, which roughly map onto the introduction and conclusion sections, respectively. At the same time, we demonstrate that students' writing style has become more similar to that of statistics experts with the addition of LLMs. We end by discussing the implications of our findings for statistics and data science educators. In particular, we propose alternative modes of assessment that still emphasize statistical thinking, such as targeted writing assignments for structuring a report introduction.

stat.AP

Selecting ChIP-seq Normalization Methods from the Perspective of their Technical Conditions

Chromatin immunoprecipitation with high-throughput sequencing (ChIP-seq) provides insights into both the genomic location occupied by the protein of interest and the difference in DNA occupancy between experimental states. Given that ChIP-seq data is collected experimentally, an important step for determining regions with differential DNA occupancy between states is between-sample normalization. While between-sample normalization is crucial for downstream differential binding analysis, the technical conditions underlying between-sample normalization methods have yet to be examined for ChIP-seq. We identify three important technical conditions underlying ChIP-seq between-sample normalization methods: balanced differential DNA occupancy, equal total DNA occupancy, and equal background binding across states. To illustrate satisfying the selected normalization method's technical conditions for downstream differential binding analysis, we simulate ChIP-seq read count data where different combinations of the technical conditions are violated. We then externally verify our simulation results using experimental data. Based on our findings, we suggest that researchers use their understanding of the ChIP-seq experiment at hand to guide their choice of between-sample normalization method. Alternatively, researchers can use a high-confidence peakset, which is the intersection of the differentially bound peaksets obtained from using different between-sample normalization methods. In our two experimental analyses, roughly half of the called peaks were called as differentially bound for every normalization method. High-confidence peaks are less sensitive to choice of between-sample normalization method and could be a more robust basis for identifying genomic regions with differential DNA occupancy between experimental states when there is uncertainty about which technical conditions are satisfied.

q-bio.GN

Philosophy within Data Science Ethics Courses

There is wide agreement that ethical considerations are a valuable aspect of a data science curriculum, and to that end, many data science programs offer courses in data science ethics. There are not always, however, explicit connections between data science ethics and the centuries-old work on ethics within the discipline of philosophy. Here, we present a framework for bringing together key data science practices with ethics topics. The ethics topics were collated from sixteen data science ethics courses with public-facing syllabi and reading lists. We encourage individuals who are teaching data science ethics to engage with the philosophical literature and its connection to current data science practices, which are rife with potentially morally charged decision points.

stat.OT