SearcharxivSearch

arXiv subjects

Oliver Gerstl

Publications and source records attributed to Oliver Gerstl.

3 recordsLinked to original sources

Inferring the Shape of Data Frames in R Programs using Abstract Interpretation

Data frames are a fundamental data structure in many data analysis tasks and are widely used in programming languages like R. Due to their omnipresence in data analysis, there are many functions that operate on their shape and content, for example, to clean and transform study data. However, languages like R do not offer static guarantees on data frames making it difficult to reason about their shape at a specific point in the program. In this paper, we present a novel static analysis to infer the shape of data frames in R programs using abstract interpretation by tracking the ensured and potential column names, as well as the potential number of columns and rows. For this, we use a reduced product domain and define abstract semantics for the most commonly used data frame operations, such as mutating, filtering, and subsetting. We evaluate the correctness and accuracy of our analysis on a selection of 78 executable real-world R scripts achieving empirical evidence for soundness by never under-approximating the data frame shape. Additionally, we demonstrate the ability of our analysis to infer the shape of data frames on a large dataset of 33,314 real-world R scripts by inferring concrete shape constraints for 42.1 % and exact shapes for 0.9 % of the data frame operations, improving to 58.7 % and 4.2 % if all datasets read in these scripts are available to our analysis. Using the inferred data frame shapes, we identified 40 real-world R scripts containing potential invalid data frame accesses. This shows the potential of our analysis to significantly support researchers in using data frames in data analysis.

cs.SE

Towards Automatically Inferring Constraints to Identify Implicit Assumptions in Data Analysis

High-level languages such as R or Python are used frequently to analyze and visualize data in the form of scripts or notebooks. However, these artifacts suffer from reproducibility issues due to what we frame as implicit assumptions made by the authors. Such assumptions range from package versions and shapes of involved data tables, to manual and often undocumented setup steps. Within this work, we provide a unified, example-driven perspective on implicit assumptions in data analysis backed by an explorative proof-of-concept implementation. With this perspective, we propose the use of static analysis techniques to identify such assumptions and to make them explicit in the form of code constraints, focusing on the inclusion of data-analysis-specific issues. Such constraints can then be used to automatically transform these scripts into executable and reproducible artifacts, to check these assumptions at runtime, and to serve as documentation to support code reuse and comprehension.

cs.SE

Supporting the Comprehension of Data Analysis Scripts

A lot of research relies on data analysis scripts to process, clean, and visualize data. However, recent studies show that these scripts are often hard to comprehend and maintain, hindering reproducibility and reuse, accompanied by a lack of tool support for handling such scripts. In this work, we focus on the R programming language, addressing this problem by presenting flowR as an extension for the common data analysis IDEs Positron and VS Code. Alongside a previously presented static backward program slicer, flowR provides an overview of data analysis scripts, interactive graph visualizations, linting, and inline value annotations to support data analysts. FlowR incrementally analyzes R projects by intertwining interprocedural data- and control-flow analyses to build a comprehensive dataflow graph, incorporating R's dynamic and explorative features. Additionally, flowR offers a plugin system and interfaces, allowing the integration of further analyses, such as new linting rules or custom visualizations. Requiring an average of 576ms to calculate the full dataflow graph of real-world projects, this enables near real-time feedback. The demonstration video is available at https://youtu.be/hJzr-r-NmMg . For the full source code and extensive documentation, refer to https://github.com/flowr-analysis/flowr . To try the docker image, use `docker run --rm -it eagleoutice/flowr`.

cs.SE