arXiv · 1905.13363
DFS: A Dataset File System for Data Discovering Users
Abstract
Many research questions can be answered quickly and efficiently using data already collected for previous research. This practice is called secondary data analysis (SDA), and has gained popularity due to lower costs and improved research efficiency. In this paper we propose DFS, a file system to standardize the metadata representation of datasets, and DDU, a scalable architecture based on DFS for semi-automated metadata generation and data recommendation on the cloud. We discuss how DFS and DDU lays groundwork for automatic dataset aggregation, how it integrates with existing data wrangling and machine learning tools, and explores their implications on datasets stored in digital libraries.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yasith Jayawardana, Sampath Jayarathna. 2019-05-31. DFS: A Dataset File System for Data Discovering Users. https://doi.org/10.1109/jcdl.2019.00068
Cite the original work for its findings. Save a collection to share your selection of sources.