SearcharxivSearch

arXiv subjects

Michael Zentner

Publications and source records attributed to Michael Zentner.

5 recordsLinked to original sources

nanoHUB services for FAIR simulations and data: ResultsDB and Sim2Ls

nanoHUB is an open cyber platform for online simulation, data, and education that seeks to make scientific software and associated data widely available and useful. This paper describes recent developments in our simulation infrastructure to address modern data needs. nanoHUB's Sim2Ls (pronounced sim tools) make simulation, modeling, and data workflows discoverable and accessible to all users for cloud computing using standard APIs. In addition, published tools are findable (with digital object identifiers), reusable (via documented requirements and services), and reproducible via containerization. In addition, all Sim2L runs are automatically cached, and their results indexed into a global and queryable database (ResultsDB). We believe this infrastructure significantly lowers the barriers towards making simulation/data workflows and their data findable, accessible, interoperable, and reusable (FAIR). This frictionless access to simulations and data enables researchers, instructors, and students to focus on the application of these products to advance their fields.

cond-mat.mtrl-sci

The Science Gateway Community Institute's Consulting Services Program: Lessons for Research Software Engineering Organizations

The Science Gateways Community Institute (SGCI) is an NSF Software Infrastructure for Sustained Innovation (S2I2) funded project that leads and supports the science gateway community. Major activities for SGCI include a) sustainability training, including the Focus Week week-long course designed to help science gateway operators develop sustainability plans, and the Jumpstart virtual short-course; b) usability and user experience consulting; c) a community catalog of science gateways and science gateway software; d) workforce development activities, including a coding institute for students, internship opportunities, and hackathons; e) an annual conference; and f) in-depth technical support for client gateway projects. The goals of SGCI's Embedded Technical Support component are to help the institute's clients to create new science gateways or to significantly enhance existing science gateways. Examples of the latter include helping to implement major new capabilities and to implement significant usability improvements suggested by SGCI's usability consultants. The Embedded Technical Support component was managed by Indiana University and involved research software engineers at San Diego Supercomputer Center, Texas Advanced Computing Center, Indiana University, and Purdue University (through 2019). Since 2016, the component has involved 20 research software engineers as consultants and has conducted 59 client consultations. This short paper provides a summary of lessons learned from the Embedded Technical Support program that may be useful for the research software engineering community.

cs.SE

Toward Interoperable Cyberinfrastructure: Common Descriptions for Computational Resources and Applications

The user-facing components of the Cyberinfrastructure (CI) ecosystem, science gateways and scientific workflow systems, share a common need of interfacing with physical resources (storage systems and execution environments) to manage data and execute codes (applications). However, there is no uniform, platform-independent way to describe either the resources or the applications. To address this, we propose uniform semantics for describing resources and applications that will be relevant to a diverse set of stakeholders. We sketch a solution to the problem of a common description and catalog of resources: we describe an approach to implementing a resource registry for use by the community and discuss potential approaches to some long-term challenges. We conclude by looking ahead to the application description language.

cs.DC

A Statistical Approach to Increase Classification Accuracy in Supervised Learning Algorithms

Probabilistic mixture models have been widely used for different machine learning and pattern recognition tasks such as clustering, dimensionality reduction, and classification. In this paper, we focus on trying to solve the most common challenges related to supervised learning algorithms by using mixture probability distribution functions. With this modeling strategy, we identify sub-labels and generate synthetic data in order to reach better classification accuracy. It means we focus on increasing the training data synthetically to increase the classification accuracy.

cs.LG