SearcharxivSearch

arXiv subjects

Kirstie Whitaker

Publications and source records attributed to Kirstie Whitaker.

3 recordsLinked to original sources

Professionalising Community Management Roles in Interdisciplinary Research Projects

In this article we discuss community management in interdisciplinary research teams, focusing on recognising and professionalising roles referred to here as the Research Community Managers (RCM). Drawing insights and examples from research and data science projects, we discuss how RCM roles address some of the research’s most pressing challenges, from promoting best practices for open research and reproducibility to engaging diverse stakeholders in community-led research and ensuring fair recognition for their contributions. We offer a Community Maturation Indicator and share examples of projects from The Alan Turing Institute, the UK's national institute for data science and Artificial Intelligence (AI), where institutionally supported RCM roles were established. With the aim to integrate RCM expertise in teams involved in data science and AI research, we provide an RCM Skills and Competencies Framework. We also propose a roadmap for professionalising RCM roles by improving recognition and rewards, potential career paths and organisational support structures. To systematically sustain and progress these roles, we recommend institutional investment in establishing RCM teams that are empowered to prioritise collaboration, transparency and community-based approaches in interdisciplinary projects, such as in data science and AI. As a team, RCMs are well placed to connect disparate teams, initiatives and resources across the organisation, building more resilient research communities that can achieve greater innovation, improved project outcomes and a strongly connected ecosystem, with impacts extending beyond their narrow contexts.

physics.soc-ph

The citation advantage of linking publications to research data

Efforts to make research results open and reproducible are increasingly reflected by journal policies encouraging or mandating authors to provide data availability statements. As a consequence of this, there has been a strong uptake of data availability statements in recent literature. Nevertheless, it is still unclear what proportion of these statements actually contain well-formed links to data, for example via a URL or permanent identifier, and if there is an added value in providing such links. We consider 531,889 journal articles published by PLOS and BMC, develop an automatic system for labelling their data availability statements according to four categories based on their content and the type of data availability they display, and finally analyze the citation advantage of different statement categories via regression. We find that, following mandated publisher policies, data availability statements become very common. In 2018 93.7% of 21,793 PLOS articles and 88.2% of 31,956 BMC articles had data availability statements. Data availability statements containing a link to data in a repository -- rather than being available on request or included as supporting information files -- are a fraction of the total. In 2017 and 2018, 20.8% of PLOS publications and 12.2% of BMC publications provided DAS containing a link to data in a repository. We also find an association between articles that include statements that link to data in a repository and up to 25.36% ($\pm$~1.07%) higher citation impact on average, using a citation prediction model. We discuss the potential implications of these results for authors (researchers) and journal publishers who make the effort of sharing their data in repositories. All our data and code are made available in order to reproduce and extend our results.

cs.DL

Design choices for productive, secure, data-intensive research at scale in the cloud

We present a policy and process framework for secure environments for productive data science research projects at scale, by combining prevailing data security threat and risk profiles into five sensitivity tiers, and, at each tier, specifying recommended policies for data classification, data ingress, software ingress, data egress, user access, user device control, and analysis environments. By presenting design patterns for security choices for each tier, and using software defined infrastructure so that a different, independent, secure research environment can be instantiated for each project appropriate to its classification, we hope to maximise researcher productivity and minimise risk, allowing research organisations to operate with confidence.

cs.CR