SearcharxivSearch

arXiv subjects

Sharif Ahmed

Publications and source records attributed to Sharif Ahmed.

11 recordsLinked to original sources

Uncovering Scientific Software Sustainability through Community Engagement and Software Quality Metrics

Scientific open-source software (Sci-OSS) projects are critical for advancing research, yet sustaining these projects long-term remains a major challenge. This paper explores the sustainability of Sci-OSS hosted on GitHub, focusing on two factors drawn from stewardship organizations: community engagement and software quality. We map sustainability to repository metrics from the literature and mined data from ten prominent Sci-OSS projects. A multimodal analysis of these projects led us to a novel visualization technique, providing a robust way to display both current and evolving software metrics over time, replacing multiple traditional visualizations with one. Additionally, our statistical analysis shows that even similar-domain projects sustain themselves differently. Natural language analysis supports claims from the literature, highlighting that project-specific feedback plays a key role in maintaining software quality. Our visualization and analysis methods offer researchers, funders, and developers key insights into long-term software sustainability.

cs.SE

Characterizing the Usefulness of Code Review Comments in Scientific Software for Software Quality and Scientific Rigor

Context: Innovation thrives on scientific software, with useful code review feedback enhancing its correctness and impact. However, unlike general-purpose commercial and open-source software, the usefulness of code review feedback (CR comment) in scientific software remains largely unstudied. Objective: This paper aims to characterize the usefulness of CR comment in scientific opens ource software (Sci-OSS), leveraging existing research on useful CR comment. Method: To achieve this objective, we mine successful Sci-OSS from GitHub, analyze their CR comments with usefulness related features, and compare the findings from prior research on general-purpose commercial and open-source CR comments. Results: The investigation on the usefulness of CR comments in SciOSS confirms many characteristics that prior research identified in general-purpose software. For example, subjective or negative CR comments remain not useful for the Sci-OSS. We also find CR comments which receive negative emoji reactions have a very small correlation with not useful comments, whereas the positive emojis show mixed correlations. Importantly, 6-33% CR comments in Sci-OSS are not useful in our mined repositories. Conclusions: Our investigation into Sci-OSS extends findings from CR comments' usefulness research on general-purpose software, benefiting developers, scientists, and researchers in the Sci-OSS community.

cs.SE

Operando study of the evolution of peritectic structures in metal solidification by quasi-simultaneous synchrotron X-ray diffraction and tomography

Using quasi-simultaneous synchrotron X-ray diffraction and tomography techniques, we have studied in-situ and in real-time the nucleation and co-growth dynamics of the peritectic structures in an Al-Mn alloy during solidification. We collected ~30 TB 4D datasets which allow us to elucidate the phases' co-growth dynamics and their spatial, crystallographic and compositional relationship. The primary Al4Mn hexagonal prisms nucleate and grow with high kinetic anisotropy -70 times faster in the axial direction than the radial direction. In all cases, a ~5 um Mn-rich diffusion layer forms at the liquid-solid interface, creating a sharp local solute gradient that governs subsequent phase transformation. The peritectic Al6Mn phases nucleate epitaxially within this diffusion zone, initially forming a thin shell surrounding the Al4Mn with an orientation relationship of {10-10}HCP // {110}O, [0001]HCP // [001]O. Such ~5 um Mn-rich diffusion layers also cause solute depletion at the liquid side of the liquid-solid interface, limiting further epitaxial phase growth, but prompting phase re-nucleation and branching at crystal edges, resulting tetragonal prism structures that no longer follow the initial orientation relationship. The anisotropic diffusion also led to the formation of core defects at the centre of both phases. Furthermore, increasing cooling rate from 0.17 to 20 °C/s can disrupt the stability of the solute diffusion zone, effectively suppressing the formation of the core defects and forcing a transition from faceted to non-faceted morphologies. Our work establishes a new theoretical framework for how to tailor and control the peritectic structures in metallic alloys through solidification processes.

cond-mat.mtrl-sci

The Moving Beam Diffraction Geometry: the DIAD Application of a Diffraction Scanning-Probe

Understanding the interactions between microstructure, strain, phase, and material behavior is crucial in many scientific fields. However, quantifying these correlations is challenging, as it requires the use of multiple instruments and techniques, often separated by space and time. The Dual Imaging And Diffraction (DIAD) beamline at Diamond is designed to address this challenge. DIAD allows its users to visualize internal structures, identify compositional/phase changes, and measure strain. DIAD provides two independent beams combined at one sample position, allowing quasi-simultaneous X-ray Computed Tomography and X-ray Powder Diffraction. A unique functionality of the DIAD configuration is the ability to perform image-guided diffraction, where the micron-sized diffraction beam is scanned over the complete area of the imaging field of view without moving the specimen. This moving beam diffraction geometry enables the study of fast-evolving and motion-susceptible processes and samples. Here, we discuss the novel moving beam diffraction geometry presenting the latest findings on the reliability of both geometry calibration and data reduction routines used. Our measures confirm diffraction is most sensitive to the moving geometry for the detector position downstream normal to the incident beam. The observed data confirm that the motion of the KB mirror coupled with a fixed aperture slit results in a rigid translation of the beam probe, without affecting the angle of the incident beam path to the sample. Our measures demonstrate a nearest-neighbour calibration can achieve the same accuracy as a self-calibrated geometry when the distance between calibrated and probed sample region is smaller or equal to the beam spot size. We show the absolute error of the moving beam diffraction geometry remains below 0.0001, which is the accuracy we observe for the beamline with stable beam operation.

cond-mat.mtrl-sci

Analyzing Social Media Engagement of Computer Science Conferences

Context: X, formerly known as Twitter, is one of the largest social media platforms and has been widely used for communication during research conferences. While previous studies have examined how users engage with X during these events, limited research has focused on analyzing the content posted by computer science conferences. Objective: This study investigates how conferences from different areas of computer science perform on social media by analyzing their activity, follower engagement, and the content posted on X. Method: We collect posts from 22 computer science conferences and conduct statistical experiments to identify variations in content. Additionally, we perform a manual analysis of the top five posts for each engagement metric. Results: Our findings indicate statistically significant differences in category, sentiment, and post length across computer science conference posts. Among all engagement metrics, likes were the most common way users interacted with conference content. Conclusion: This study provides insights into the social media presence of computer science conferences, highlighting key differences in content, sentiment, and engagement patterns across different venues.

cs.SI

Hold On! Is My Feedback Useful? Evaluating the Usefulness of Code Review Comments

Context: In collaborative software development, the peer code review process proves beneficial only when the reviewers provide useful comments. Objective: This paper investigates the usefulness of Code Review Comments (CR comments) through textual feature-based and featureless approaches. Method: We select three available datasets from both open-source and commercial projects. Additionally, we introduce new features from software and non-software domains. Moreover, we experiment with the presence of jargon, voice, and codes in CR comments and classify the usefulness of CR comments through featurization, bag-of-words, and transfer learning techniques. Results: Our models outperform the baseline by achieving state-of-the-art performance. Furthermore, the result demonstrates that the commercial gigantic LLM, GPT-4o, or non-commercial naive featureless approach, Bag-of-Word with TF-IDF, is more effective for predicting the usefulness of CR comments. Conclusion: The significant improvement in predicting usefulness solely from CR comments escalates research on this task. Our analyses portray the similarities and differences of domains, projects, datasets, models, and features for predicting the usefulness of CR comments.

cs.SE

Three-Dimensional, Multimodal Synchrotron Data for Machine Learning Applications

Machine learning techniques are being increasingly applied in medical and physical sciences across a variety of imaging modalities; however, an important issue when developing these tools is the availability of good quality training data. Here we present a unique, multimodal synchrotron dataset of a bespoke zinc-doped Zeolite 13X sample that can be used to develop advanced deep learning and data fusion pipelines. Multi-resolution micro X-ray computed tomography was performed on a zinc-doped Zeolite 13X fragment to characterise its pores and features, before spatially resolved X-ray diffraction computed tomography was carried out to characterise the homogeneous distribution of sodium and zinc phases. Zinc absorption was controlled to create a simple, spatially isolated, two-phase material. Both raw and processed data is available as a series of Zenodo entries. Altogether we present a spatially resolved, three-dimensional, multimodal, multi-resolution dataset that can be used for the development of machine learning techniques. Such techniques include development of super-resolution, multimodal data fusion, and 3D reconstruction algorithm development.

cs.LG

Decade-long Utilization Patterns of ICSE Technical Papers and Associated Artifacts

Context: Annually, ICSE acknowledges a range of papers, a subset of which are paired with research artifacts such as source code, datasets, and supplementary materials, adhering to the Open Science Policy. However, no prior systematic inquiry dives into gauging the influence of ICSE papers using artifact attributes. Objective: We explore the mutual impact between artifacts and their associated papers presented at ICSE over ten years. Method: We collect data on usage attributes from papers and their artifacts, conduct a statistical assessment to identify differences, and analyze the top five papers in each attribute category. Results: There is a significant difference between paper citations and the usage of associated artifacts. While statistical analyses show no notable difference between paper citations and GitHub stars, variations exist in views and/or downloads of papers and artifacts. Conclusion: We provide a thorough overview of ICSE's accepted papers from the last decade, emphasizing the intricate relationship between research papers and their artifacts. To enhance the assessment of artifact influence in software research, we recommend considering key attributes that may be present in one platform but not in another.

cs.SE

Understanding Emojis :) in Useful Code Review Comments

Emojis and emoticons serve as non-verbal cues and are increasingly prevalent across various platforms, including Modern Code Review. These cues often carry emotive or instructive weight for developers. Our study dives into the utility of Code Review comments (CR comments) by scrutinizing the sentiments and semantics conveyed by emojis within these comments. To assess the usefulness of CR comments, we augment traditional 'textual' features and pre-trained embeddings with 'emoji-specific' features and pre-trained embeddings. To fortify our inquiry, we expand an existing dataset with emoji annotations, guided by existing research on GitHub emoji usage, and re-evaluate the CR comments accordingly. Our models, which incorporate textual and emoji-based sentiment features and semantic understandings of emojis, substantially outperform baseline metrics. The often-overlooked emoji elements in CR comments emerge as key indicators of usefulness, suggesting that these symbols carry significant weight.

cs.SE

Exploring the Advances in Identifying Useful Code Review Comments

Effective peer code review in collaborative software development necessitates useful reviewer comments and supportive automated tools. Code review comments are a central component of the Modern Code Review process in the industry and open-source development. Therefore, it is important to ensure these comments serve their purposes. This paper reflects the evolution of research on the usefulness of code review comments. It examines papers that define the usefulness of code review comments, mine and annotate datasets, study developers' perceptions, analyze factors from different aspects, and use machine learning classifiers to automatically predict the usefulness of code review comments. Finally, it discusses the open problems and challenges in recognizing useful code review comments for future research.

cs.SE

Automatic Transformation of Natural to Unified Modeling Language: A Systematic Review

Context: Processing Software Requirement Specifications (SRS) manually takes a much longer time for requirement analysts in software engineering. Researchers have been working on making an automatic approach to ease this task. Most of the existing approaches require some intervention from an analyst or are challenging to use. Some automatic and semi-automatic approaches were developed based on heuristic rules or machine learning algorithms. However, there are various constraints to the existing approaches of UML generation, such as restriction on ambiguity, length or structure, anaphora, incompleteness, atomicity of input text, requirements of domain ontology, etc. Objective: This study aims to better understand the effectiveness of existing systems and provide a conceptual framework with further improvement guidelines. Method: We performed a systematic literature review (SLR). We conducted our study selection into two phases and selected 70 papers. We conducted quantitative and qualitative analyses by manually extracting information, cross-checking, and validating our findings. Result: We described the existing approaches and revealed the issues observed in these works. We identified and clustered both the limitations and benefits of selected articles. Conclusion: This research upholds the necessity of a common dataset and evaluation framework to extend the research consistently. It also describes the significance of natural language processing obstacles researchers face. In addition, it creates a path forward for future research.

cs.SE