SearcharxivSearch

arXiv subjects

Anna-Lena Roth

Publications and source records attributed to Anna-Lena Roth.

3 recordsLinked to original sources

Performance Analysis in Parallel Programming Education: A Comparative Usability Study

Parallel programming curricula encompass not only the development of parallel code and algorithm design but also emphasize efficiency, optimization, and performance analysis. To equip students with the skills necessary for writing efficient parallel code using message passing with MPI, practical experience on HPC environments is essential. Performance analysis tools assist in identifying issues such as load imbalances or bottlenecks. Despite their use by experienced developers, these tools' complexity and required knowledge of cluster architectures, resource management, MPI, and common parallel issues hinder their educational integration. To address these barriers, we developed EduMPI, a learning support tool designed to simplify cluster usage and performance analysis for students. EduMPI offers an intuitive GUI that automates program execution on clusters and delivers near-real-time visualizations of MPI communication. This enables students to track process communication according to their physical placement within the cluster and detect performance problems interactively. This paper presents a user study comparing EduMPI with established professional performance analysis tools, demonstrating that EduMPI lowers entry barriers and fosters an intuitive understanding of parallel program performance, thereby enhancing its educational value.

cs.DC

An Empirical Analysis of High-Performance Computing Education in Germany

The growing importance of High-Performance Computing (HPC) requires the systematic integration of parallel programming and performance-oriented competencies into computational science curricula. Effective HPC education combines theoretical foundations with practical experience on real cluster infrastructures, enabling students to understand scalability, efficiency, and architectural differences between shared and distributed memory systems. However, cross-institutional evidence on how HPC education is implemented, and how curricula relate to locally available infrastructure, remains limited. We address this gap through a systematic empirical assessment of HPC education at 102 academic institutions in Germany. Based on module handbooks and course catalogs, we identified 178 HPC-related courses and evaluated their competency coverage and curricular placement. We additionally assessed local academic HPC cluster infrastructures with respect to availability, size, and documented accessibility for teaching. The results show that 67.6% of institutions offer at least one HPC-related course, but these offerings are predominantly elective modules at the master's level, with limited integration in bachelor's programs. Although 61.8% of institutions operate HPC clusters, only 23.0% explicitly document their availability for educational use, as infrastructures are mainly reserved for research. Statistical analysis indicates a significant association between restricted teaching access and reduced curricular emphasis on practical competencies such as resource management, cluster usage, parallel debugging, and performance analysis. Overall, the findings reveal a structural imbalance between theoretical instruction and the development of practical HPC competencies in German higher education.

cs.DC

Generated, Parallel, Scalable? A Study of Agentic AI-Generated Julia Code on Supercomputers

Julia is increasingly used in HPC as a single-language alternative to combining high-level scripting with low-level systems languages, but achieving scalable performance still requires expertise in parallel programming. LLMs are increasingly used for code generation and are advancing rapidly with each new version. Yet, existing studies focus on single-shot prompting rather than agentic settings, in which an LLM autonomously plans, generates, and refines code through tool use. Using an OpenCode-based agent extended with a Julia-documentation MCP server, we study agentic generation of parallel Julia code, focusing on task-based execution with Dagger$.$jl. We evaluate three LLMS, OpenAI GPT-5.5, Anthropic Claude Opus 4.7, and the open-weight Qwen3-Coder-Next, on three problems with distinct parallel structures: Pi approximation, tiled general matrix multiplication, and tiled Cholesky decomposition. The generated Dagger$.$jl implementations are compared against agent-generated Base$.$Threads and MPI$.$jl baselines, with shared-memory experiments scaling to 192 cores and distributed-memory experiments on two nodes. The agents reliably produce executable code for small inputs but fail at larger scales due to deadlocks, oversubscription, or out-of-memory errors, with the open-weight model affected most severely. The two commercial models scale comparably on Base$.$Threads and MPI$.$jl, while their Dagger$.$jl implementations expose recurring weaknesses in task dependencies, granularity, and scheduling. Agentic AI is promising for producing parallel Julia code, but generating robust, performance-aware implementations for large-scale HPC systems remains an open challenge.

cs.DC