SearcharxivSearch

arXiv subjects

Murat Keceli

Publications and source records attributed to Murat Keceli.

11 recordsLinked to original sources

Overcoming Orchestration Bottlenecks at Exascale: A Decentralized, Policy-Driven Approach for Sim-AI Ensembles

Scientific computing is increasingly shifting from monolithic applications to coupled simulation-AI workflows composed of highly heterogeneous tasks with diverse hardware, scale, and runtime requirements. As these workflows scale to leadership-class systems, the resulting extreme ensemble sizes and task variability can create orchestration bottlenecks. System-level schedulers are often configured for limited throughput, while workflow tools face scalability issues due to rigid control-plane topologies and static scheduling heuristics. We introduce EnsembleLauncher, a recursively hierarchical workflow orchestrator for exascale systems, featuring a fully decentralized control plane and a programmable scheduling policy interface. On the Aurora supercomputer, EnsembleLauncher successfully scales to the entire machine with up to eight million serial tasks, outperforming state-of-the-art tools by more than four times. Additionally, we implement a programmable scheduling interface and demonstrate a significant impact of scheduling policies on resource utilization for high-variance ensembles and active learning pipelines representative of modern coupled simulation-AI workflows.

cs.DC

ChemGraph-XANES: An Agentic Framework for XANES Simulation and Curation

Computational X-ray absorption near-edge structure (XANES) is widely used to interpret local coordination environments, oxidation states, and electronic structure, but large computational campaigns are often limited by workflow complexity. We present ChemGraph-XANES, a large language model (LLM)-based agentic framework that combines documentation-grounded parameter retrieval via retrieval-augmented generation (RAG), schema-constrained tool execution, deterministic FDMNES input generation, Parsl-backed execution, and provenance-aware data curation. Scripted and natural-language interfaces share a common scientific backend for structure handling, parameterization, execution, spectral extraction, and optional post-processing. We evaluate three workflow modes: documentation-grounded parameter propagation, structure-file-based execution, and composition-based execution from a chemistry-level request. Repeated trials yielded end-to-end completion in 10/10 composition-based runs, 10/10 structure-file-based runs, and 9/10 documentation-grounded RAG runs. In every RAG run, the energy-grid specification retrieved from the FDMNES manual was correctly propagated, with the single end-to-end failure occurring downstream during multi-structure handling. In a separate task-parallel demonstration, the framework retrieved 21 TiO$_2$ structures from the Materials Project and submitted one FDMNES calculation per structure. All calculations completed successfully, with Parsl distributing the independent tasks across the user-configured worker pool. Together, these results show that ChemGraph-XANES provides a constrained and reproducible orchestration layer for computational spectroscopy, supporting consistent execution of representative tasks, documentation-linked parameter selection, and task-parallel generation of structure-linked XANES collections.

cond-mat.mtrl-sci

From Atomistic Models to Machine Learning: Predictive Design of Nanocarbons under Extreme Conditions

The formation of technologically valuable nanocarbon structures under extreme conditions, such as those produced during high-explosive detonations, remains poorly understood but holds significant potential for the development of controlled synthesis pathways. While detonation shockwaves provide the HPHT environment required for nanodiamond formation, subsequent cooling and decompression dictate whether the diamond phase is preserved or transformed into other nanocarbon structures. Here, we employ GPU-accelerated ReaxFF simulations to investigate the graphitization and structural remodeling of detonation nanodiamond under nonlinear quench and pressure-release conditions. We further investigate how the initial nanodiamond morphology influences the resulting transformation products. Evolution of nanostructure, allotrope, carbon hybridization, and ring statistics are tracked. Rapid cooling combined with slow decompression optimizes cubic diamond retention, whereas slow cooling with rapid pressure release promotes surface-to-core graphitization, producing concentric sp2 layers and hollowed inner shells. Octahedral nanodiamonds evolve into carbon nano-onions, initially forming bucky diamonds that progressively transform into full sp2 structures, while hexagonal prisms preferentially form parallel-stacked graphite layers resembling carbon dots. Lonsdaleite emerges as an interfacial phase, suggesting potential reversibility in the shock-induced graphite-to-diamond transformation pathway transformation route. To extend predictive capabilities, we trained MLP regressors on over 10^5 node-hours of simulations. The model reliably predicts the number of graphitized layers from T-P trajectories with R^2 exceeding 0.90. Collectively, morphological control combined with optimized quench-decompression conditions promote the selective synthesis of nanocarbon allotropes.

cond-mat.mtrl-sci

An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc

While LLMs have accelerated scientific code generation, comprehensively evaluating generated code remains challenging. Many benchmarks emphasize functional correctness or task completion, which is insufficient for code built on production HPC libraries, where solver selection, API conventions, memory management, parallel awareness, and performance also matter. We introduce PETSCAgent-Bench, a multidimensional benchmark and agent-based framework for assessing whether AI-generated scientific code uses a production HPC library as an expert would. A tool-augmented evaluator compiles, executes, and measures code and combines deterministic checks with LLM-based assessments in a 14-evaluator pipeline spanning five categories: correctness, performance, code quality, algorithmic appropriateness, and library-specific conventions. A2A and MCP enable black-box evaluation of compatible coding agents. Across realistic PETSc problems, frontier models generate readable, well-structured code but struggle with correctness on challenging problems and with library-specific conventions even when code compiles and runs---limitations that conventional pass/fail evaluation does not capture.

cs.AI

Thermal Conductivity Limits of MoS$_2$ and MoSe$_2$: Revisiting High-Order Anharmonic Lattice Dynamics with Machine Learning Potentials

Group-VI transition metal dichalcogenides (TMDs), MoS$_2$ and MoSe$_2$, have emerged as prototypical low-dimensional systems with distinctive phononic and electronic properties, making them attractive for applications in nanoelectronics, optoelectronics, and thermoelectrics. Yet, their reported lattice thermal conductivities ($\kappa$) remain highly inconsistent, with experimental values and theoretical predictions differing by more than an order of magnitude. These discrepancies stem from uncertainties in measurement techniques, variations in computational protocols, and ambiguities in the treatment of higher-order anharmonic processes. In this study, we critically review these inconsistencies, first by mapping the spread of experimental and modeling results, and then by identifying the methodological origins of divergence. To this end, we bridge first-principles calculations, molecular dynamics simulations, and state-of-the-art machine learning force fields (MLFFs) including recently developed foundation models. %MACE-OMAT-0, UMA, and NEP89. We train and benchmark GAP, MACE, NEP, and \textsc{HIPHIVE} against density functional theory (DFT) and rigorously evaluate the impact of third- and fourth-order phonon scattering processes on $\kappa$. The computational efficiency of MLFFs enables us to extend convergence tests beyond conventional limits and to validate predictions through homogeneous nonequilibrium molecular dynamics as well. Our analysis demonstrates that, contrary to some recent claims, fully converged four-phonon processes contribute negligibly to the intrinsic thermal conductivity of both MoS$_2$ and MoSe$_2$. These findings not only refine the intrinsic transport limits of 2D TMDs but also establish MLFF-based approaches as a robust and scalable framework for predictive modeling of phonon-mediated thermal transport in low-dimensional materials.

cond-mat.mtrl-sci

Red Teaming for Generative AI, Report on a Copyright-Focused Exercise Completed in an Academic Medical Center

Background: Generative artificial intelligence (AI) deployment in academic medical settings raises copyright compliance concerns. Dana-Farber Cancer Institute implemented GPT4DFCI, an internal generative AI tool utilizing OpenAI models, that is approved for enterprise use in research and operations. Given (1) the exceptionally broad adoption of the tool in our organization, (2) our research mission, and (3) the shared responsibility model required to benefit from Customer Copyright Commitment in Azure OpenAI Service products, we deemed rigorous copyright compliance testing necessary. Case Description: We conducted a structured red teaming exercise in Nov. 2024, with 42 participants from academic, industry, and government institutions. Four teams attempted to extract copyrighted content from GPT4DFCI across four domains: literary works, news articles, scientific publications, and access-restricted clinical notes. Teams successfully extracted verbatim book dedications and near-exact passages through various strategies. News article extraction failed despite jailbreak attempts. Scientific article reproduction yielded only high-level summaries. Clinical note testing revealed appropriate privacy safeguards. Discussion: The successful extraction of literary content indicates potential copyrighted material presence in training data, necessitating inference-time filtering. Differential success rates across content types suggest varying protective mechanisms. The event led to implementation of a copyright-specific meta-prompt in GPT4DFCI; this mitigation has been in production since Jan. 2025. Conclusion: Systematic red teaming revealed specific vulnerabilities in generative AI copyright compliance, leading to concrete mitigation strategies. Academic medical institutions deploying generative AI should implement continuous testing protocols to ensure legal and ethical compliance.

cs.CY

AI Assistants to Enhance and Exploit the PETSc Knowledge Base

Generative AI, especially through large language models (LLMs), is transforming how technical knowledge can be accessed, reused, and extended. PETSc, a widely used numerical library for high-performance scientific computing, has accumulated a rich but fragmented knowledge base over its three decades of development, spanning source code, documentation, mailing lists, GitLab issues, Discord conversations, technical papers, and more. Much of this knowledge remains informal and inaccessible to users and new developers. To activate and utilize this knowledge base more effectively, the PETSc team has begun building an LLM-powered system that combines PETSc content with custom LLM tools -- including retrieval-augmented generation (RAG), reranking algorithms, and chatbots -- to assist users, support developers, and propose updates to formal documentation. This paper presents initial experiences designing and evaluating these tools, focusing on system architecture, using RAG and reranking for PETSc-specific information, evaluation methodologies for various LLMs and embedding models, and user interface design. Leveraging the Argonne Leadership Computing Facility resources, we analyze how LLM responses can enhance the development and use of numerical software, with an initial focus on scalable Krylov solvers. Our goal is to establish an extensible framework for knowledge-centered AI in scientific software, enabling scalable support, enriched documentation, and enhanced workflows for research and development. We conclude by outlining directions for expanding this system into a robust, evolving platform that advances software ecosystems to accelerate scientific discovery.

cs.AI

EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants

Recent advancements have positioned AI, and particularly Large Language Models (LLMs), as transformative tools for scientific research, capable of addressing complex tasks that require reasoning, problem-solving, and decision-making. Their exceptional capabilities suggest their potential as scientific research assistants but also highlight the need for holistic, rigorous, and domain-specific evaluation to assess effectiveness in real-world scientific applications. This paper describes a multifaceted methodology for Evaluating AI models as scientific Research Assistants (EAIRA) developed at Argonne National Laboratory. This methodology incorporates four primary classes of evaluations. 1) Multiple Choice Questions to assess factual recall; 2) Open Response to evaluate advanced reasoning and problem-solving skills; 3) Lab-Style Experiments involving detailed analysis of capabilities as research assistants in controlled environments; and 4) Field-Style Experiments to capture researcher-LLM interactions at scale in a wide range of scientific domains and applications. These complementary methods enable a comprehensive analysis of LLM strengths and weaknesses with respect to their scientific knowledge, reasoning abilities, and adaptability. Recognizing the rapid pace of LLM advancements, we designed the methodology to evolve and adapt so as to ensure its continued relevance and applicability. This paper describes the methodology state at the end of February 2025. Although developed within a subset of scientific domains, the methodology is designed to be generalizable to a wide range of scientific domains.

cs.AI

Toward an Automated HPC Pipeline for Processing Large Scale Electron Microscopy Data

We present a fully modular and scalable software pipeline for processing electron microscope (EM) images of brain slices into 3D visualization of individual neurons and demonstrate an end-to-end segmentation of a large EM volume using a supercomputer. Our pipeline scales multiple packages used by the EM community with minimal changes to the original source codes. We tested each step of the pipeline individually, on a workstation, a cluster, and a supercomputer. Furthermore, we can compose workflows from these operations using a Balsam database that can be triggered during the data acquisition or with the use of different front ends and control the granularity of the pipeline execution. We describe the implementation of our pipeline and modifications required to integrate and scale up existing codes. The modular nature of our environment enables diverse research groups to contribute to the pipeline without disrupting the workflow, i.e. new individual codes can be easily integrated for each step on the pipeline.

cs.DC

Scaling Distributed Training of Flood-Filling Networks on HPC Infrastructure for Brain Mapping

Mapping all the neurons in the brain requires automatic reconstruction of entire cells from volume electron microscopy data. The flood-filling network (FFN) architecture has demonstrated leading performance for segmenting structures from this data. However, the training of the network is computationally expensive. In order to reduce the training time, we implemented synchronous and data-parallel distributed training using the Horovod library, which is different from the asynchronous training scheme used in the published FFN code. We demonstrated that our distributed training scaled well up to 2048 Intel Knights Landing (KNL) nodes on the Theta supercomputer. Our trained models achieved similar level of inference performance, but took less training time compared to previous methods. Our study on the effects of different batch sizes on FFN training suggests ways to further improve training efficiency. Our findings on optimal learning rate and batch sizes agree with previous works.

cs.DC

Ansatz from Non-Linear Optics Applied to Trapped Bose-Einstein Condensates

A simple analytical ansatz, which has been used to describe the intensity profile of the similariton laser (a laser with self-similar propagation of ultrashort pulses), is used as a variational wave function to solve the Gross-Pitaevskii equation for a wide range of interaction parameters. The variational form interpolates between the noninteracting density profile and the strongly interacting Thomas-Fermi profile smoothly. The simple form of the ansatz is modified for both cylindrically symmetric and completely anisotropic harmonic traps. The resulting ground-state density profile and energy are in very good agreement with both the analytical solutions in the limiting cases of interaction and the numerical solutions in the intermediate regime.

cond-mat.other