Searcharxiv⌕ Search

arXiv subjects

Niaz Abdolrahim

Publications and source records attributed to Niaz Abdolrahim.

2 recordsLinked to original sources

ERAF4XRD: A multimodal agentic framework for constructing validated experimental X-ray diffraction databases from scientific literature

The scientific literature contains decades of experimental measurements that remain difficult to access as structured data for modern AI and data-driven research. Much of this information is distributed across figures, captions, text, and tables, requiring experimental data and their context to be identified, connected, and verified before they can be reused. Here we introduce ERAF4XRD (Experiment Reader Agentic Framework for X-Ray Diffraction), a fully automated multimodal (i.e., image and text), multi-agent framework that reconstructs validated X-ray diffraction (XRD) records from scientific publications. ERAF4XRD downloads and screens documents, identifies XRD figures, extracts and links metadata to the corresponding experimental data, and validates outputs against source evidence using an independent validation agent. On a manually curated benchmark of 273 scientific publications containing 3,150 candidate figures, ERAF4XRD achieved up to 98.7% accuracy for XRD figure identification and generated 1,400 metadata values across 22 fields. Independent manual assessment of the final validated records yielded 98.5% precision and 90.7% recall, with no unsupported metadata observed among 443 evaluated fields. By moving beyond information extraction to the reconstruction and validation of linked experimental records, ERAF4XRD establishes an automated approach for transforming published scientific information into machine-readable experimental datasets for AI and data-driven science.

cond-mat.mtrl-sci↗

OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering

We introduce OPENXRD, a comprehensive benchmarking framework for evaluating large language models (LLMs) and multimodal LLMs (MLLMs) in crystallography question answering. The framework measures context assimilation, or how models use fixed, domain-specific supporting information during inference. The framework includes 217 expert-curated X-ray diffraction (XRD) questions covering fundamental to advanced crystallographic concepts, each evaluated under closed-book (without context) and open-book (with context) conditions, where the latter includes concise reference passages generated by GPT-4.5 and refined by crystallography experts. We benchmark 74 state-of-the-art LLMs and MLLMs, including GPT-4, GPT-5, O-series, LLaVA, LLaMA, QWEN, Mistral, and Gemini families, to quantify how different architectures and scales assimilate external knowledge. Results show that mid-sized models (7B--70B parameters) gain the most from contextual materials, while very large models often show saturation or interference and the largest relative gains appear in small and mid-sized models. Expert-reviewed materials provide significantly higher improvements than AI-generated ones even when token counts are matched, confirming that content quality, not quantity, drives performance. OPENXRD offers a reproducible diagnostic benchmark for assessing reasoning, knowledge integration, and guidance sensitivity in scientific domains, and provides a foundation for future multimodal and retrieval-augmented crystallography systems.

cs.CL↗