SearcharxivSearch

arXiv subjects

Namsoo Shin

Publications and source records attributed to Namsoo Shin.

8 recordsLinked to original sources

Confusion-Aware Rubric Optimization for LLM-based Automated Grading

Accurate and unambiguous guidelines are critical for large language model (LLM) based graders, yet manually crafting these prompts is often sub-optimal as LLMs can misinterpret expert guidelines or lack necessary domain specificity. Consequently, the field has moved toward automated prompt optimization to refine grading guidelines without the burden of manual trial and error. However, existing frameworks typically aggregate independent and unstructured error samples into a single update step, resulting in "rule dilution" where conflicting constraints weaken the model's grading logic. To address these limitations, we introduce Confusion-Aware Rubric Optimization (CARO), a novel framework that enhances accuracy and computational efficiency by structurally separating error signals. CARO leverages the confusion matrix to decompose monolithic error signals into distinct modes, allowing for the diagnosis and repair of specific misclassification patterns individually. By synthesizing targeted "fixing patches" for dominant error modes and employing a diversity-aware selection mechanism, the framework prevents guidance conflict and eliminates the need for resource-heavy nested refinement loops. Empirical evaluations on teacher education and STEM datasets demonstrate that CARO significantly outperforms existing SOTA methods. These results suggest that replacing mixed-error aggregation with surgical, mode-specific repair yields robust improvements in automated assessment scalability and precision.

cs.AI

From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG

Automated short answer grading (ASAG) is critical for scaling educational assessment, yet large language models (LLMs) often struggle with hallucinations and strict rubric adherence due to their reliance on generalized pre-training. While Rretrieval-Augmented Generation (RAG) mitigates these issues, standard "flat" vector retrieval mechanisms treat knowledge as isolated fragments, failing to capture the structural relationships and multi-hop reasoning essential for complex educational content. To address this limitation, we introduce a Graph Retrieval-Augmented Generation (GraphRAG) framework that organizes reference materials into a structured knowledge graph to explicitly model dependencies between concepts. Our methodology employs a dual-phase pipeline: utilizing Microsoft GraphRAG for high-fidelity graph construction and the HippoRAG neurosymbolic algorithm to execute associative graph traversals, thereby retrieving comprehensive, connected subgraphs of evidence. Experimental evaluations on a Next Generation Science Standards (NGSS) dataset demonstrate that this structural approach significantly outperforms standard RAG baselines across all metrics. Notably, the HippoRAG implementation achieved substantial improvements in evaluating Science and Engineering Practices (SEP), confirming the superiority of structural retrieval in verifying the logical reasoning chains required for higher-order academic assessment.

cs.CL

How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment

The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptability to diverse question types and flexibility in output formats, they also introduce new challenges related to output uncertainty, stemming from the inherently probabilistic nature of LLMs. Output uncertainty is an inescapable challenge in automatic assessment, as assessment results often play a critical role in informing subsequent pedagogical actions, such as providing feedback to students or guiding instructional decisions. Unreliable or poorly calibrated uncertainty estimates can lead to unstable downstream interventions, potentially disrupting students' learning processes and resulting in unintended negative consequences. To systematically understand this challenge and inform future research, we benchmark a broad range of uncertainty quantification methods in the context of LLM-based automatic assessment. Although the effectiveness of these methods has been demonstrated in many tasks across other domains, their applicability and reliability in educational settings, particularly for automatic grading, remain underexplored. Through comprehensive analyses of uncertainty behaviors across multiple assessment datasets, LLM families, and generation control settings, we characterize the uncertainty patterns exhibited by LLMs in grading scenarios. Based on these findings, we evaluate the strengths and limitations of different uncertainty metrics and analyze the influence of key factors, including model families, assessment tasks, and decoding strategies, on uncertainty estimates. Our study provides actionable insights into the characteristics of uncertainty in LLM-based automatic assessment and lays the groundwork for developing more reliable and effective uncertainty-aware grading systems in the future.

cs.AI

Machine Learning Assisted Design and Optimization of Transition Metal-Incorporated Carbon Quantum Dot Catalysts for Hydrogen Evolution Reaction

Development of cost-effective hydrogen evolution reaction (HER) catalysts with outstanding catalytic activity, replacing cost-prohibitive noble metal-based catalysts, is critical for practical green hydrogen production. A popular strategy for promoting the catalytic performance of noble metal-free catalysts is to incorporate earth-abundant transition metal (TM) atoms into nanocarbon platforms such as carbon quantum dots (CQDs). Although data-driven catalyst design methods can significantly accelerate the rational design of TM element-doped CQD (M@CQD) catalysts, they suffer from either a simplified theoretical model or the prohibitive cost and complexity of experimental data generation. In this study, we propose an effective and facile HER catalyst design strategy based on machine learning (ML) and ML model verification using electrochemical methods accompanied with density functional theory (DFT) simulations. Based on a Bayesian genetic algorithm (BGA) ML model, the Ni@CQD catalyst on a three-dimensional reduced graphene oxide (3D rGO) conductor is proposed as the best HER catalyst under the optimal conditions of catalyst loading, electrode type, and temperature and pH of electrolyte. We validate the ML results with electrochemical experiments, where the Ni@CQD catalyst exhibited superior HER activity, requiring an overpotential of 189 mV to achieve 10 mA cm-2 with a Tafel slope of 52 mV dec-1 and impressive durability in acidic media. We expect that this methodology and the excellent performance of the Ni@CQD catalyst provide an effective route for the rational design of highly active electrocatalysts for commercial applications.

physics.chem-ph

Confirmation of the monoclinic Cc space group for the ground state phase of Pb(Zr0.525Ti0.475)O3 (PZT525): A Combined Synchrotron X-Ray and Neutron Powder Diffraction Study

The low temperature antiferrodistortive phase transition in a pseudo-tetragonal composition of PZT with x=0.525 is investigated through a combined synchrotron x-ray and neutron powder diffraction study. It is shown that the superlattice peaks cannot be correctly accounted for in the Rietveld refinement using R3c or R3c+Cm structural models, whereas the Cc space group gives excellent fits to the superlattice peaks as well as to the perovskite peaks. This settles at rest the existing controversies about the structure of the ground state phase of PZT in the MPB region.

cond-mat.mtrl-sci

Origin of high piezoelectric response of Pb(Zr_xTi_1-x)O_3 at the morphotropic phase boundary: Role of elastic instability

Temperature dependent structural changes in a nearly pure monoclinic phase composition (x=0.525) of Pb(Zr_xTi_1-x)O_3 (PZT) have been investigated using Rietveld analysis of high-resolution synchrotron powder x-ray diffraction data and correlated with changes in the dielectric constant and planar electromechanical coupling coefficient. Our results show that the intrinsic piezoelectric response of the tetragonal phase of PZT is higher than that of the monoclinic phase. It is also shown that the high piezoelectric response of PZT may be linked with an anomalous softening of the elastic modulus (1/S_11) of the tetragonal compositions closest to the morphotropic phase boundary.

cond-mat.mtrl-sci

High-resolution synchrotron XRD study of Zr-rich compositions of Pb(Zr_xTi_1-x)O_3 (0.525\leq x \leq 0.60): evidence for the absence of the rhombohedral phase

Results of Rietveld analysis of the synchrotron XRD data on Pb(Zr_xTi_1-x)O_3 (PZT) for 0.525\leqx\leq0.60 are presented to show the absence of rhombohedral phase on the Zr-rich side of the morphotropic phase boundary (MPB). Our results reveal that the structure of PZT is monoclinic in the Cm space group for 0.525\leq x\leq 0.60. The nature of the monoclinic distortion changes from pseudo-tetragonal for 0.525\leqx\leq0.54 to pseudo-rhombohedral for x>0.54.

cond-mat.mtrl-sci

Evidence for monoclinic crystal structure and negative thermal expansion below magnetic transition temperature in Pb(Fe_1/2Nb_1/2)O_3

The existing controversy about the room temperature structure of multiferroic Pb(Fe_1/2Nb_1/2)O_3 is settled using synchrotron powder x-ray diffraction data. Results of Rietveld refinements in the temperature range 300 to 12K reveal that the structure remains monoclinic in the Cm space group down to 12K, but the lattice parameters show anomalies at the magnetic transition temperature (T_N) due to spin lattice coupling. The lattice volume exhibits negative thermal expansion behaviour, with Alpha = - 4.64 x 10^-6 K^-1, below T_N.

cond-mat.mtrl-sci