SearcharxivSearch

arXiv subjects

Songhee Han

Publications and source records attributed to Songhee Han.

4 recordsLinked to original sources

Why teaching resists automation in an AI-inundated era: Human judgment, non-modular work, and the limits of delegation

Debates about artificial intelligence (AI) in education often portray teaching as a modular and procedural job that can increasingly be automated or delegated to technology. This brief communication paper argues that such claims depend on treating teaching as more separable than it is in practice. Drawing on recent literature and empirical studies of large language models and retrieval-augmented generation systems, I argue that although AI can support some bounded functions, instructional work remains difficult to automate in meaningful ways because it is inherently interpretive, relational, and grounded in professional judgment. More fundamentally, teaching and learning are shaped by human cognition, behavior, motivation, and social interaction in ways that cannot be fully specified, predicted, or exhaustively modeled. Tasks that may appear separable in principle derive their instructional value in practice from ongoing contextual interpretation across learners, situations, and relationships. As long as educational practice relies on emergent understanding of human cognition and learning, teaching remains a form of professional work that resists automation. AI may improve access to information and support selected instructional activities, but it does not remove the need for human judgment and relational accountability that effective teaching requires.

cs.CL

How Trustworthy Are LLM-as-Judge Ratings for Interpretive Responses? Implications for Qualitative Research Workflows

As qualitative researchers show growing interest in using automated tools to support interpretive analysis, a large language model (LLM) is often introduced into an analytic workflow as is, without systematic evaluation of interpretive quality or comparison across models. This practice leaves model selection largely unexamined despite its potential influence on interpretive outcomes. To address this gap, this study examines whether LLM-as-judge evaluations meaningfully align with human judgments of interpretive quality and can inform model-level decision making. Using 712 conversational excerpts from semi-structured interviews with K-12 mathematics teachers, we generated one-sentence interpretive responses using five widely adopted inference models: Command R+ (Cohere), Gemini 2.5 Pro (Google), GPT-5.1 (OpenAI), Llama 4 Scout-17B Instruct (Meta), and Qwen 3-32B Dense (Alibaba). Automated evaluations were conducted using AWS Bedrock's LLM-as-judge framework across five metrics, and a stratified subset of responses was independently rated by trained human evaluators on interpretive accuracy, nuance preservation, and interpretive coherence. Results show that LLM-as-judge scores capture broad directional trends in human evaluations at the model level but diverge substantially in score magnitude. Among automated metrics, Coherence showed the strongest alignment with aggregated human ratings, whereas Faithfulness and Correctness revealed systematic misalignment at the excerpt level, particularly for non-literal and nuanced interpretations. Safety-related metrics were largely irrelevant to interpretive quality. These findings suggest that LLM-as-judge methods are better suited for screening or eliminating underperforming models than for replacing human judgment, offering practical guidance for systematic comparison and selection of LLMs in qualitative research workflows.

cs.CL

Determination of the thickness and orientation of few-layer tungsten ditelluride using polarized Raman spectroscopy

Orthorhombic tungsten ditelluride (WTe2), with a distorted 1T structure, exhibits a large magnetoresistance that depends on the orientation, and its electrical characteristics changes rom semimetallic to insulating as the thickness decreases. Through polarized Raman spectroscopy in combination with transmission electron diffraction, we establish a reliable method to determine the thickness and crystallographic orientation of few-layer WTe2. The Raman spectrum shows a pronounced dependence on the polarization of the excitation laser. We found that the separation between two Raman peaks at ~90 cm-1 and at 80-86 cm-1, depending on thickness, is a reliable fingerprint for determination of the thickness. For determination of the crystallographic orientation, the polarization dependence of the A1 modes, measured with the 632.8-nm excitation, turns out to be the most reliable. We also discovered that the polarization behaviors of some of the Raman peaks depend on the excitation wavelength as well as thickness, indicating a close interplay between the band structure and anisotropic Raman scattering cross section.

cond-mat.mes-hall

Raman Signatures of Polytypism in Molybdenum Disulfide

Since the stacking order sensitively affects various physical properties of layered materials, accurate determination of the stacking order is important for studying the basic properties of these materials as well as for device applications. Because 2H-molybdenum disulfide (MoS2) is most common in nature, most studies so far have focused on 2H-MoS2. However, we found that the 2H, 3R, and mixed stacking sequences exist in few-layer MoS2 exfoliated from natural molybdenite crystals. The crystal structures are confirmed by HR-TEM measurements. The Raman signatures of different polytypes are investigated by using 3 different excitation energies which are non-resonant and resonant with A and C excitons, respectively. The low-frequency breathing and shear modes show distinct differences for each polytype whereas the high-frequency intra-layer modes show little difference. For resonant excitations at 1.96 and 2.81 eV, distinct features are observed which enable determination of the stacking order.

cond-mat.mtrl-sci