SearcharxivSearch

arXiv subjects

Haruto Suzuki

Publications and source records attributed to Haruto Suzuki.

5 recordsLinked to original sources

Revision-Aware Success Prediction from Multi-Attempt Programming Trajectories

Programming outcome prediction plays a central role in data-driven programming education, supporting learner modeling, timely intervention, and adaptive assistance. Yet predicting submission success is difficult due to heterogeneous error states, short-term revisions, and uneven future-horizon availability in programming trajectories. This study examines three prediction tasks under a unified formulation: whether the current attempt is accepted (Task~1), whether the next attempt is accepted (Task~2), and whether acceptance is reached within a three-attempt recovery window (Task~3). Each task is evaluated across current-only, pairwise, and multi-step input regimes using ML, DL, and transformer-based pretrained models (PTM), represented by LinearSVM, XGBoost, BiGRU, BiLSTM, GraphCodeBERT, and CodeT5+. Results show a consistent pattern: the current-only regime is the most reliable, while pairwise and multi-step history provide no consistent gain. ML models are the strongest and most stable overall, particularly in Tasks~1 and~3, and Task~2 is the hardest across all model families. DL and PTMs perform well on Task~3 but are more task-dependent. In the Task~3 current-only setting, XGBoost achieves AP/PR-AUC of 99.09% and MCC of 0.6325, while GraphCodeBERT and CodeT5+ reach F1 scores of 80.00% and 73.68%, respectively. A sensitivity analysis confirms that Task~3 conclusions hold most robustly for ML models under stricter future-horizon control. Across all settings, ML models remain highly effective for programming success prediction, while complex models offer value in specific settings. This work provides a systematic comparison across predictive formulations and offers robust modeling guidance for submission-aware analytics in programming education, where near-future success prediction can inform timely intervention in online judge platforms and adaptive programming support systems.

cs.CY

Large Language Models as Modal Models in Linguistics

The rapid advancement of large language models (LLMs) has intensified debates about their significance for linguistic theory. These debates are commonly divided into three positions: insulationism, which regards LLMs as irrelevant to human language; eliminativism, which claims that LLMs can replace traditional linguistic theories; and conciliationism, which views them as useful tools for linguistic research. To clarify these positions, this paper applies the framework of modal modeling from the philosophy of science. We argue that LLMs possess genuine epistemic value as minimal models, even without structural correspondence to human cognition. In particular, they can provide how-possibly explanations (HPEs) by testing modal claims about language acquisition and linguistic competence. We then examine the conditions under which LLMs could qualify as how-actually explanations (HAEs) of human language, drawing on the mechanistic account of scientific explanation. We argue that current LLMs do not yet satisfy these requirements. On the basis of this analysis, we propose understanding the explanatory power of LLMs as lying on a continuum between HPEs and HAEs. This framework avoids both overstating and understating their explanatory significance and offers a more precise basis for evaluating the role of LLMs in the scientific study of language.

cs.CL

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding

LLMs are increasingly employed both as judges for evaluating open-ended outputs and as co-creation partners in AI-assisted programming; yet rigorous evaluation in human-AI co-creation settings remains underdeveloped as judgments must be reliable, comparable across models, and interpretable over multi-turn interaction. To address this gap, a rubric-driven LLM-as-a-Judge framework is presented for contest-style human-AI co-creation in coding and software engineering (SE). The framework is built around schema-constrained judge outputs, validation and repair mechanisms, grouped and split by user and problem to prevent trajectory leakage, and participant-level NONBLIND context. Multiple LLM judges are assessed through a multi-metric protocol covering discrimination (ROC-AUC, PR-AUC), thresholded decision quality (MCC), probabilistic reliability (LogLoss, Brier score, ECE), and inter-judge agreement (Cohen's and Fleiss' k). Human-AI co-creation is further examined through trajectory-level signals, including turn-wise confidence, Success-at-Turn, time-to-success, revision churn, and CodeBLEU. Co-creation success is found to concentrate early, with Success-at-Turn rising to 0.8533 at the first observed turn and stabilizing at 0.8641 by turn 6. Revision behavior, however, remains heterogeneous, suggesting that productive progress can emerge through either incremental refinement or broader restructuring. On the judging side, the best held-out scores reach 0.5937 for ROC-AUC, 0.6904 for PR-AUC, and 0.5000 for MCC test, while inter-judge consistency remains modest overall (mean pairwise Cohen's k = 0.1592, Fleiss' k = 0.0696). Taken together, this work offers an auditable and reproducible evaluation methodology that links reliability-aware LLM judging with trajectory-based analysis of human-AI co-creation, providing a practical evaluation template for future AI-assisted coding and SE.

cs.SE

AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation

As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a novel two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (approx. 94 hours) collected with the Human Support Robot (HSR) and is fully standardized in the LeRobot v2.1 format. By uniquely integrating mobile manipulation, contact-rich interaction, and long-horizon structure, AIRoA MoMa provides a critical benchmark for advancing the next generation of Vision-Language-Action models. The first version of our dataset is now available at https://huggingface.co/datasets/airoa-org/airoa-moma .

cs.RO

Translational-Symmetry-Broken Magnetization Plateaux of the $S=3/2$ Anisotropic Antiferromagnetic Chain

The magnetization process of the $S=3/2$ quantum spin chain with the $XXZ$ anisotropy and the single-ion anisotropy $D$ is investigated using the numerical diagonalization of finite-size clusters and the level spectroscopy analysis. We obtain the phase diagrams at 1/3 and 2/3 of the saturation magnetization to find that the translational-symmetry-broken magnetization plateau appears for the first time. The similarity and the difference between the phase diagrams of the present model and the related models are discussed by use of the discrete parameters of the models. In addition several typical magnetization curves are presented.

cond-mat.str-el