Searcharxiv⌕ Search

arXiv subjects

Hiroaki Ogata

Publications and source records attributed to Hiroaki Ogata.

13 recordsLinked to original sources

Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing

Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipeline that isolates structural errors from unnaturalness by generating literal error corrections and idiomatic revisions for 3,830 English writing samples from 120 Japanese junior high school students. Applying the regex-based CEFR-J grammar extractor, we quantify two diagnostic measures: accuracy gaps (structures attempted but incorrectly produced) and idiomatic gaps (grammatically correct structures underused or overused relative to native norms). Results reveal distinct patterns: definite articles, third-person singular -s, and modals (would, could) exhibit significant accuracy difficulties, while -ing forms and hypothetical modals (would) show the largest idiomatic underuse, with simple present verbs, subject-verb-object patterns, and modal can conversely exhibiting the most pronounced overuse. A two-dimensional instructional typology maps error rates against idiomatic gaps, distinguishing accurate but overused grammar items from error-prone or avoided complex forms requiring targeted production practice. This framework advances pedagogical feedback by enabling teachers to diagnose whether learner difficulties arise from inaccurate execution, structural avoidance, or L1-mapped overreliance, supporting evidence-based interventions tailored to the specific needs of each learner.

cs.CL↗

Penny: Transition Network Analysis of Learner-Chatbot Interactions in Scaffolded EFL Writing

Generative AI chatbots promise to transform English as a Foreign Language (EFL) writing by providing immediate, personalised feedback. However, their pedagogical value depends on how learners engage with them - a process often treated as a "black box." This study uses Transition Network Analysis to model the temporal dynamics of Japanese EFL learners using "Penny," an LLM-powered writing chatbot. Analysis of over 4,500 writing sessions and 21,000 chatbot interactions reveals two dominant behavioural loops: a "Revision Loop," where feedback leads directly to successful error correction, and a "Chat Loop," where learners engage in sustained dialogue with the chatbot following feedback. Crucially, EFL proficiency significantly shapes interaction: high-proficiency learners engage more in open dialogue and negotiation with the chatbot, while low-proficiency learners rely more heavily on repetitive corrective feedback cycles. The findings demonstrate that AI-scaffolded writing is a non-linear, dialogic process and highlight the need for differentiated chatbot design to move beyond simple error correction and foster deeper cognitive engagement for all learners.

cs.CY↗

Agentic AI and Pedagogical Best Practice: The Tension Between Automation and Learning

Artificial intelligence in education is evolving from passive chatbots to proactive AI agents capable of initiation and goal-directed interactions. While offering opportunities for personalised learning, this shift risks undermining learner agency and cognitive effort. This paper reviews six pedagogical principles-prior knowledge activation, collaborative learning, problem-based learning, formative assessment, scaffolding, and metacognition-through the lens of agentic AI. We discuss the tension between automation and learning, proposing design recommendations that prioritise intentional friction, dynamic scaffolding, human-in-the-loop oversight, and considered AI utilisation to ensure AI supports rather than supplants human learning.

cs.CY↗

Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises

This study evaluates the pedagogical viability of LLM-generated English as a Foreign Language (EFL) learning content. Utilising log data from Japanese junior high school students practicing on a grammar drilling application, we analysed how different question modalities impact student performance and whether theoretical localised CEFR difficulty tiers accurately predict empirical task difficulty. Results reveal a clear performance hierarchy: multiple-choice questions carried the lowest cognitive load, cloze tasks posed the greatest barrier to active recall, and drag-and-drop exercises incurred the heaviest time penalties. Furthermore, learner data validated the CEFR-J grammar framework, showing a steady decline in accuracy and increased response times as proficiency levels advanced. These findings demonstrate that LLMs can successfully generate learning content, while highlighting the need for developers to strategically sequence question modalities to transition learners from passive recognition to active linguistic production.

cs.CY↗

Training-Free Private Synthesis with Validation: A New Frontier for Practical Educational Data Sharing

While secondary use of real-world data (RWD) in education offers substantial research opportunities, data sharing is often limited by privacy constraints. Differentially private synthetic data generation (DP-SDG) has emerged as a possible solution. However, educational RWD is fragmented across platforms and institutions and stored in different formats, so DP-SDG must be tailored to each dataset, requiring substantial engineering effort. In addition, such data are often small-sample and high-dimensional, making deep learning (DL)-based methods common but difficult to implement without specialist expertise. In this setting, it is also hard to achieve practically useful downstream utility. As a result, despite its theoretical promise, DP-SDG remains far from a practical solution in education. To address this issue, we propose a more practical two-stage method: (1) training-free, LLM-based DP-SDG is performed for sharing synthetic data and (2) on-demand real-data validation, where researchers submit code for remote validation of results. This simple method is designed for individual data custodians without extensive DP-SDG expertise. It can also be adapted to multi-shot synthesis, where data from different learner cohorts are synthesised regularly. We evaluate this method experimentally in both the one-shot and multi-shot synthesis settings using RWD collected over three years and conduct a case study with real researchers. Results show that LLM-based DP-SDG performs comparably to a DL-based baseline while greatly reducing engineering costs, and that non-DP validation causes measurable but moderate privacy leakage. Nonetheless, in the case study researchers reported that on average only 36% of synthetic findings are validated on real data. Overall, the paper provides a practical method for sharing educational RWD, while highlighting challenges in risk mitigation and epistemic precision.

cs.CY↗

Cyclic Adaptive Private Synthesis for Sharing Real-World Data in Education

The rapid adoption of digital technologies has greatly increased the volume of real-world data (RWD) in education. While these data offer significant opportunities for advancing learning analytics (LA), secondary use for research is constrained by privacy concerns. Differentially private synthetic data generation is regarded as the gold-standard approach to sharing sensitive data, yet studies on the private synthesis of educational data remain very scarce and rely predominantly on large, low-dimensional open datasets. Educational RWD, however, are typically high-dimensional and small in sample size, leaving the potential of private synthesis underexplored. Moreover, because educational practice is inherently iterative, data sharing is continual rather than one-off, making a traditional one-shot synthesis approach suboptimal. To address these challenges, we propose the Cyclic Adaptive Private Synthesis (CAPS) framework and evaluate it on authentic RWD. By iteratively sharing RWD, CAPS not only fosters open science, but also offers rich opportunities of design-based research (DBR), thereby amplifying the impact of LA. Our case study using actual RWD demonstrates that CAPS outperforms a one-shot baseline while highlighting challenges that warrant further investigation. Overall, this work offers a crucial first step towards privacy-preserving sharing of educational RWD and expands the possibilities for open science and DBR in LA.

cs.CY↗

The Third-Party Access Effect: An Overlooked Challenge in Secondary Use of Educational Real-World Data

Secondary use of growing real-world data (RWD) in education offers significant opportunities for research, yet privacy practices intended to enable third-party access to such RWD are rarely evaluated for their implications for downstream analyses. As a result, potential problems introduced by otherwise standard privacy practices may remain unnoticed. To address this gap, we investigate potential issues arising from common practices by assessing (1) the re-identification risk of fine-grained RWD, (2) how communicating such risks influences learners' privacy behaviour, and (3) the sensitivity of downstream analytical conclusions to resulting changes in the data. We focus on these practices because re-identification risk and stakeholder communication can jointly influence the data shared with third parties. We find that substantial re-identification risk in RWD, when communicated to stakeholders, can induce opt-outs and non-self-disclosure behaviours. Sensitivity analysis demonstrates that these behavioural changes can meaningfully alter the shared data, limiting validity of secondary-use findings. We conceptualise this phenomenon as the third-party access effect (3PAE) and discuss implications for trustworthy secondary use of educational RWD.

cs.CY↗

Defining the Scope of Learning Analytics: An Axiomatic Approach for Analytic Practice and Measurable Learning Phenomena

Learning Analytics (LA) has rapidly expanded through practical and technological innovation, yet its foundational identity has remained theoretically under-specified. This paper addresses this gap by proposing the first axiomatic theory that formally defines the essential structure, scope, and limitations of LA. Derived from the psychological definition of learning and the methodological requirements of LA, the framework consists of five axioms specifying discrete observation, experience construction, state transition, and inference. From these axioms, we derive a set of theorems and propositions that clarify the epistemological stance of LA, including the inherent unobservability of learner states, the irreducibility of temporal order, constraints on reachable states, and the impossibility of deterministically predicting future learning. We further define LA structure and LA practice as formal objects, demonstrating the sufficiency and necessity of the axioms and showing that diverse LA approaches -- such as Bayesian Knowledge Tracing and dashboards -- can be uniformly explained within this framework. The theory provides guiding principles for designing analytic methods and interpreting learning data while avoiding naive behaviorism and category errors by establishing an explicit theoretical inference layer between observations and states. This work positions LA as a rigorous science of state transition systems based on observability, establishing the theoretical foundation necessary for the field's maturation as a scholarly discipline.

cs.CY↗

Local Fr'echet Regression via RKHS embedding and Its Applications to Data Analysis on Manifolds

Local Fr'echet Regression (LFR) is a nonparametric regression method for settings in which the explanatory variable lies in a Euclidean space and the response variable lies in a metric space. It is used to estimate smooth trajectories in general metric spaces from noisy observations of random objects taking values in such spaces. Since metric spaces form a broad class of spaces that often lack algebraic structures such as addition or scalar multiplication characteristics typical of vector spaces the asymptotic theory for conventional random variables cannot be directly applied. As a result, deriving the asymptotic distribution of the LFR estimator is challenging. In this paper, we first extend nonparametric regression models for real-valued responses to Hilbert spaces and derive the asymptotic distribution of the LFR estimator in a Hilbert space setting. Furthermore, we propose a new estimator based on the LFR estimator in a reproducing kernel Hilbert space (RKHS), by mapping data from a general metric space into an RKHS. Finally, we consider applications of the proposed method to data lying on manifolds and construct confidence regions in metric spaces based on the derived asymptotic distribution.

math.ST↗

How Good is ChatGPT in Giving Adaptive Guidance Using Knowledge Graphs in E-Learning Environments?

E-learning environments are increasingly harnessing large language models (LLMs) like GPT-3.5 and GPT-4 for tailored educational support. This study introduces an approach that integrates dynamic knowledge graphs with LLMs to offer nuanced student assistance. By evaluating past and ongoing student interactions, the system identifies and appends the most salient learning context to prompts directed at the LLM. Central to this method is the knowledge graph's role in assessing a student's comprehension of topic prerequisites. Depending on the categorized understanding (good, average, or poor), the LLM adjusts its guidance, offering advanced assistance, foundational reviews, or in-depth prerequisite explanations, respectively. Preliminary findings suggest students could benefit from this tiered support, achieving enhanced comprehension and improved task outcomes. However, several issues related to potential errors arising from LLMs were identified, which can potentially mislead students. This highlights the need for human intervention to mitigate these risks. This research aims to advance AI-driven personalized learning while acknowledging the limitations and potential pitfalls, thus guiding future research in technology and data-driven education.

cs.AI↗

A mixture transition distribution modeling for higher-order circular Markov processes

The stationary higher-order Markov process for circular data is considered. We employ the mixture transition distribution (MTD) model to express the transition density of the process on the circle. The underlying circular transition distribution is based on Wehrly and Johnson's bivariate joint circular models. The structures of the circular autocorrelation function together with the circular partial autocorrelation function are found to be similar to those of the autocorrelation and partial autocorrelation functions of the real-valued autoregressive process when the underlying binding density has zero sine moments. The validity of the model is assessed by applying it to some Monte Carlo simulations and real directional data.

stat.ME↗

Pair circulas modelling for multivariate circular time series

Modelling multivariate circular time series is considered. The cross-sectional and serial dependence is described by circulas, which are analogs of copulas for circular distributions. In order to obtain a simple expression of the dependence structure, we decompose a multivariate circula density to a product of several pair circula densities. Moreover, to reduce the number of pair circula densities, we consider strictly stationary multi-order Markov processes. The real data analysis, in which the proposed model is fitted to multivariate time series wind direction data is also given.

stat.ME↗

Copula bounds for circular data

We propose the extension of Fréchet-Hoeffding copula bounds for circular data. The copula is a powerful tool for describing the dependency of random variables. In two dimensions, the Fréchet-Hoeffding upper (lower) bound indicates the perfect positive (negative) dependence between two random variables. However, for circular random variables, the usual concept of dependency is not accepted because of their periodicity. In this work, we redefine Fréchet-Hoeffding bounds and consider modified Fréchet and Mardia families of copulas for modelling the dependency of two circular random variables. Simulation studies are also given to demonstrate the behavior of the model.

math.ST↗