SearcharxivSearch

arXiv subjects

Ick Hoon Jin

Publications and source records attributed to Ick Hoon Jin.

At least 19 recordsLinked to original sources

Constructing Reliable Social Networks from Conversational Data: An Ensemble Prompt Engineering Approach with Uncertainty Quantification

Conversational data are central to the study of interaction dynamics and social structures across psychological research. However, constructing structured social networks from unstructured conversational data remains a major methodological challenge. This study presents a pipeline for network construction using prompt engineering. We employ an ensemble of five Large Language Models (LLMs) with majority voting to automate utterance classification, reducing dependence on manual coding without task-specific parameter fine-tuning. Classification uncertainty is assessed through an uncertainty quantification framework based on Shannon entropy, which can be used to prioritize ambiguous cases for review. The classified utterances are used to construct directed interaction networks for subsequent analysis. Reliability and accuracy are established relative to the criterion-referenced and human comparisons reported here rather than as unconditional guarantees across settings. We demonstrate the utility of this approach through two illustrative applications to classroom interaction data: network centrality analysis to characterize participant roles, and network mediation analysis using the additive and multiplicative effects network (AMEN) model to examine how interaction structures mediate the relationship between gender and mathematics performance. This pipeline provides a scalable foundation for automated network construction from conversational data across diverse research contexts.

stat.AP

Issue-Specific Polarization and Cohesion in a Multi-Party Legislature: Integrating the Latent Space Item Response Model with Topic-Based Regression

We develop a two-stage, cut-posterior Bayesian framework for quantifying issue-specific legislative alignment in multi-party systems. The approach integrates a Latent Space Item Response Model (LSIRM), embedding legislators and bills in a shared Euclidean space, with Bayesian beta regression using text-derived topic proportions as bill-level covariates. The resulting legislator- and issue-specific coefficients allow polarization and cohesion to be compared across policy domains. The beta-regression layer must not feed back into the estimated latent geometry, so that the two components target a cut posterior. We estimate it with a two-stage Multiple-Imputation procedure that propagates uncertainty in the latent positions into every downstream quantity. In the 17th Korean National Assembly, fiscal domains such as Taxation and Grants and Local Government Budget show sharp polarization with tight within-party clustering, whereas Armed Services, Patriots, and Veterans exhibits weak party structuring and greater intra-party variability. The Democratic Labor Party forms a distinct cluster on several issues even where the two major parties are not strongly polarized, showing that legislative conflict escapes a single left--right ordering. The framework supports analysis of issue-structured voting in legislatures where one-dimensional ideal point models are unreliable.

stat.AP

Modeling Transition Dynamics and Network Structure in Cross-National Process Data: A Hierarchical Multi-State Survival Framework

Process data from computer-based assessments record the sequence and timing of actions through which respondents solve a task, providing information about both the pace and structure of problem-solving behavior. Modeling such processes across countries is challenging because country-by-response-group cells are often small and unbalanced and the observed transition supports can differ substantially across countries. We propose a hierarchical framework that integrates a Bayesian multi-state survival model with a network-based representation of transition structure. Partial pooling across countries yields country-specific covariate and key-action effects, transition speed, and estimates of between-country heterogeneity. Posterior transition probability networks are embedded in a common latent space using a directed graph auto-encoder adapted to heterogeneous supports, and 1-Wasserstein distances between node-role distributions are evaluated across posterior draws to characterize global network structure while propagating estimation uncertainty. We apply the framework to two problem-solving items from the Programme for the International Assessment of Adult Competencies across 14 countries. The results reveal cross-country heterogeneity in transition speed, systematic response-group differences in global network organization, item-dependent variation in within-group dispersion across countries, and local differences in routing around shared intermediate actions.

stat.AP

A Representation-Learning Item Response Model for Identifying Behaviorally Important Actions in PIAAC Process Data

Problem-solving log process data from computer-based assessments provide detailed information about how respondents approach and complete tasks. However, the resulting action sequences are complex and noisy, making it difficult to identify specific behaviors associated with successful performance. This paper proposes a representation-learning item response modeling (IRT) framework for identifying behaviorally important actions while accounting for respondent proficiency and item-level differences. Raw log sequences and timing information are first transformed into action representations that incorporate the hierarchical structure of action labels and the sequential and temporal context in which each action occurs. These respondent-specific representations are then entered as covariates in an extended IRT model, with spike-and-slab priors used to identify action-item combinations associated with response accuracy. The framework therefore evaluates actions contextually rather than as simple occurrence indicators and provides posterior uncertainty for their associations with performance. We apply the approach to problem-solving process data from the OECD Programme for the International Assessment of Adult Competencies (PIAAC). The analysis identifies a sparse set of actions associated with successful and unsuccessful performance and reveals differences across items in where behavioral information occurs within the problem-solving process.

stat.AP

Analyzing Process Data from Computer-Based Assessments: A Tutorial on Preprocessing, Feature Extraction, and Model-Based Inference

Computer-based assessments routinely generate detailed interaction logs -- commonly referred to as process data -- that record every action a respondent performs during task completion, yet systematic preprocessing guidance, integrated analytical workflows, and cross-method consistency checks remain scarce in the literature. This paper provides a unified, end-to-end analytical framework for analyzing process data from large-scale assessments -- covering the full pipeline from raw log preprocessing to model-based inference -- using the Programme for the International Assessment of Adult Competencies (PIAAC) Problem Solving in Technology-Rich Environments (PS-TRE) domain as an illustrative example. We first present a systematic preprocessing pipeline -- including timestamp correction, duplicate removal, action block consolidation, and LLM-assisted standardization -- that transforms raw event-level logs into analysis-ready action sequences. We then review and demonstrate two complementary families of analytical methods. The first consists of feature-based methods and their downstream applications, including descriptive process indicators, n-gram analysis with TF--IDF weighting, multidimensional scaling, and process data-informed differential item functioning (DIF) analysis. The second consists of model-based approaches, namely hidden Markov models and the subtask identification procedure. Empirical illustrations using the United States sample illustrate that n-gram-based behavioral clusters carry differential diagnostic information primarily among incorrect respondents, that multidimentionsl scaling-derived features comprehensively reconstruct observed behavioral variables, and that process-informed DIF analyses can identify and mitigate construct-irrelevant sources of group differences. Reproducible R code implementations are provided for all major techniques.

stat.AP

Euclidean Ideal Point Estimation From Roll-Call Data via Distance-Based Bipartite Network Models

Conventional ideal point models rely on Gaussian or quadratic utility functions that violate the triangle inequality, producing non-metric distances that complicate geometric interpretation and undermine clustering and dispersion-based analyses. We introduce a distance-based alternative that adapts the Latent Space Item Response Model (LSIRM) to roll-call data, treating legislators and bills as nodes in a bipartite network jointly embedded in a Euclidean metric space. Through controlled simulations, Euclidean LSIRM consistently recovers latent coalition structure with superior cluster separation relative to existing methods. Applied to the 118th U.S. House, the model provides competitive predictive performance while yielding bill embeddings that clarify cross-cutting issue alignments. The results show that restoring metric structure to ideal point estimation provides a clearer and more coherent inference about party cohesion, factional divisions, and multidimensional legislative behavior.

stat.AP

Hierarchical Latent Space Item Response Model for Analyzing Mental Health Vulnerability of Elementary School Students in South Korea

Mental health difficulties among elementary school students represent a growing public health concern in South Korea, yet analytical tools for identifying school-specific vulnerability patterns from item response data remain limited. We propose the hierarchical latent space item response model (HLSIRM), which adds hierarchical respondent effects and an inner-product latent interaction for signed respondent-item associations, yielding a unified interaction map that separates school, individual main effects from school/individual-item interactions. We apply HLSIRM to mental health vulnerability data from 2,210 elementary school students across 35 schools in Incheon, South Korea. Clustering item vectors by directional similarity identifies four empirically derived vulnerability domains. School-level analysis reveals that the absence of counseling experience is the primary vulnerability domain aligned with most school vectors, while stress, depression, and smartphone dependency concentrate in specific schools. Within-school analysis demonstrates how individual student positions in the interaction map translate into targeted intervention strategies that address school-specific needs.

stat.AP

Analysis of Log Data from an International Online Educational Assessment System: A Multi-state Survival Modeling Approach to Reaction Time between and across Action Sequence

With increasingly available computer-based or online assessments, researchers have shown keen interest in analyzing log data to improve our understanding of test takers' problem-solving processes. In this paper, we propose a multi-state survival model (MSM) to action sequence data from log files, focusing on modeling test takers' reaction times between actions, in order to investigate which factors and how they influence test takers' transition speed between actions. We specifically identify the key actions that differentiate correct and incorrect answers, compare transition probabilities between these groups, and analyze their distinct problem-solving patterns. Through simulation studies and sensitivity analyses, we evaluate the robustness of our proposed model. We demonstrate the proposed approach using problem-solving items from the Programme for the International Assessment of Adult Competencies (PIAAC).

stat.AP

lsirm12pl: An R package for latent space item response modeling

The item response model in latent space (LSIRM; Jeon et al., 2021) uncovers unobserved interactions between respondents and items in the item response data by embedding both in a shared latent metric space. The R package lsirm12pl implements Bayesian estimation of the LSIRM and its extensions for various response types, base model specifications, and missing data handling. Furthermore, lsirm12pl package provides methods to improve model utilization and interpretation, such as clustering item positions on an estimated interaction map. The package also offers convenient summary and plotting options to evaluate and process the estimated results. In this paper, we provide an overview of the LSIRM's methodological foundation and describe several extensions included in the package. We then demonstrate the use of the package with real data examples contained within it.

stat.ME

Impacts of Innovation School System in Korea: A Latent Space Item Response Model with Neyman-Scott Point Process

South Korea's educational system has faced criticism for its lack of focus on critical thinking and creativity, resulting in high levels of stress and anxiety among students. As part of the government's effort to improve the educational system, the innovation school system was introduced in 2009, which aims to develop students' creativity as well as their non-cognitive skills. To better understand the differences between innovation and regular school systems in South Korea, we propose a novel method that combines the latent space item response model (LSIRM) with the Neyman-Scott (NS) point process model. Our method accounts for the heterogeneity of items and students, captures relationships between respondents and items, and identifies item and student clusters that can provide a comprehensive understanding of students' behaviors/perceptions on non-cognitive outcomes. Our analysis reveals that students in the innovation school system show a higher sense of citizenship, while those in the regular school system tend to associate confidence in appearance with social ability. We compare our model with exploratory item factor analysis in terms of item clustering and find that our approach provides a more detailed and automated analysis. A comparison with exploratory item factor analysis highlights our method's advantages in terms of uncertainty quantification of the clustering process and more detailed and nuanced clustering results. Our method is made available to an existing R package, lsirm12pl.

stat.AP

Network-based Topic Structure Visualization

In the real world, many topics are inter-correlated, making it challenging to investigate their structure and relationships. Understanding the interplay between topics and their relevance can provide valuable insights for researchers, guiding their studies and informing the direction of research. In this paper, we utilize the topic-words distribution, obtained from topic models, as item-response data to model the structure of topics using a latent space item response model. By estimating the latent positions of topics based on their distances toward words, we can capture the underlying topic structure and reveal their relationships. Visualizing the latent positions of topics in Euclidean space allows for an intuitive understanding of their proximity and associations. We interpret relationships among topics by characterizing each topic based on representative words selected using a newly proposed scoring scheme. Additionally, we assess the maturity of topics by tracking their latent positions using different word sets, providing insights into the robustness of topics. To demonstrate the effectiveness of our approach, we analyze the topic composition of COVID-19 studies during the early stage of its emergence using biomedical literature in the PubMed database. The software and data used in this paper are publicly available at https://github.com/jeon9677/gViz .

stat.AP

A Latent Space Accumulator Model for Response Time: Applications to Cognitive Assessment Data

Response time has attracted increased interest in educational and psychological assessment for, e.g., measuring test takers' processing speed, improving the measurement accuracy of ability, and understanding aberrant response behavior. Most models for response time analysis are based on a parametric assumption about the response time distribution. The Cox proportional hazard model has been utilized for response time analysis for the advantages of not requiring a distributional assumption of response time and enabling meaningful interpretations with respect to response processes. In this paper, we present a new version of the proportional hazard model, called a latent space accumulator model, for cognitive assessment data based on accumulators for two competing response outcomes, such as correct vs. incorrect responses. The proposed model extends a previous accumulator model by capturing dependencies between respondents and test items across accumulators in the form of distances in a two-dimensional Euclidean space. A fully Bayesian approach is developed to estimate the proposed model. The utilities of the proposed model are illustrated with two real data examples.

stat.ME

How social networks influence human behavior: An integrated latent space approach for differential social influence

How social networks influence human behavior has been an interesting topic in applied research. Existing methods often utilized scale-level behavioral data to estimate the influence of a social network on human behavior. This study proposes a novel approach to studying social influence that utilizes item-level behavioral measures. Under the latent space modeling framework, we integrate the two interaction maps for respondents' social network data and item-level behavior measures. The interaction map visualizes the association between the latent homophily of the respondents and their behaviors measured at the item level in a low-dimensional latent space, revealing the potential, differential social influence effects across specific behaviors measured at the item level. We also measure overall social influence as the impact of the interaction map configuration contributed by the social network data on the behavior data. The performance and properties of the proposed approach are evaluated via simulation studies. We apply the proposed model to an empirical dataset to demonstrate how the students' friendship network influences their participation in school activities.

cs.SI

Network-based Topic Interaction Map for Big Data Mining of COVID-19 Biomedical Literature

Since the emergence of the worldwide pandemic of COVID-19, relevant research has been published at a dazzling pace, which yields an abundant amount of big data in biomedical literature. Due to the high volum of relevant literature, it is practically impossible to follow up the research manually. Topic modeling is a well-known unsupervised learning that aims to reveal latent topics from text data. In this paper, we propose a novel analytical framework for estimating topic interactions and effective visualization to improve topics' relationships. We first estimate topic-word distributions using the biterm topic model and estimate the topics' interaction based on the word distribution using the latent space item response model. We mapped these latent topics onto networks to visualize relationships among the topics. Moreover, in the proposed approach, we developed a score that is helpful in selecting meaningful words that characterize the topic. We figure out how topics are related by looking at how their relationships change. We do this with a "trajectory plot" that is made with different levels of word richness. These findings provide a thoroughly mined and intuitive representation of relationships between topics related to a specific research area. The application of this proposed framework to the PubMed literature demonstrates utility of our approach in understanding of the topic composition related to COVID-19 studies in the stage of its emergence.

cs.IR

Quantile Regression with Multiple Proxy Variables

Data integration has become increasingly popular owing to the availability of multiple data sources. This study considered quantile regression estimation when a key covariate had multiple proxies across several datasets. In a unified estimation procedure, the proposed method incorporates multiple proxies that have various relationships with the unobserved covariates. The proposed approach allows the inference of both the quantile function and unobserved covariates. Moreover, it does not require the quantile function's linearity and, simultaneously, accommodates both the linear and nonlinear proxies. Simulation studies have demonstrated that this methodology successfully integrates multiple proxies and revealed quantile relationships for a wide range of nonlinear data. The proposed method is applied to administrative data obtained from the Survey of Household Finances and Living Conditions provided by Statistics Korea, to specify the relationship between assets and salary income in the presence of multiple income records.

stat.ME

Comparing multiple latent space embeddings using topological analysis

The latent space model is one of the well-known methods for statistical inference of network data. While the model has been much studied for a single network, it has not attracted much attention to analyze collectively when multiple networks and their latent embeddings are present. We adopt a topology-based representation of latent space embeddings to learn over a population of network model fits, which allows us to compare networks of potentially varying sizes in an invariant manner to label permutation and rigid motion. This approach enables us to propose algorithms for clustering and multi-sample hypothesis tests by adopting well-established theories for Hilbert space-valued analysis. After the proposed method is validated via simulated examples, we apply the framework to analyze educational survey data from Korean innovative school reform.

stat.ME

A Bayesian Precision Response-adaptive Phase II Clinical Trial Design for Radiotherapies with Competing Risk Survival Outcomes

Many phase II clinical trials have used survival outcomes as the primary endpoints in recent decades. Suppose the radiotherapy is evaluated in a phase II trial using survival outcomes. In that case, the competing risk issue often arises because the time to disease progression can be censored by the time to normal tissue complications, and vice versa. Besides, much literature has examined that patients receiving the same radiotherapy dose may yield distinct responses due to their heterogeneous radiation susceptibility statuses. Therefore, the "one-dose-fit-all" strategy often fails, and it is more relevant to evaluate the subgroup-specific treatment effect with the subgroup defined by the radiation susceptibility status. In this paper, we propose a Bayesian precision phase II trial design evaluating the subgroup-specific treatment effects of radiotherapy. We use the cause-specific hazard approach to model the competing risk survival outcomes. We propose restricting the candidate radiation doses based on each patient's radiation susceptibility status. Only the clinically feasible personalized dose will be considered, which enhances the benefit for the patients in the trial. In addition, we propose a stratified Bayesian adaptive randomization scheme such that more patients will be randomized to the dose reporting more favorable survival outcomes. Numerical studies have shown that the proposed design performed well and outperformed the conventional design ignoring the competing risk issue.

stat.AP

Multilevel Network Item Response Modeling for Discovering Differences Between Innovation and Regular School Systems in Korea

The innovation school system in South Korea has been developed in response to the traditional high-pressure school system in South Korea, with a view to cultivating a bottom-up and student-centered educational culture. Despite its ambitious goals, questions have been raised about the success of the innovation school system. Leveraging data from the Gyeonggi Education Panel Study (GEPS) along with advances in the statistical analysis of network data and educational data, we compare the two school systems in more depth. We find that some schools are indeed different from others, and those differences are not detected by conventional multilevel models. Having said that, we do not find much evidence that the innovation school system differs from the regular school system in terms of self-reported mental well-being, although we do detect differences among some schools that appear to be unrelated to the school system.

stat.AP