SearcharxivSearch

arXiv subjects

Augustin Kelava

Publications and source records attributed to Augustin Kelava.

6 recordsLinked to original sources

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead

Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive and psychological traits, such as intelligence and personality. While these results are often interpreted as strong evidence of human-like characteristics in LLMs, this paper argues that such interpretations constitute an ontological error. Human psychological and educational tests are theory-driven measurement instruments, calibrated to a specific human population. Applying these tests to non-human subjects without empirical validation, risks mischaracterizing what is being measured. Furthermore, a growing trend frames AI performance on benchmarks as measurements of traits such as ``intelligence'', despite known issues with validity, data contamination, cultural bias and sensitivity to superficial prompt changes. We argue that interpreting benchmark performance as measurements of human-like traits, lacks sufficient theoretical and empirical justification. This leads to our position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead. We call for the development of principled, AI-specific evaluation frameworks tailored to AI systems. Such frameworks might build on existing frameworks for constructing and validating psychometrics tests, or could be created entirely from scratch to fit the unique context of AI.

cs.LG

Reconciling Latent Variables and Networks: Exploring and extending the Psychometric-Toolbox

Since the introduction of network psychometrics, several connections to statistical models in "classical" psychometrics (i.e., IRT, SEM, GLM) as well as to approaches from other research fields have been established. In this paper, these developments have been reviewed and synthesized and, based on an exploratory literature search, further advanced and presented in an accessible visual format. This perspective opens up promising opportunities to extend the psychometric-toolbox by incorporating and learning from statistical methodologies developed in other research domains, which often address similar or even identical problems. Highlighting these methodological commonalities may also foster collaboration across research fields that have traditionally remained largely independent. Moreover, awareness of these connections may render methodological development more systematic and goal-directed and may enable a meaningful division of labor, for example between the development of statistical methodology and its practical implementation for empirical research through software tools. Finally, these methodological advances provide new opportunities for empirical research and may contribute to a reconciliation with longstanding conceptual issues concerning psychometric constructs and, more broadly, psychological phenomena.

stat.ME

Frequentist forecasting in regime-switching models with extended Hamilton filter

Psychological change processes, such as university student dropout in math, often exhibit discrete latent state transitions and can be studied using regime-switching models with intensive longitudinal data (ILD). Recently, regime-switching state-space (RSSS) models have been extended to allow for latent variables and their autoregressive effects. Despite this progress, estimation methods for handling both intra-individual changes and inter-individual differences as predictors of regime-switches need further exploration. Specifically, there's a need for frequentist estimation methods in dynamic latent variable frameworks that allow real-time inferences and forecasts of latent or observed variables during ongoing data collection. Building on Chow and Zhang's (2013) extended Kim filter, we introduce a first frequentist filter for RSSS models which allows hidden Markov(-switching) models to depend on both latent within- and between-individual characteristics. As a counterpart of Kelava et al.'s (2022) Bayesian forecasting filter for nonlinear dynamic latent class structural equation models (NDLC-SEM), our proposed method is the first frequentist approach within this general class of models. In an empirical study, the filter is applied to forecast emotions and behavior related to student dropout in math. Parameter recovery and prediction of regime and dynamic latent variables are evaluated through simulation study.

stat.ME

On an EM-based closed-form solution for 2 parameter IRT models

It is a well-known issue that in Item Response Theory models there is no closed-form for the maximum likelihood estimators of the item parameters. Parameter estimation is therefore typically achieved by means of numerical methods like gradient search. The present work has a two-fold aim: On the one hand, we revise the fundamental notions associated to the item parameter estimation in 2 parameter Item Response Theory models from the perspective of the complete-data likelihood. On the other hand, we argue that, within an Expectation-Maximization approach, a closed-form for discrimination and difficulty parameters can actually be obtained that simply corresponds to the Ordinary Least Square solution.

stat.ME

Challenging the Validity of Personality Tests for Large Language Models

With large language models (LLMs) like GPT-4 appearing to behave increasingly human-like in text-based interactions, it has become popular to attempt to evaluate personality traits of LLMs using questionnaires originally developed for humans. While reusing measures is a resource-efficient way to evaluate LLMs, careful adaptations are usually required to ensure that assessment results are valid even across human subpopulations. In this work, we provide evidence that LLMs' responses to personality tests systematically deviate from human responses, implying that the results of these tests cannot be interpreted in the same way. Concretely, reverse-coded items ("I am introverted" vs. "I am extraverted") are often both answered affirmatively. Furthermore, variation across prompts designed to "steer" LLMs to simulate particular personality types does not follow the clear separation into five independent personality factors from human samples. In light of these results, we believe that it is important to investigate tests' validity for LLMs before drawing strong conclusions about potentially ill-defined concepts like LLMs' "personality".

cs.CL

Forecasting intra-individual changes of affective states taking into account inter-individual differences using intensive longitudinal data from a university student drop out study in math

The longitudinal process that leads to university student drop out in STEM subjects can be described by referring to a) inter-individual differences (e.g., cognitive abilities) as well as b) intra-individual changes (e.g., affective states), c) (unobserved) heterogeneity of trajectories, and d) time-dependent variables. Large dynamic latent variable model frameworks for intensive longitudinal data (ILD) have been proposed which are (partially) capable of simultaneously separating the complex data structures (e.g., DLCA; Asparouhov, Hamaker, & Muthén, 2017; DSEM; Asparouhov, Hamaker, & Muthén, 2018; NDLC-SEM, Kelava & Brandt, 2019). From a methodological perspective, forecasting in dynamic frameworks allowing for real-time inferences on latent or observed variables based on ongoing data collection has not been an extensive research topic. From a practical perspective, there has been no empirical study on student drop out in math that integrates ILD, dynamic frameworks, and forecasting of critical states of the individuals allowing for real-time interventions. In this paper, we show how Bayesian forecasting of multivariate intra-individual variables and time-dependent class membership of individuals (affective states) can be performed in these dynamic frameworks. To illustrate our approach, we use an empirical example where we apply forecasting methodology to ILD from a large university student drop out study in math with multivariate observations collected over 50 measurement occasions from multiple students (N = 122). More specifically, we forecast emotions and behavior related to drop out. This allows us to model (i) just-in-time interventions, (ii) detection of heterogeneity in trajectories, and (iii) prediction of emerging dynamic states (e.g. critical stress levels or pre-decisional states).

stat.ME