SearcharxivSearch

arXiv subjects

Dae-Jin Lee

Publications and source records attributed to Dae-Jin Lee.

12 recordsLinked to original sources

Modelling Athletic Ageing Relative to an Estimated Performance Envelope

Athletic careers yield sparse, irregular longitudinal series: few seasons per athlete, incomplete paths, and selection into continued play. Scientific interest often centres on proximity to peak attainable performance-a ceiling-rather than on the mean trajectory, and on how that proximity co-varies with the rate of decline. Standard tools address only pieces of this problem. Linear mixed models describe average ageing; functional principal components describe dominant modes of variation; shape-invariant models register curves about a mean template; growth charts estimate population centiles but stop short of individual latent trajectories. We develop RACE (Relative Aging Curves via Envelopes), a two-stage framework that estimates a population performance envelope as an age-conditional high centile and then models each athlete as a low-dimensional geometric transformation of that envelope. STAR (Shape Translation And Rotation) is the Stage 2 mixed model, with parameters for level, timing, and tempo. Envelope geometry determines which of these parameters are identifiable when careers are short: near-linear envelopes identify level and tempo only, whereas curved envelopes identify all three. Embedding STAR in a nonlinear mixed-effects hierarchy makes the level-tempo association a parameter of the random-effect covariance rather than a post-hoc correlation of separate fits. Simulations ask whether that association is recoverable under sparsity, how geometry governs identifiability, and how sensitive results are to envelope misspecification. Applied to Major League Baseball Statcast sprint speed and bolt rate, the analysis demonstrates how the proposed embedding separates athletic level and ageing tempo in Functional Ageing Space, and shows why timing is identifiable for bolt rate but not for sprint speed.

stat.AP

Learning Latent Memory States from Longitudinal Athlete Monitoring Data

We propose a new unit of analysis for longitudinal data: the Latent Memory Table. The scientific contribution is not the encoder. It is that table, treated as a reusable statistical object on the same footing as a matrix of principal-component scores, a table of estimated random effects, or a table of predicted probabilities. We estimate a statistical table that summarizes recent longitudinal history and is intended to be stored, queried, analysed and reused throughout the statistical workflow. A memory operator maps each masked windowed history to a finite-dimensional state; collecting those states with uncertainty yields the Latent Memory Table. Validation is organized around six properties---recoverability, personalization, temporal coherence, interpretability, stability and reusability---summarized by a composite quality index \(Q\); the Transformer, the SoccerMon case study and the simulations exist to argue that this table deserves that status. Classical exponentially weighted moving averages and related short- and long-horizon scalar summaries arise as restricted, typically univariate special cases of the same operator class. A simulation study with known memory mechanisms shows that \(Q\) and rotation-invariant recovery scores discriminate genuine multivariate or personalized memory from negative controls and from misspecified windows, whereas regime classification accuracy alone does not. SoccerMon serves as an empirical case study: a constructed Latent Memory Table attains \(Q\approx 0.73\) versus about \(0.40\) for classical and lagged principal-component baselines, with incremental held-out value for some wellness targets and Procrustes ensembles for row-wise reliability.

stat.CO

Soil Texture Prediction with Bayesian Generalized Additive Models for Spatial Compositional Data

Compositional data (CoDa) play an important role in many fields such as ecology, geology, and biology. The most widely used modeling approaches are based on the Dirichlet and the logistic-normal formulation under Aitchison geometry. Recent developments in simplex geometry allow regression models to be expressed in terms of coordinates and their coefficients to be estimated. Once the model is projected into real space, a multivariate Gaussian regression can be employed. However, most existing methods focus on linear models, and there is a lack of flexible alternatives such as additive or spatial models, especially within a Bayesian framework and with practical implementation details. In this work, we present a geoadditive regression model for CoDa from a Bayesian perspective using the \texttt{brms} package in R. The model applies the isometric log-ratio (ilr) transformation and penalized splines to incorporate nonlinear effects. We also propose two new Bayesian goodness-of-fit measures for CoDa regression: BR-CoDa-$R^2$ and BM-CoDa-$R^2$, extending the Bayesian $R^2$ to the compositional setting. {These measures, alongside WAIC, support the descriptive assessment of explained variability and complement model evaluation.} The methodology is validated through simulation studies and applied to predict soil texture composition in the Basque Country. Results demonstrate good performance, interpretable spatial patterns, and reliable quantification of explained variability in compositional outcomes.

stat.ME

Stochastic EM Estimation and Inference for Zero-Inflated Beta-Binomial Mixed Models for Longitudinal Count Data

Analyzing overdispersed, zero-inflated, longitudinal count data poses significant modeling and computational challenges, which standard count models (e.g., Poisson or negative binomial mixed effects models) fail to adequately address. We propose a Zero-Inflated Beta-Binomial Mixed Effects Regression (ZIBBMR) model that augments a beta-binomial count model with a zero-inflation component, fixed effects for covariates, and subject-specific random effects, accommodating excessive zeros, overdispersion, and within-subject correlation. Maximum likelihood estimation is performed via a Stochastic Approximation EM (SAEM) algorithm with latent variable augmentation, which circumvents the model's intractable likelihood and enables efficient computation. Simulation studies show that ZIBBMR achieves accuracy comparable to leading mixed-model approaches in the literature and surpasses simpler zero-inflated count formulations, particularly in small-sample scenarios. As a case study, we analyze longitudinal microbiome data, comparing ZIBBMR with an external Zero-Inflated Beta Regression (ZIBR) benchmark; the results indicate that applying both count- and proportion-based models in parallel can enhance inference robustness when both data types are available.

stat.ME

Modeling Psychological Profiles in Volleyball via Mixed-Type Bayesian Networks

Psychological attributes rarely operate in isolation: coaches reason about networks of related traits. We analyze a new dataset of 164 female volleyball players from Italy's C and D leagues that combines standardized psychological profiling with background information. To learn directed relationships among mixed-type variables (ordinal questionnaire scores, categorical demographics, continuous indicators), we introduce latent MMHC, a hybrid structure learner that couples a latent Gaussian copula and a constraint-based skeleton with a constrained score-based refinement to return a single DAG. We also study a bootstrap-aggregated variant for stability. In simulations spanning sample size, sparsity, and dimension, latent Max-Min Hill-Climbing (MMHC) attains lower structural Hamming distance and higher edge recall than recent copula-based learners while maintaining high specificity. Applied to volleyball, the learned network organizes mental skills around goal setting and self-confidence, with emotional arousal linking motivation and anxiety, and locates Big-Five traits (notably neuroticism and extraversion) upstream of skill clusters. Scenario analyses quantify how improvements in specific skills propagate through the network to shift preparation, confidence, and self-esteem. The approach provides an interpretable, data-driven framework for profiling psychological traits in sport and for decision support in athlete development.

cs.LG

Will AI Take My Job? Evolving Perceptions of Automation and Labor Risk in Latin America

As artificial intelligence and robotics increasingly reshape the global labor market, understanding public perceptions of these technologies becomes critical. We examine how these perceptions have evolved across Latin America, using survey data from the 2017, 2018, 2020, and 2023 waves of the Latinobarómetro. Drawing on responses from over 48,000 individuals across 16 countries, we analyze fear of job loss due to artificial intelligence and robotics. Using statistical modeling and latent class analysis, we identify key structural and ideological predictors of concern, with education level and political orientation emerging as the most consistent drivers. Our findings reveal substantial temporal and cross-country variation, with a notable peak in fear during 2018 and distinct attitudinal profiles emerging from latent segmentation. These results offer new insights into the social and structural dimensions of AI anxiety in emerging economies and contribute to a broader understanding of public attitudes toward automation beyond the Global North.

cs.CY

Understanding support for AI regulation: A Bayesian network perspective

As artificial intelligence (AI) becomes increasingly embedded in public and private life, understanding how citizens perceive its risks, benefits, and regulatory needs is essential. To inform ongoing regulatory efforts such as the European Union's proposed AI Act, this study models public attitudes using Bayesian networks learned from the nationally representative 2023 German survey Current Questions on AI. The survey includes variables on AI interest, exposure, perceived threats and opportunities, awareness of EU regulation, and support for legal restrictions, along with key demographic and political indicators. We estimate probabilistic models that reveal how personal engagement and techno-optimism shape public perceptions, and how political orientation and age influence regulatory attitudes. Sobol indices and conditional inference identify belief patterns and scenario-specific responses across population profiles. We show that awareness of regulation is driven by information-seeking behavior, while support for legal requirements depends strongly on perceived policy adequacy and political alignment. Our approach offers a transparent, data-driven framework for identifying which public segments are most responsive to AI policy initiatives, providing insights to inform risk communication and governance strategies. We illustrate this through a focused analysis of support for AI regulation, quantifying the influence of political ideology, perceived risks, and regulatory awareness under different scenarios.

cs.CY

Deep-SITAR: A SITAR-Based Deep Learning Framework for Growth Curve Modeling via Autoencoders

Several approaches have been developed to capture the complexity and nonlinearity of human growth. One widely used is the Super Imposition by Translation and Rotation (SITAR) model, which has become popular in studies of adolescent growth. SITAR is a shape-invariant mixed-effects model that represents the shared growth pattern of a population using a natural cubic spline mean curve while incorporating three subject-specific random effects -- timing, size, and growth intensity -- to account for variations among individuals. In this work, we introduce a supervised deep learning framework based on an autoencoder architecture that integrates a deep neural network (neural network) with a B-spline model to estimate the SITAR model. In this approach, the encoder estimates the random effects for each individual, while the decoder performs a fitting based on B-splines similar to the classic SITAR model. We refer to this method as the Deep-SITAR model. This innovative approach enables the prediction of the random effects of new individuals entering a population without requiring a full model re-estimation. As a result, Deep-SITAR offers a powerful approach to predicting growth trajectories, combining the flexibility and efficiency of deep learning with the interpretability of traditional mixed-effects models.

stat.ML

Multidimensional Adaptive Penalised Splines with Application to Neurons' Activity Studies

P-spline models have achieved great popularity both in statistical and in applied research. A possible drawback of P-spline is that they assume a smooth transition of the covariate effect across its whole domain. In some practical applications, however, it is desirable and needed to adapt smoothness locally to the data, and adaptive P-splines have been suggested. Yet, the extra flexibility afforded by adaptive P-spline models is obtained at the cost of a high computational burden, especially in a multidimensional setting. Furthermore, to the best of our knowledge, the literature lacks proposals for adaptive P-splines in more than two dimensions. Motivated by the need for analysing data derived from experiments conducted to study neurons' activity in the visual cortex, this work presents a novel locally adaptive anisotropic P-spline model in two (e.g., space) and three (space and time) dimensions. Estimation is based on the recently proposed SOP (Separation of Overlapping Precision matrices) method, which provides the speed we look for. The practical performance of the proposal is evaluated through simulations, and comparisons with alternative methods are reported. In addition to the spatio-temporal analysis of the data that motivated this work, we also discuss an application in two dimensions on the absenteeism of workers.

stat.ME

On the estimation of variance parameters in non-standard generalised linear mixed models: Application to penalised smoothing

We present a novel method for the estimation of variance parameters in generalised linear mixed models. The method has its roots in Harville (1977)'s work, but it is able to deal with models that have a precision matrix for the random-effect vector that is linear in the inverse of the variance parameters (i.e., the precision parameters). We call the method SOP (Separation of Overlapping Precision matrices). SOP is based on applying the method of successive approximations to easy-to-compute estimate updates of the variance parameters. These estimate updates have an appealing form: they are the ratio of a (weighted) sum of squares to a quantity related to effective degrees of freedom. We provide the sufficient and necessary conditions for these estimates to be strictly positive. An important application field of SOP is penalised regression estimation of models where multiple quadratic penalties act on the same regression coefficients. We discuss in detail two of those models: penalised splines for locally adaptive smoothness and for hierarchical curve data. Several data examples in these settings are presented.

stat.ME

Spatio-temporal adaptive penalized splines with application to Neuroscience

Data analysed here derive from experiments conducted to study neurons' activity in the visual cortex of behaving monkeys. We consider a spatio-temporal adaptive penalized spline (P-spline) approach for modelling the firing rate of visual neurons. To the best of our knowledge, this is the first attempt in the statistical literature for locally adaptive smoothing in three dimensions. Estimation is based on the Separation of Overlapping Penalties (SOP) algorithm, which provides the stability and speed we look for.

stat.ME

Fast estimation of multidimensional adaptive P-spline models

A fast and stable algorithm for estimating multidimensional adaptive P-spline models is presented. We call it as Separation of Overlapping Penalties (SOP) as it is an extension of the \textit{Separation of Anisotropic Penalties} (SAP) algorithm. SAP was originally derived for the estimation of the smoothing parameters of a multidimensional tensor product P-spline model with anisotropic penalties.

stat.ME