SearcharxivSearch

arXiv subjects

Juan Sosa

Publications and source records attributed to Juan Sosa.

At least 19 recordsLinked to original sources

Structural Analysis of a Dynamic Multilayer Network via Matrix Autoregressive Models: A Case Study of International Interactions between Countries

Dynamic and multilayer networks have been widely studied separately, but their joint analysis remains comparatively underdeveloped. Because the relational information of a dynamic multilayer network can be represented as a tensor at each time point $t$, each layer can be summarized through a set of structural statistics, yielding a matrix-valued observation and, consequently, a matrix-valued time series. To exploit this structure, we propose the use of matrix autoregressive (MAR) models, which simultaneously characterize temporal dependence across relational layers and structural statistics. We apply this framework to the ICEWS dataset, which records international interactions among countries under four relational domains and therefore naturally defines a dynamic multilayer network. The results indicate that negative verbal interactions (\textit{Verbal-}) play a prominent role in the subsequent structural reconfiguration of the material-interaction layers, while mean strength exhibits the strongest temporal persistence and reciprocity the broadest cross-statistic influence. These findings illustrate the usefulness of MAR models for providing a parsimonious and interpretable characterization of temporal and cross-layer dependence in dynamic multilayer networks.

stat.AP

Statistics in the Age of AI

Artificial intelligence (AI) can automate programming, model fitting, visualization, simulation, literature synthesis, and increasingly sophisticated methodological tasks, but it cannot remove the logical conditions under which data support scientific claims or consequential decisions. We formalize these conditions through statistical warrant, which connects data to a claim through the target, observation regime, assumptions, procedure, uncertainty assessment, validation criterion, loss structure, governance and accountability. No algorithm can consistently recover a target that is not identified by the observation regime without additional information or assumptions. From this principle, we organize the argument around five statements. Questions and targets are integral to statistical methods. Data acquire evidential meaning only through design, provenance, and assumptions. Description, prediction, causal inference, and decision are mathematically distinct tasks. Analytical abundance requires accounting for how analyses are selected, uncertainty across the analytical system, and deployment validation. The statistician's fundamental role is therefore to construct, criticize, and safeguard statistical warrant, including by developing new methodology when existing theory is inadequate. This role requires statistical reasoning and attributable human and institutional responsibility.

stat.OT

Latent space models for networks with nodal multiplicative effects

Latent space models represent network nodes as points in a geometric space, with connection probabilities determined by distances between latent positions under a fixed metric, typically Euclidean, spherical, or hyperbolic. We generalize the classical formulation by introducing nodal multiplicative effects motivated by a local deformation of the latent metric. This modification approximates a conformal deformation of the metric tensor while preserving the logistic predictor and the geometric interpretability of the model, thereby capturing additional structural heterogeneity without altering the global reference geometry. We study the generative behavior of the proposed model through simulation experiments and develop an optimization-based inference scheme derived from a hierarchical Bayesian formulation with parameter regularization. Applications to eight real networks show that the proposed approach increases the generative flexibility of classical latent space models and more accurately reproduces several topological properties observed in real networks.

stat.ME

Uncovering latent territorial structure in ICFES Saber 11 performance with Bayesian multilevel spatial models

This article develops a Bayesian hierarchical framework to analyze academic performance in the 2022 second semester Saber 11 examination in Colombia. Our approach combines multilevel regression with municipal and departmental spatial random effects, and it incorporates Ridge and Lasso regularization priors to compare the contribution of sociodemographic covariates. Inference is implemented in a fully open source workflow using Markov chain Monte Carlo methods, and model behavior is assessed through synthetic data that mirror key features of the observed data. Simulation results indicate that Ridge provides the most balanced performance in parameter recovery, predictive accuracy, and sampling efficiency, while Lasso shows weaker fit and posterior stability, with gains in predictive accuracy under stronger multicollinearity. In the application, posterior rankings show a strong centralization of performance, with higher scores in central departments and lower scores in peripheral territories, and the strongest correlates of scores are student level living conditions, maternal education, access to educational resources, gender, and ethnic background, while spatial random effects capture residual regional disparities. A hybrid Bayesian segmentation based on K means propagates posterior uncertainty into clustering at departmental, municipal, and spatial scales, revealing multiscale territorial patterns consistent with structural inequalities and informing territorial targeting in education policy.

stat.AP

The Colombian legislative process, 2014-2025: networks, topics, and polarization

The legislative output of Colombia's House of Representatives between 2014 and 2025 is analyzed using 4,083 bills. Bipartite networks are constructed between parties and bills, and between representatives and bills, along with their projections, to characterize co-sponsorship patterns, centrality, and influence, and to assess whether political polarization is reflected in legislative collaboration. In parallel, the content of the initiatives is studied through semantic networks based on co-occurrences extracted from short descriptions, and topics by party and period are identified using a stochastic block model for weighted networks, with additional comparison using Latent Dirichlet Allocation. In addition, a Bayesian sociability model is applied to detect terms with robust connectivity and to summarize discursive cores. Overall, the approach integrates relational and semantic structure to describe thematic shifts across administrations, identify influential actors and collectives, and provide a reproducible synthesis that promotes transparency and citizen oversight of the legislative process.

physics.soc-ph

Variational Inference for Fully Bayesian Hierarchical Linear Models

Bayesian hierarchical linear models provide a natural framework to analyze nested and clustered data. Classical estimation with Markov chain Monte Carlo produces well calibrated posterior distributions but becomes computationally expensive in high dimensional or large sample settings. Variational Inference and Stochastic Variational Inference offer faster optimization based alternatives, but their accuracy in hierarchical structures is uncertain when group separation is weak. This paper compares these two paradigms across three model classes, the Linear Regression Model, the Hierarchical Linear Regression Model, and a Clustered Hierarchical Linear Regression Model. Through simulation studies and an application to real data, the results show that variational methods recover global regression effects and clustering structure with a fraction of the computing time, but distort posterior dependence and yield unstable values of information criteria such as WAIC and DIC. The findings clarify when variational methods can serve as practical surrogates for Markov chain Monte Carlo and when their limitations make full Bayesian sampling necessary, and they provide guidance for extending the same variational framework to generalized linear models and other members of the exponential family.

stat.ME

Reyes's I: Measuring Spatial Autocorrelation in Compositions

Compositional observations arise when measurements are recorded as parts of a whole, so that only relative information is meaningful and the natural sample space is the simplex equipped with Aitchison geometry. Despite extensive development of compositional methods, a direct analogue of Moran's \(I\) for assessing spatial autocorrelation in areal compositional data has been lacking. We propose Reyes's \(I\), a Moran type statistic defined through the Aitchison inner product and norm, which is invariant to scale, to permutations of the parts, and to the choice of the \(\operatorname{ilr}\) contrast matrix. Under the randomization assumption, we derive an upper bound, the expected value, and the noncentral second moment, and we describe exact and Monte Carlo permutation procedures for inference. Through simulations covering identical, independent, and spatially correlated compositions under multiple covariance structures and neighborhood definitions, we show that Reyes's \(I\) provides stable behavior, competitive calibration, and improved efficiency relative to a naive alternative based on averaging componentwise Moran statistics. We illustrate practical utility by studying the spatial dependence of a composition measuring COVID-19 severity across Colombian departments during January 2021, documenting significant positive autocorrelation early in the month that attenuates over time.

stat.ME

Peace Sells, But Whose Songs Connect? Bayesian Multilayer Network Analysis of the Big 4 of Thrash Metal

We propose a Bayesian framework for multilayer song similarity networks and apply it to the complete studio discographies of the "Big 4" of thrash metal (Metallica, Slayer, Megadeth, Anthrax). Starting from raw audio, we construct four feature-specific layers (loudness, brightness, tonality, rhythm), augment them with song exogenous information, and represent each layer as a k-nearest neighbor graph. We then fit a family of hierarchical probit models with global and layer-specific baselines, node- and layer-specific sociability effects, dyadic covariates, and alternative forms of latent structure (bilinear, distance-based, and stochastic block communities), comparing increasingly flexible specifications using posterior predictive checks, discrimination and calibration metrics (AUC, Brier score, log-loss), and information criteria (DIC, WAIC). Across all bands, the richest stochastic block specification attains the best predictive performance and posterior predictive fit, while revealing sparse but structured connectivity, interpretable covariate effects (notably album membership and temporal proximity), and latent communities and hubs that cut across albums and eras. Taken together, these results illustrate how Bayesian multilayer network models can help organize high-dimensional audio and text features into coherent, musically meaningful patterns.

stat.ME

The Bayesian Way: Uncertainty, Learning, and Statistical Reasoning

This paper offers a comprehensive introduction to Bayesian inference, combining historical context, theoretical foundations, and core analytical examples. Beginning with Bayes' theorem and the philosophical distinctions between Bayesian and frequentist approaches, we develop the inferential framework for estimation, interval construction, hypothesis testing, and prediction. Through canonical models, we illustrate how prior information and observed data are formally integrated to yield posterior distributions. We also explore key concepts including loss functions, credible intervals, Bayes factors, identifiability, and asymptotic behavior. While emphasizing analytical tractability in classical settings, we outline modern extensions that rely on simulation-based methods and discuss challenges related to prior specification and model evaluation. Though focused on foundational ideas, this paper sets the stage for applying Bayesian methods in contemporary domains such as hierarchical modeling, nonparametrics, and structured applications in time series, spatial data, networks, and political science. The goal is to provide a rigorous yet accessible entry point for students and researchers seeking to adopt a Bayesian perspective in statistical practice.

stat.ME

Network Dynamics and Spatial Shifts in Civilian Targeting: A Stochastic Block Model Analysis of the Colombian Armed Conflict

In this article, we explore how the escalating victimization of civilians during civil wars is mirrored in the fragmented distribution of territorial control, focusing on the Colombian armed conflict. Through an exhaustive characterization of the topology of bipartite and projected networks of municipalities, we describe changes in territorial configurations across different periods between 1978 and 2007. By employing stochastic block models for count data, we show that, during periods dominated by a small set of actors, the networks adopt a centralized node periphery structure, whereas during times of widespread conflict, areas of influence overlap in complex ways. Our findings also suggest the existence of cohesive municipal communities shaped by both geographic proximity and affinities between armed structures, as well as internally dispersed groups with a high likelihood of interaction. As the spatial distribution shifts toward a more fragmented arrangement, the average interaction intensity between communities predicted by the stochastic block model approaches that within communities, indicating a weakening of modular structure and increased inter community connectivity.

stat.AP

Spherical latent space models for social network analysis

This article introduces a spherical latent space model for social network analysis, embedding actors on a hypersphere rather than in Euclidean space as in standard latent space models. The spherical geometry facilitates the representation of transitive relationships and community structure, naturally captures cyclical patterns, and ensures bounded distances, thereby mitigating degeneracy issues common in traditional approaches. Bayesian inference is performed via Markov chain Monte Carlo methods to estimate both latent positions and other model parameters. The approach is demonstrated using two benchmark social network datasets, yielding improved model fit and interpretability relative to conventional latent space models.

stat.ME

Hybrid Bayesian Models for Community Detection with Application to a Colombian Conflict Network

We introduce a flexible Bayesian framework for clustering nodes in undirected binary networks, motivated by the need to uncover structural patterns in complex environments. Building on the stochastic block model, we develop two hybrid extensions: the Class-Distance Model, which governs interaction probabilities through Euclidean distances between cluster-level latent positions, and the Class-Bilinear Model, which captures more complex relational patterns via bilinear interactions. We apply this framework to a novel network derived from the Colombian armed conflict, where municipalities are connected through the co-presence of armed actors, violence, and illicit economies. The resulting clusters align with empirical patterns of territorial control and trafficking corridors, highlighting the models' capacity to recover and explain complex dynamics. Full Bayesian inference is carried out via MCMC under both finite and nonparametric clustering priors. While the main application centers on the Colombian conflict, we also assess model performance using synthetic data as well as other two benchmark datasets.

stat.ME

Constructing the Truth: Text Mining and Linguistic Networks in Public Hearings of Case 03 of the Special Jurisdiction for Peace (JEP)

Case 03 of the Special Jurisdiction for Peace (JEP), focused on the so-called false positives in Colombia, represents one of the most harrowing episodes of the Colombian armed conflict. This article proposes an innovative methodology based on natural language analysis and semantic co-occurrence models to explore, systematize, and visualize narrative patterns present in the public hearings of victims and appearing parties. By constructing skipgram networks and analyzing their modularity, the study identifies thematic clusters that reveal regional and procedural status differences, providing empirical evidence on dynamics of victimization, responsibility, and acknowledgment in this case. This computational approach contributes to the collective construction of both judicial and extrajudicial truth, offering replicable tools for other transitional justice cases. The work is grounded in the pillars of truth, justice, reparation, and non-repetition, proposing a critical and in-depth reading of contested memories.

cs.CL

Parapolitics and Roll-Call Voting in Colombia: A Bayesian Euclidean and Spherical Spatial Analysis

This study presents a Bayesian spatial voting analysis of the Colombian Senate during the 2006-2010 legislative period, leveraging a newly constructed roll-call dataset comprising 147 senators and 136 plenary votes. We estimate legislators' ideal points under two alternative geometric frameworks: A traditional Euclidean model and a circular model that embeds preferences on the unit circle. Both models are implemented using Markov Chain Monte Carlo methods, with the circular specification capturing geodesic distances and von Mises-distributed latent traits. The results reveal a latent structure in voting behavior best characterized not by a conventional left-right ideological continuum but by an opposition-non-opposition alignment. Using Bayesian logistic regression, we further investigate the association between senators' ideal points and their involvement in the para-politics scandal. Findings indicate a significant and robust relationship between political alignment and para-politics implication, suggesting that extralegal influence was systematically related to senators' legislative behavior during this period.

stat.ME

Bayesian Sociality Models: A Scalable and Flexible Alternative for Network Analysis

Bayesian sociality models provide a scalable and flexible alternative for network analysis, capturing degree heterogeneity through actor-specific parameters while mitigating the identifiability challenges of latent space models. This paper develops a comprehensive Bayesian inference framework, leveraging Markov chain Monte Carlo and variational inference to assess their efficiency-accuracy trade-offs. Through empirical and simulation studies, we demonstrate the model's robustness in goodness-of-fit, predictive performance, clustering, and other key network analysis tasks. The Bayesian paradigm further enhances uncertainty quantification and interpretability, positioning sociality models as a powerful and generalizable tool for modern network science.

stat.ME

Bayesian Inference of Geometric Brownian Motion: An Extension with Jumps

This analysis derives the maximum likelihood estimator and applies Bayesian inference to model geometric Brownian motion, incorporating jump diffusion to account for sudden market shifts. The Bayesian approach is implemented using Markov Chain Monte Carlo simulations on S\&P 500 stock data from 2009 to 2014, providing a robust framework for analyzing stock dynamics and forecasting future trends. Exact solutions are obtained for both the standard Geometric Brownian Motion (GBM) model and the GBM model with Poisson jumps. Although both models yield reasonable results and fit the data well, the GBM with Poisson jumps exhibits superior performance, significantly enhancing model fit and capturing more complex market dynamics.

stat.AP

Advances in Bayesian Modeling: Applications and Methods

This paper explores the versatility and depth of Bayesian modeling by presenting a comprehensive range of applications and methods, combining Markov chain Monte Carlo (MCMC) techniques and variational approximations. Covering topics such as hierarchical modeling, spatial modeling, higher-order Markov chains, and Bayesian nonparametrics, the study emphasizes practical implementations across diverse fields, including oceanography, climatology, epidemiology, astronomy, and financial analysis. The aim is to bridge theoretical underpinnings with real-world applications, illustrating the formulation of Bayesian models, elicitation of priors, computational strategies, and posterior and predictive analyses. By leveraging different computational methods, this paper provides insights into model fitting, goodness-of-fit evaluation, and predictive accuracy, addressing computational efficiency and methodological challenges across various datasets and domains.

stat.AP

An unified approach to link prediction in collaboration networks

This article investigates and compares three approaches to link prediction in colaboration networks, namely, an ERGM (Exponential Random Graph Model; Robins et al. 2007), a GCN (Graph Convolutional Network; Kipf and Welling 2017), and a Word2Vec+MLP model (Word2Vec model combined with a multilayer neural network; Mikolov et al. 2013a and Goodfellow et al. 2016). The ERGM, grounded in statistical methods, is employed to capture general structural patterns within the network, while the GCN and Word2Vec+MLP models leverage deep learning techniques to learn adaptive structural representations of nodes and their relationships. The predictive performance of the models is assessed through extensive simulation exercises using cross-validation, with metrics based on the receiver operating characteristic curve. The results clearly show the superiority of machine learning approaches in link prediction, particularly in large networks, where traditional models such as ERGM exhibit limitations in scalability and the ability to capture inherent complexities. These findings highlight the potential benefits of integrating statistical modeling techniques with deep learning methods to analyze complex networks, providing a more robust and effective framework for future research in this field.

stat.AP