SearcharxivSearch

arXiv subjects

Mendeli Vainstein

Publications and source records attributed to Mendeli Vainstein.

3 recordsLinked to original sources

A Shared IPTC Topic Space for Cross-Source Topic Modelling

Comparing topic attention across different media is hindered by a fundamental modelling problem: topic models fitted separately to each corpus produce corpus-specific topic spaces that cannot be aligned directly. This paper presents a reproducible framework that places corpora in a single shared topic space defined by a taxonomy. Discovered topics are obtained with guided BERTopic, scored against the ninety-four IPTC Media Topics' taxonomy topics (level-1) through weighted keyword and target centroids, and then collapsed upward to seventeen IPTC parent topics by a maximum-similarity rule. The framework was developed and selected on a controlled New York Times 2011 corpus through a narrowing sequence: a broad model screen, a focused mapping refinement, a strict finalist comparison, a target-construction ablation, and a threshold calibration. In this corpus, the guided family retained substantially stronger mapped coverage than a zero-shot benchmark under stricter assignment thresholds, a parent-enriched target construction improved both coverage and parent consistency, and coverage declined gradually rather than collapsing as the assignment threshold was tightened. The contribution is an externally anchored method for constructing a shared topic space that enables reproducible cross-source topic comparison.

cs.IR

Exact solution for the Anisotropic Ornstein-Uhlenbeck Process

Active Matter models commonly consider particles with overdamped dynamics subject to a force (speed) with constant modulus and random direction. Some models include also random noise in particle displacement (Wiener process) resulting in a diffusive motion at short time scales. On the other hand, Ornstein-Uhlenbeck processes consider Langevin dynamics for the particle velocity and predict a motion that is not diffusive at short time scales. However, experiments show that migrating cells may present a varying speed as well as a short-time diffusive behavior. While Ornstein-Uhlenbeck processes can describe the varying speed, Active Mater models can explain the short-time diffusive behavior. Isotropic models cannot explain both: short-time diffusion renders instantaneous velocity ill-defined, hence impeding dynamical equations that consider velocity time-derivatives. On the other hand, both models apply for migrating biological cells and must, in some limit, yield the same observable predictions. Here we propose and analytically solve an Anisotropic Ornstein-Uhlenbeck process that considers polarized particles, with a Langevin dynamics for the particle movement in the polarization direction while following a Wiener process for displacement in the orthogonal direction. Our characterization provides a theoretically robust way to compare movement in dimensionless simulations to movement in dimensionful experiments, besides proposing a procedure to deal with inevitable finite precision effects in experiments or simulations.

physics.bio-ph

A Simple Non-Markovian Computational Model of the Statistics of Soccer Leagues: Emergence and Scaling effects

We propose a novel algorithm that outputs the final standings of a soccer league, based on a simple dynamics that mimics a soccer tournament. In our model, a team is created with a defined potential(ability) which is updated during the tournament according to the results of previous games. The updated potential modifies a teams' future winning/losing probabilities. We show that this evolutionary game is able to reproduce the statistical properties of final standings of actual editions of the Brazilian tournament (Brasileirão). However, other leagues such as the Italian and the Spanish tournaments have notoriously non-Gaussian traces and cannot be straightforwardly reproduced by this evolutionary non-Markovian model. A complete understanding of these phenomena deserves much more attention, but we suggest a simple explanation based on data collected in Brazil: Here several teams were crowned champion in previous editions corroborating that the champion typically emerges from random fluctuations that partly preserves the gaussian traces during the tournament. On the other hand, in the Italian and Spanish leagues only a few teams in recent history have won their league tournaments. These leagues are based on more robust and hierarchical structures established even before the beginning of the tournament. For the sake of completeness, we also elaborate a totally Gaussian model (which equalizes the winning, drawing, and losing probabilities) and we show that the scores of the "Brasileirão" cannot be reproduced. Such aspects stress that evolutionary aspects are not superfluous in our modeling. Finally, we analyse the distortions of our model in situations where a large number of teams is considered, showing the existence of a transition from a single to a double peaked histogram of the final classification scores. An interesting scaling is presented for different sized tournaments.

physics.data-an