SearcharxivSearch

arXiv subjects

Roy Cerqueti

Publications and source records attributed to Roy Cerqueti.

At least 19 recordsLinked to original sources

Manipulation testing based on Benford's Law for discrete scores

This paper addresses the problem of running variable manipulation in Regression Discontinuity Designs. Leveraging the observation that manipulation often alters the density balance around the cutoff, we detect these structural imbalances using Benford's Law -a natural statistical regularity widely applied in fraud detection. Our framework serves as a vital precautionary safeguard alongside traditional McCrary-type tests. It eliminates researcher-chosen parameters that can skew outcomes, while delivering a deeper diagnostic breakdown of the density's behavior. Crucially, whereas the classic McCrary test can overlook systemic imbalances due to its rigid symmetric setup, our method separates the data into directional components. This allows researchers to pinpoint the exact origin of a deviation and spot hidden manipulation that standard frameworks fail to capture. To achieve this, we introduce an innovative method for selecting a bandwidth consistent with BL, and construct two distinct, complementary tests using threshold values adapted from Nigrini (2012) that successfully transition the law's application from digits to probabilities. Empirical applications confirm the enhanced protective value of this diagnostic framework.

stat.ME

Z-Dip: a standardized measure for data modality assessment

Detecting multimodality in empirical distributions is a fundamental problem in statistics and data analysis, with applications ranging from clustering to the study of complex systems. In practice, however, assessing departures from unimodality in a consistent and comparable way remains challenging. Widely used methods such as Hartigan and Hartigan's Dip Test illustrate these difficulties, as the interpretation of their statistics depends strongly on sample size, requires calibration to determine significance, and, for large samples, exhibit increasing sensitivity, leading to rejection of unimodality for arbitrarily small deviations from the null. We introduce Z-Dip, a standardized measure of multimodality that addresses these limitations. By treating the Dip statistic as a random variable under the null hypothesis of unimodality and standardizing its observed value, the proposed approach yields scores that are directly comparable across datasets of different sizes. Using simulation-based calibration, we derive a universal decision threshold that closely reproduces classical Dip Test decisions without requiring sample-size-specific adjustments. Extensive validation on simulated data and on more than 88,000 empirical opinion distributions shows near-perfect agreement with the classical Dip Test while providing a more interpretable and comparable measure of modality. Finally, we propose a downsampling-based correction that mitigates residual sensitivity in extremely large samples. Open-source software and reference tables are provided to facilitate practical adoption.

stat.ME

How to Detect Information Voids Using Longitudinal Data from Social Media and Web Searches

The model of the attention economy, where content producers compete for the attention of users, relies on two key forces: information supply and demand. This study leverages the feedback loop between these forces to develop a method for detecting and quantifying information voids, i.e., periods in which little or no reliable information is available on a given topic. Using a case study on COVID-19 vaccines rollout in six European countries, and drawing on data from multiple platforms including Facebook, Google, Twitter, Wikipedia, and online news outlets, we examine how information voids emerge, persist and correlate with a decline in the proportion of high-quality information circulating online. By conceptualising information voids as a specific regime of information spreading, we also quantify their counterpart, information overabundance, which constitute a central component of the current definition of infodemic. We show that information voids are associated with a higher prevalence of misinformation, thus representing problematic hotspots in which individuals are more likely to be misled by low-quality online content. Overall, our findings provide empirical support for the inclusion of information voids in mechanistic explanations of misinformation emergence.

cs.CY

Hybrid Galam--Bass Model for Technology Innovation

This work proposes a hybrid model that combines the Galam model of opinion dynamics with the Bass diffusion model used in technology adoption on Barabasi-Albert complex networks. The main idea is to advance a version of the Bass model that can suitably describe an opinion formation context while introducing irreversible transitions from group B (opponents) to group A (supporters). Moreover, we extend the model to take into account the presence of a charismatic competitor, which fosters conversion back to the old technology. The approach is different from the introduction of a mean field due to the interactions driven by the network structure. Additionally, we introduce the Kolmogorov-Sinai entropy to quantify the system's unpredictability and information loss over time. The results show an increase in the regularity of the trajectories as the preferential attachment parameter increases.

physics.soc-ph

Quantifying Polarization: A Comparative Study of Measures and Methods

Political polarization, a key driver of social fragmentation, has drawn increasing attention for its role in shaping online and offline discourse. Despite significant efforts, accurately measuring polarization within ideological distributions remains a challenge. This study evaluates five widely used polarization measures, testing their strengths and weaknesses with synthetic datasets and a real-world case study on YouTube discussions during the 2020 U.S. Presidential Election. Building on these findings, we present a novel adaptation of Kleinberg's burst detection algorithm to improve mode detection in polarized distributions. By offering both a critical review and an innovative methodological tool, this work advances the analysis of ideological patterns in social media discourse.

cs.CY

Spatially-clustered spatial autoregressive models with application to agricultural market concentration in Europe

In this paper, we present an extension of the spatially-clustered linear regression models, namely, the spatially-clustered spatial autoregression (SCSAR) model, to deal with spatial heterogeneity issues in clustering procedures. In particular, we extend classical spatial econometrics models, such as the spatial autoregressive model, the spatial error model, and the spatially-lagged model, by allowing the regression coefficients to be spatially varying according to a cluster-wise structure. Cluster memberships and regression coefficients are jointly estimated through a penalized maximum likelihood algorithm which encourages neighboring units to belong to the same spatial cluster with shared regression coefficients. Motivated by the increase of observed values of the Gini index for the agricultural production in Europe between 2010 and 2020, the proposed methodology is employed to assess the presence of local spatial spillovers on the market concentration index for the European regions in the last decade. Empirical findings support the hypothesis of fragmentation of the European agricultural market, as the regions can be well represented by a clustering structure partitioning the continent into three-groups, roughly approximated by a division among Western, North Central and Southeastern regions. Also, we detect heterogeneous local effects induced by the selected explanatory variables on the regional market concentration. In particular, we find that variables associated with social, territorial and economic relevance of the agricultural sector seem to act differently throughout the spatial dimension, across the clusters and with respect to the pooled model, and temporal dimension.

stat.ME

Evaluating the effect of viral news on social media engagement

This study examines Facebook and YouTube content from over a thousand news outlets in four European languages from 2018 to 2023, using a Bayesian structural time-series model to evaluate the impact of viral posts. Our results show that most viral events do not significantly increase engagement and rarely lead to sustained growth. The virality effect usually depends on the engagement trend preceding the viral post, typically reversing it. When news emerges unexpectedly, viral events enhances users' engagement, reactivating the collective response process. In contrast, when virality manifests after a sustained growth phase, it represents the final burst of that growth process, followed by a decline in attention. Moreover, quick viral effects fade faster, while slower processes lead to more persistent growth. These findings highlight the transient effect of viral events and underscore the importance of consistent, steady attention-building strategies to establish a solid connection with the user base rather than relying on sudden visibility spikes.

cs.SI

Followers do not dictate the virality of news outlets on social media

Initially conceived for entertainment, social media platforms have profoundly transformed the dissemination of information and consequently reshaped the dynamics of agenda-setting. In this scenario, understanding the factors that capture audience attention and drive viral content is crucial. Employing Gibrat's Law, which posits that an entity's growth rate is unrelated to its size, we examine the engagement growth dynamics of news outlets on social media. Our analysis encloses the Facebook historical data of over a thousand news outlets, encompassing approximately 57 million posts in four European languages from 2008 to the end of 2022. We discover universal growth dynamics according to which news virality is independent of the traditional size or engagement with the outlet. Moreover, our analysis reveals a significant long-term impact of news source reliability on engagement growth, with engagement induced by unreliable sources decreasing over time. We conclude the paper by presenting a statistical model replicating the observed growth dynamics.

cs.SI

A theory of best choice selection through objective arguments grounded in Linear Response Theory concepts

In this paper, we propose how to use objective arguments grounded in statistical mechanics concepts in order to obtain a single number, obtained after aggregation, which would allow to rank "agents", "opinions", ..., all defined in a very broad sense. We aim toward any process which should a priori demand or lead to some consensus in order to attain the presumably best choice among many possibilities. In order to precise the framework, we discuss previous attempts, recalling trivial "means of scores", - weighted or not, Condorcet paradox, TOPSIS, etc. We demonstrate through geometrical arguments on a toy example, with 4 criteria, that the pre-selected order of criteria in previous attempts makes a difference on the final result. However, it might be unjustified. Thus, we base our "best choice theory" on the linear response theory in statistical mechanics: we indicate that one should be calculating correlations functions between all possible choice evaluations, thereby avoiding an arbitrarily ordered set of criteria. We justify the point through an example with 6 possible criteria. Applications in many fields are suggested. Beside, two toy models serving as practical examples and illustrative arguments are given in an Appendix.

physics.soc-ph

Higher order assortativity for directed weighted networks and Markov chains

This paper proposes a new class of assortativity measures for weighted and directed networks. We extend the classical Newman's degree-degree assortativity by considering nodes' attributes different from the degree. Moreover, we propose connections among the nodes through directed paths of length greater than one, thus obtaining higher-order assortativity. We provide an empirical application of these measures for the paradigmatic case of the trade network. Importantly, we show how this global network indicator is strongly related to the autocorrelations of the states of a Markov chain.

physics.soc-ph

Markov Chain Monte Carlo for generating ranked textual data

This paper faces a central theme in applied statistics and information science, which is the assessment of the stochastic structure of rank-size laws in text analysis. We consider the words in a corpus by ranking them on the basis of their frequencies in descending order. The starting point is that the ranked data generated in linguistic contexts can be viewed as the realisations of a discrete states Markov chain, whose stationary distribution behaves according to a discretisation of the best fitted rank-size law. The employed methodological toolkit is Markov Chain Monte Carlo, specifically referring to the Metropolis-Hastings algorithm. The theoretical framework is applied to the rank-size analysis of the hapax legomena occurring in the speeches of the US Presidents. We offer a large number of statistical tests leading to the consistency of our methodological proposal. To pursue our scopes, we also offer arguments supporting that hapaxes are rare (``extreme") events resulting from memory-less-like processes. Moreover, we show that the considered sample has the stochastic structure of a Markov chain of order one. Importantly, we discuss the versatility of the method, which is considered suitable for deducing similar outcomes for other applied science contexts.

stat.ME

Severe testing of Benford's law

Benford's law is often used as a support to critical decisions related to data quality or the presence of data manipulations or even fraud. However, many authors argue that conventional statistical tests will reject the null of data "Benford-ness" if applied in samples of the typical size in this kind of applications, even in the presence of tiny and practically unimportant deviations from Benford's law. Therefore, they suggest using alternative criteria that, however, lack solid statistical foundations. This paper contributes to the debate on the "large $n$" (or "excess power") problem in the context of Benford's law testing. This issue is discussed in relation with the notion of severity testing for goodness of fit tests, with a specific focus on tests for conformity with Benford's law. To do so, we also derive the asymptotic distribution of the mean absolute deviation ($MAD$) statistic as well as an asymptotic standard normal test. Finally, the severity testing principle is applied to six controversial data sets to assess their "Benford-ness".

stat.ME

Rational expectations as a tool for predicting failure of weighted k-out-of-n reliability systems

Here we introduce the idea of using rational expectations, a core concept in economics and finance, as a tool to predict the optimal failure time for a wide class of weighted k-out-of-n reliability systems. We illustrate the concept by applying it to systems which have components with heterogeneous failure times. Depending on the heterogeneous distributions of component failure, we find different measures to be optimal for predicting the failure time of the total system. We give examples of how, as a given system deteriorates over time, one can issue different optimal predictions of system failure by choosing among a set of time-dependent measures.

physics.soc-ph

Tsallis entropy for cross-shareholding network configurations

In this work, we develop the Tsallis entropy approach for examining the cross-shareholding network of companies traded on the Italian stock market. In such a network, the nodes represent the companies, and the links represent the ownership. Within this context, we introduce the out-degree of the nodes -- which represents the diversification -- and the in-degree of them -- capturing the integration. Diversification and integration allow a clear description of the industrial structure formed by the considered companies. The stochastic dependence of diversification and integration is modelled through copulas. We argue that copulas are well suited for modelling the joint distribution. The analysis of the stochastic dependence between integration and diversification by means of the Tsallis entropy gives a crucial information on the reaction of the market structure to the external shocks, - on the basis of some relevant cases of dependence between the considered variables. In this respect, the considered entropy framework provides insights on the relationship between in-degree and out-degree dependence structure and market polarisation or fairness. Moreover, the interpretation of the results in the light of the Tsallis entropy parameter gives relevant suggestions for policymakers who aim at shaping the industrial context for having high polarisation or fair joint distribution of diversification and integration. Furthermore, a discussion of possible parametrisations of the in-degree and out-degree marginal distribution, -- by means of power laws or exponential functions, -- is also carried out. An empirical experiment on a large dataset of Italian companies validates the theoretical framework.

physics.soc-ph

Model-based fuzzy time series clustering of conditional higher moments

This paper develops a new time series clustering procedure allowing for heteroskedasticity, non-normality and model's non-linearity. At this aim, we follow a fuzzy approach. Specifically, considering a Dynamic Conditional Score (DCS) model, we propose to cluster time series according to their estimated conditional moments via the Autocorrelation-based fuzzy C-means (A-FCM) algorithm. The DCS parametric modelling is appealing because of its generality and computational feasibility. The usefulness of the proposed procedure is illustrated using an experiment with simulated data and several empirical applications with financial time series assuming both linear and nonlinear models' specification and under several assumptions about time series density function.

stat.ME

The sooner the better: lives saved by the lockdown during the COVID-19 outbreak. The case of Italy

This paper estimates the effects of non-pharmaceutical interventions - mainly, the lockdown - on the COVID-19 mortality rate for the case of Italy, the first Western country to impose a national shelter-in-place order. We use a new estimator, the Augmented Synthetic Control Method (ASCM), that overcomes some limits of the standard Synthetic Control Method (SCM). The results are twofold. From a methodological point of view, the ASCM outperforms the SCM in that the latter cannot select a valid donor set, assigning all the weights to only one country (Spain) while placing zero weights to all the remaining. From an empirical point of view, we find strong evidence of the effectiveness of non-pharmaceutical interventions in avoiding losses of human lives in Italy: conservative estimates indicate that for each human life actually lost, in the absence of lockdown there would have been on average other 1.15, the policy saved in total 20,400 human lives.

econ.EM

Simple approaches on how to discover promising strategies for efficient enterprise performance, at time of crisis in the case of SMEs : Voronoi clustering and outlier effects perspective

This paper analyzes the connection between innovation activities of companies -- implemented before a financial crisis -- and their performance -- measured after such a time of crisis. Pertinent data about companies listed in the STAR Market Segment of the Italian Stock Exchange is analyzed. Innovation is measured through the level of investments in total tangible and intangible fixed assets in 2006-2007, while performance is captured through growth -- expressed by variations of sales or of total assets, -- profitability -- through ROI or ROS evolution, - and productivity -- through asset turnover or sales/employee in the period 2008-2010. The variables of interest are analyzed and compared through statistical techniques and by adopting a cluster analysis. In particular, a Voronoi tessellation is implemented in a varying centroids framework. In accord with a large part of the literature, we find that the behavior of the performance of the companies is not univocal when they innovate. The statistical outliers are the best cases in order to suggest efficient strategies. In brief, it is found that a positive rate of investments is preferable.

q-fin.ST

Anxiety for the pandemic and trust in financial markets

The COVID-19 pandemic has generated disruptive changes in many fields. Here we focus on the relationship between the anxiety felt by people during the pandemic and the trust in the future performance of financial markets. Precisely, we move from the idea that the volume of Google searches about "coronavirus" can be considered as a proxy of the anxiety and, jointly with the stock index prices, can be used to produce mood indicators -- in terms of pessimism and optimism -- at country level. We analyse the "very high human developed countries" according to the Human Development Index plus China and their respective main stock market indexes. Namely, we propose both a temporal and a global measure of pessimism and optimism and provide accordingly a classification of indexes and countries. The results show the existence of different clusters of countries and markets in terms of pessimism and optimism. Moreover, specific regimes along the time emerge, with an increasing optimism spreading during the mid of June 2020. Furthermore, countries with different government responses to the pandemic have experienced different levels of mood indicators, so that countries with less strict lockdown had a higher level of optimism.

q-fin.ST