SearcharxivSearch

arXiv subjects

Lucio Barabesi

Publications and source records attributed to Lucio Barabesi.

At least 19 recordsLinked to original sources

Exploring the Shape of Economics: A Multilayer Network Analysis of Social Communities and Intellectual Similarity Among Journals Before and After the 2008 Financial Crisis

This paper develops a multilayer network approach for exploring the evolution of scientific disciplines, using the case of economics before and after the 2008 global financial crisis as a large-scale empirical testing ground. The units of analysis are journals, linked by social and intellectual relationships. The analysis covers all journals indexed in EconLit across three years (2006, 2012 and 2019). In the most recent year (2019), the dataset includes 909 journals, over 30,000 editorial board members, more than 260,000 authors, 134,000 articles, and nearly 2 million cited references. For each period, we model journals as connected in a four-layer multiplex network: the social relationships are based on shared editors (interlocking editorship) and shared authors (interlocking authorship), while the intellectual ones are based on shared references (bibliographic coupling) and textual similarity between articles. These four layers are integrated using Similarity Network Fusion to produce unified similarity networks from which journal communities are identified. Comparing the field across the three periods reveals a high degree of structural continuity. Although research topics changed after the crisis, the fundamental social and intellectual relationships among journals remained remarkably stable. A major result of the analysis is that editorial networks play the dominant role in shaping hierarchies and legitimize knowledge production within the discipline. Whether this finding holds in other scientific disciplines remains an open question for future research.

econ.GN

Robust inference under Benford's law

We address the task of identifying anomalous observations by analyzing digits under the lens of Benford's law. Motivated by the crucial objective of providing reliable statistical analysis of customs declarations, we answer one major and still open question: How can we detect the behavior of operators who are aware of the prevalence of the Benford's pattern in the digits of regular observations and try to manipulate their data in such a way that the same pattern also holds after data fabrication? This challenge arises from the ability of highly skilled and strategically minded manipulators in key organizational positions or criminal networks to exploit statistical knowledge and evade detection. For this purpose, we write a specific contamination model for digits, obtain new relevant distributional results and derive appropriate goodness-of-fit statistics for the considered adversarial testing problem. Along our path, we also unveil the peculiar relationship between two simple conformance tests based on the distribution of the first digit. We show the empirical properties of the proposed tests through a simulation exercise and application to data from international trade transactions. Although we cannot claim that our results are able to anticipate data fabrication with certainty, they surely point to situations where more substantial controls are needed. Furthermore, our work can reinforce trust in data integrity in many critical domains where mathematically informed misconduct is suspected.

stat.ME

Goodness-of-fit test for count distributions with finite second moment

A goodness-of-fit test for one-parameter count distributions with finite second moment is proposed. The test statistic is derived from the $L^1$ distance of a function of the probability generating function of the model under the null hypothesis and that of the random variable actually generating data, when the latter belongs to a suitable wide class of alternatives. The test statistic has a rather simple form and it is asymptotically normally distributed under the null hypothesis, allowing a straightforward implementation of the test. Moreover, the test is consistent for alternative distributions belonging to the class, but also for all the alternative distributions whose probability of zero is different from that under the null hypothesis. Thus, the use of the test is proposed and investigated also for alternatives not in the class. The finite-sample properties of the test are assessed by means of an extensive simulation study.

math.ST

Estimation and goodness-of-fit testing for non-negative random variables with explicit Laplace transform

Many flexible families of positive random variables exhibit non-closed forms of the density and distribution functions and this feature is considered unappealing for modelling purposes. However, such families are often characterized by a simple expression of the corresponding Laplace transform. Relying on the Laplace transform, we propose to carry out parameter estimation and goodness-of-fit testing for a general class of non-standard laws. We suggest a novel data-driven inferential technique, providing parameter estimators and goodness-of-fit tests, whose large-sample properties are derived. The implementation of the method is specifically considered for the positive stable and Tweedie distributions. A Monte Carlo study shows good finite-sample performance of the proposed technique for such laws.

math.ST

Fine-grained classification of journal articles by relying on multiple layers of information through similarity network fusion: the case of the Cambridge Journal of Economics

In order to explore the suitability of a fine-grained classification of journal articles by exploiting multiple sources of information, articles are organized in a two-layer multiplex. The first layer conveys similarities based on the full-text of articles, and the second similarities based on cited references. The information of the two layers are only weakly associated. The Similarity Network Fusion process is adopted to combine the two layers into a new single-layer network. A clustering algorithm is applied to the fused network and the classification of articles is obtained. In order to evaluate its coherence, this classification is compared with the ones obtained by applying the same algorithm to each of two layers. Moreover, the classification obtained for the fused network is also compared with the classifications obtained when the layers of information are integrated using different methods available in literature. In the case of the Cambridge Journal of Economics, Similarity Network Fusion appears to be the best option. Moreover, the achieved classification appears to be fine-grained enough to represent the extreme heterogeneity characterizing the contributions published in the journal.

cs.DL

Similarity network aggregation for the analysis of glacier ecosystems

The synthesis of information deriving from complex networks is a topic receiving increasing relevance in ecology and environmental sciences. In particular, the aggregation of multilayer networks, i.e. network structures formed by multiple interacting networks (the layers), constitutes a fast-growing field. In several environmental applications, the layers of a multilayer network are modelled as a collection of similarity matrices describing how similar pairs of biological entities are, based on different types of features (e.g. biological traits). The present paper first discusses two main techniques for combining the multi-layered information into a single network (the so-called monoplex), i.e. Similarity Network Fusion (SNF) and Similarity Matrix Average (SMA). Then, the effectiveness of the two methods is tested on a real-world dataset of the relative abundance of microbial species in the ecosystems of nine glaciers (four glaciers in the Alps and five in the Andes). A preliminary clustering analysis on the monoplexes obtained with different methods shows the emergence of a tightly connected community formed by species that are typical of cryoconite holes worldwide. Moreover, the weights assigned to different layers by the SMA algorithm suggest that two large South American glaciers (Exploradores and Perito Moreno) are structurally different from the smaller glaciers in both Europe and South America. Overall, these results highlight the importance of integration methods in the discovery of the underlying organizational structure of biological entities in multilayer ecological networks.

cs.SI

Similarity matrix average for aggregating multiplex networks

We introduce a methodology based on averaging similarity matrices with the aim of integrating the layers of a multiplex network into a single monoplex network. Multiplex networks are adopted for modelling a wide variety of real-world frameworks, such as multi-type relations in social, economic and biological structures. More specifically, multiplex networks are used when relations of different nature (layers) arise between a set of elements from a given population (nodes). A possible approach for investigating multiplex networks consists in aggregating the different layers in a single network (monoplex) which is a valid representation -- in some sense -- of all the layers. In order to obtain such an aggregated network, we propose a theoretical approach -- along with its practical implementation -- which stems on the concept of similarity matrix average. This methodology is finally applied to a multiplex similarity network of statistical journals, where the three considered layers express the similarity of the journals based on co-citations, common authors and common editors, respectively.

physics.soc-ph

Tempered positive Linnik processes and their representations

This paper analyzes various classes of processes associated with the tempered positive Linnik (TPL) distribution. We provide several subordinated representations of TPL Lévy processes and in particular establish a stochastic self-similarity property with respect to negative binomial subordination. In finite activity regimes we show that the explicit compound Poisson representations gives rise to innovations following Mittag-Leffler type laws which are apparently new. We characterize two time-inhomogeneous TPL processes, namely the Ornstein-Uhlenbeck (OU) Lévy-driven processes with stationary distribution and the additive process determined by a TPL law. We finally illustrate how the properties studied come together in a multivariate TPL Lévy framework based on a novel negative binomial mixing methodology. Some potential applications are outlined in the contexts of statistical anti-fraud and financial modelling.

math.PR

Similarity network fusion for scholarly journals

This paper explores intellectual and social proximity among scholarly journals by using network fusion techniques. Similarities among journals are initially represented by means of a three-layer network based on co-citations, common authors and common editors. The information contained in the three layers is then combined by building a fused similarity network. The fusion consists in an unsupervised process that exploits the structural properties of the layers. Subsequently, partial distance correlations are adopted for measuring the contribution of each layer to the structure of the fused network. Finally, the community morphology of the fused network is explored by using modularity. In the three fields considered (i.e. economics, information and library sciences and statistics) the major contribution to the structure of the fused network arises from editors. This result suggests that the role of editors as gatekeepers of journals is the most relevant in defining the boundaries of scholarly communities. In information and library sciences and statistics, the clusters of journals reflect sub-field specializations. In economics, clusters of journals appear to be better interpreted in terms of alternative methodological approaches. Thus, the graphs representing the clusters of journals in the fused network are powerful instruments for exploring research fields.

cs.DL

On the agreement between bibliometrics and peer review: evidence from the Italian research assessment exercises

This paper appraises the concordance between bibliometrics and peer review, by drawing evidence from the data of two experiments realized by the Italian governmental agency for research evaluation. The experiments were performed for validating the dual system of evaluation, consisting in the interchangeable use of bibliometyrics and peer review, adopted by the agency in the research assessment exercises. The two experiments were based on stratified random samples of journal articles. Each article was scored by bibliometrics and by peer review. The degree of concordance between the two evaluations is then computed. The correct setting of the experiments is defined by developing the design-based estimation of the Cohen's kappa coefficient and some testing procedures for assessing the homogeneity of missing proportions between strata. The results of both experiments show that for each research areas of hard sciences, engineering and life sciences, the degree of agreement between bibliometrics and peer review is -- at most -- weak at an individual article level. Thus, the outcome of the experiments does not validate the use of the dual system of evaluation in the Italian research assessments. More in general, the very weak concordance indicates that metrics should not replace peer review at the level of individual article. Hence, the use of the dual system of evaluation for reducing costs might introduce unknown biases in a research assessment exercise.

stat.AP

Intellectual and social similarity among scholarly journals: an exploratory comparison of the networks of editors, authors and co-citations

This paper explores, by using suitable quantitative techniques, to what extent the intellectual proximity among scholarly journals is also a proximity in terms of social communities gathered around the journals. Three fields are considered: statistics, economics and information and library sciences. Co-citation networks (CC) represent the intellectual proximity among journals. The academic communities around the journals are represented by considering the networks of journals generated by authors writing in more than one journal (interlocking authorship: IA), and the networks generated by scholars sitting in the editorial board of more than one journal (interlocking editorship: IE). For comparing the whole structure of the networks, the dissimilarity matrices are considered. The CC, IE and IA networks appear to be correlated for the three fields. The strongest correlations is between CC and IA for the three fields. Lower and similar correlations are obtained for CC and IE, and for IE and IA. The CC, IE and IA networks are then partitioned in communities. Information and library sciences is the field where communities are more easily detectable, while the most difficult field is economics. The degrees of association among the detected communities show that they are not independent. For all the fields, the strongest association is between CC and IA networks; the minimum level of association is between IE and CC. Overall, these results indicate that the intellectual proximity is also a proximity among authors and among editors of the journals. Thus, the three maps of editorial power, intellectual proximity and authors communities tell similar stories.

cs.SI

The tempered discrete Linnik distribution

A tempered version of the discrete Linnik distribution is introduced in order to obtain integer-valued distribution families connected to stable laws. The proposal constitutes a generalization of the well-known Poisson-Tweedie law, which is actually a tempered discrete stable law. The features of the new tempered discrete Linnik distribution are explored by providing a series of identities in law - which describe its genesis in terms of mixture and compound Poisson law, as well as in terms of mixture discrete stable law. A manageable expression of the corresponding probability function is also provided and several special cases are analysed.

math.ST

Crossing the hurdle: the determinants of individual scientific performance

An original cross sectional dataset referring to a medium sized Italian university is implemented in order to analyze the determinants of scientific research production at individual level. The dataset includes 942 permanent researchers of various scientific sectors for a three year time span (2008 - 2010). Three different indicators - based on the number of publications or citations - are considered as response variables. The corresponding distributions are highly skewed and display an excess of zero - valued observations. In this setting, the goodness of fit of several Poisson mixture regression models are explored by assuming an extensive set of explanatory variables. As to the personal observable characteristics of the researchers, the results emphasize the age effect and the gender productivity gap, as previously documented by existing studies. Analogously, the analysis confirm that productivity is strongly affected by the publication and citation practices adopted in different scientific disciplines. The empirical evidence on the connection between teaching and research activities suggests that no univocal substitution or complementarity thesis can be claimed: a major teaching load does not affect the odds to be a non-active researcher and does not significantly reduce the number of publications for active researchers. In addition, new evidence emerges on the effect of researchers administrative tasks, which seem to be negatively related with researcher's productivity, and on the composition of departments. Researchers' productivity is apparently enhanced by operating in department filled with more administrative and technical staff, and it is not significantly affected by the composition of the department in terms of senior or junior researchers.

physics.soc-ph

A functional derivative useful for the linearization of inequality indexes in the design-based framework

Linearization methods are customarily adopted in sampling surveys to obtain approximated variance formulae for estimators of nonlinear functions of finite population totals - such as ratios, correlation coefficients or measures of income inequality - which can be usually rephrased in terms of statistical functionals. In the present paper, by considering the Deville (1991) approach stemming on the concept of design-based influence curve, we provide a general result for linearizing large families of inequality indexes. As an example, the achievement is applied to the Gini, the Amato, the Zenga and the Atkinson indexes, respectively.

stat.ME

A note on a universal random variate generator for integer-valued random variables

A universal generator for integer-valued square-integrable random variables is introduced. The generator relies on a rejection technique based on a generalization of the inversion formula for integer-valued random variables. The proposal gives rise to a simple algorithm which may be implemented in a few code lines and which may show good performance when the classical families of distributions - such as the Poisson and the Binomial - are considered. In addition, the method is suitable for the computer generation of integer-valued random variables which display closed-form characteristic functions, but do not possess a probability function expressible in a simple analytical way. As an example of such a framework, an application to the Poisson-Tweedie distribution is provided.

stat.CO

Statistical inference on the h-index with an application to top-scientist performance

Despite the huge amount of literature on h-index, few papers have been devoted to the statistical analysis of h-index when a probabilistic distribution is assumed for citation counts. The present contribution relies on showing the available inferential techniques, by providing the details for proper point and set estimation of the theoretical h-index. Moreover, some issues on simultaneous inference - aimed to produce suitable scholar comparisons - are carried out. Finally, the analysis of the citation dataset for the Nobel Laureates (in the last five years) and for the Fields medallists (from 2002 onward) is proposed.

stat.AP

Properties of design-based estimation under stratified spatial sampling with application to canopy coverage estimation

The estimation of the total of an attribute defined over a continuous planar domain is required in many applied settings, such as the estimation of canopy coverage in the Monterano Nature Reserve in Italy. If the design-based approach is considered, the scheme for the placement of the sample sites over the domain is fundamental in order to implement the survey. In real situations, a commonly adopted scheme is based on partitioning the domain into suitable strata, in such a way that a single sample site is uniformly placed (i.e., selected with uniform probability density) in each stratum and sample sites are independently located. Under mild conditions on the function representing the target attribute, it is shown that this scheme gives rise to an unbiased spatial total estimator which is "superefficient" with respect to the estimator based on the uniform placement of independent sample sites over the domain. In addition, the large-sample normality of the estimator is proven and variance estimation issues are discussed.

stat.AP

Statistical analysis of the Hirsch Index

The Hirsch index (commonly referred to as h-index) is a bibliometric indicator which is widely recognized as effective for measuring the scientific production of a scholar since it summarizes size and impact of the research output. In a formal setting, the h-index is actually an empirical functional of the distribution of the citation counts received by the scholar. Under this approach, the asymptotic theory for the empirical h-index has been recently exploited when the citation counts follow a continuous distribution and, in particular, variance estimation has been considered for the Pareto-type and the Weibull-type distribution families. However, in bibliometric applications, citation counts display a distribution supported by the integers. Thus, we provide general properties for the empirical h-index under the small- and large-sample settings. In addition, we also introduce consistent nonparametric variance estimation, which allows for the implemention of large-sample set estimation for the theoretical h-index.

math.ST