SearcharxivSearch

arXiv subjects

Dario Mazzilli

Publications and source records attributed to Dario Mazzilli.

At least 19 recordsLinked to original sources

An exact and fast solution of the inverse Regularized Optimal Transport problem

Optimal transport describes the most efficient way to move mass between two distributions, given a cost matrix for moving mass between each pair of locations. Entropic optimal transport, solved via the Sinkhorn algorithm, is a widely used regularized version of this problem. Its inverse problem asks the opposite question: given an observed transport plan, what cost matrix produced it? This is difficult because the cost is identifiable only up to an additive gauge freedom. Here we show that this freedom can be fixed exactly by a single double-centering operation applied to the observed plan, yielding the true cost matrix in closed form, with no iterative optimization required. When a modest number of true cost entries are known, the same approach lets us jointly estimate the temperature parameter controlling the entropic regularization, together with a diagnostic for the reliability of this estimate. We further show that the method is not specific to the entropic optimal transport, but extends to a broader class of transport models defined by an invertible relation between cost and plan.

cond-mat.stat-mech

Statistical Mechanics of the Sub-Optimal Transport

Statistical mechanics is a powerful framework for analyzing optimization yielding analytical results for matching, optimal transport, and other combinatorial problems. However, these methods typically target the zero-temperature limit, where systems collapse onto optimal configurations, a.k.a. the ground states. Real-world systems often occupy intermediate regimes where entropy and cost minimization genuinely compete, producing configurations that are structured yet sub-optimal. The Sub-Optimal Transport (SOT) model captures this competition through an ensemble of weighted bipartite graphs: a coupling parameter interpolates between entropy-dominated dense configurations and cost-dominated sparse structures. This crossover has been observed numerically but lacked analytical understanding. Here we develop a mean-field theory that characterizes this transition. We show that local fluctuations in Lagrange multipliers become sub-extensive in the thermodynamic limit, reducing the full model with strength constraints to an effective single-constraint problem admitting an exact solution in some intermediate regime. The resulting free energy is analytic in the coupling parameter, confirming a smooth crossover rather than a phase transition. We derive closed-form expressions for thermodynamic observables and weight distributions, validated against numerical simulations. These results establish the first analytical description of the SOT model, extending statistical mechanics methods beyond the zero-temperature regime.

cond-mat.stat-mech

Maximum entropy modeling of Optimal Transport: the sub-optimality regime and the transition from dense to sparse networks

We present a bipartite network model that captures intermediate stages of optimization by blending the Maximum Entropy approach with Optimal Transport. In this framework, the network's constraints define the total mass each node can supply or receive, while an external cost field favors a minimal set of links, driving the system toward a sparse, tree-like structure. By tuning the control parameter, one transitions from uniformly distributed weights to an optimal transport regime in which weights condense onto cost-favorable edges. We quantify this dense-to-sparse transition, showing with numerical analyses that the process does not hinge on specific assumptions about the node-strength or cost distributions. Finite-size analysis confirms that the results persist in the thermodynamic limit. Because the model offers explicit control over the degree of sub-optimality, this approach lends to practical applications in link prediction, network reconstruction, and statistical validation, particularly in systems where partial optimization coexists with other noise-like factors.

cond-mat.stat-mech

Vulnerabilities and capabilities in the EU Automotive industry: Leveraging Input-Output Analysis and Economic Complexity

This paper investigates the structural vulnerabilities and competitive dynamics of the EU27 automotive sector, with a focus on the complexity and the fragmentation of production processes across global value chains. Employing a mixed-methods approach, our analysis integrates input-output tables to quantify the sector's reliance on non-EU economic branches, alongside an economic complexity framework to assess the underlying productive capabilities of European countries in automotive-related industries. The findings indicate an increasing dependency on extra-EU suppliers, particularly China, for critical components such as lithium-ion batteries, which heightens supply chain risks. Currently, Eastern European countries-most notably Poland, Czechia, and Hungary-have enhanced their competitiveness in the production of automotive components, surpassing traditional leaders such as Germany. The paper advances the literature by providing a novel, granular list of 6-digit products within the automotive supply chain and offers new insights into the challenges posed by the ongoing electric mobility transition in the European Union, particularly in relation to electric accumulators.

econ.GN

Follow the money: a startup-based measure of AI exposure across occupations, industries and regions

The integration of artificial intelligence (AI) into the workplace is advancing rapidly, necessitating robust metrics to evaluate its tangible impact on the labour market. Existing measures of AI occupational exposure largely focus on AI's theoretical potential to substitute or complement human labour on the basis of technical feasibility, providing limited insight into actual adoption and offering inadequate guidance for policymakers. To address this gap, we introduce the AI Startup Exposure (AISE) index-a novel metric based on occupational descriptions from O*NET and AI applications developed by startups funded by the Y Combinator accelerator. Our findings indicate that while high-skilled professions are theoretically highly exposed according to conventional metrics, they are heterogeneously targeted by startups. Roles involving routine organizational tasks-such as data analysis and office management-display significant exposure, while occupations involving tasks that are less amenable to AI automation due to ethical or high-stakes, more than feasibility, considerations -- such as judges or surgeons -- present lower AISE scores. By focusing on venture-backed AI applications, our approach offers a nuanced perspective on how AI is reshaping the labour market. It challenges the conventional assumption that high-skilled jobs uniformly face high AI risks, highlighting instead the role of today's AI players' societal desirability-driven and market-oriented choices as critical determinants of AI exposure. Contrary to fears of widespread job displacement, our findings suggest that AI adoption will be gradual and shaped by social factors as much as by the technical feasibility of AI applications. This framework provides a dynamic, forward-looking tool for policymakers and stakeholders to monitor AI's evolving impact and navigate the changing labour landscape.

econ.GN

Structural Change, Employment, and Inequality in Europe: an Economic Complexity Approach

Structural change consists of industrial diversification towards more productive, knowledge intensive activities. However, changes in the productive structure bear inherent links with job creation and income distribution. In this paper, we investigate the consequences of structural change, defined in terms of labour shifts towards more complex industries, on employment growth, wage inequality, and functional distribution of income. The analysis is conducted for European countries using data on disaggregated industrial employment shares over the period 2010-2018. First, we identify patterns of industrial specialisation by validating a country-industry industrial employment matrix using a bipartite weighted configuration model (BiWCM). Secondly, we introduce a country-level measure of labour-weighted Fitness, which can be decomposed in such a way as to isolate a component that identifies the movement of labour towards more complex industries, which we define as structural change. Thirdly, we link structural change to i) employment growth, ii) wage inequality, and iii) labour share of the economy. The results indicate that our structural change measure is associated negatively with employment growth. However, it is also associated with lower income inequality. As countries move to more complex industries, they drop the least complex ones, so the (low-paid) jobs in the least complex sectors disappear. Finally, structural change predicts a higher labour ratio of the economy; however, this is likely to be due to the increase in salaries rather than by job creation.

econ.GN

Economic complexity and the sustainability transition: A review of data, methods, and literature

Economic Complexity (EC) methods have gained increasing popularity across fields and disciplines. In particular, the EC toolbox has proved particularly promising in the study of complex and interrelated phenomena, such as the transition towards a greener economy. Using the EC approach, scholars have been investigating the relationship between EC and sustainability, proposing to identify the distinguishing characteristics of green products and to assess the readiness of productive and technological structures for the sustainability transition. This article proposes to review and summarize the data, methods, and empirical literature that are relevant to the study of the sustainability transition from an EC perspective. We review three distinct but connected blocks of literature on EC and environmental sustainability. First, we survey the evidence linking measures of EC to indicators related to environmental sustainability. Second, we review articles that strive to assess the green competitiveness of productive systems. Third, we examine evidence on green technological development and its connection to non-green knowledge bases. Finally, we summarize the findings for each block and identify avenues for further research in this recent and growing body of empirical literature.

econ.GN

Ranking species in complex ecosystems through nestedness maximization

Identifying the rank of species in a social or ecological network is a difficult task, since the rank of each species is invariably determined by complex interactions stipulated with other species. Simply put, the rank of a species is a function of the ranks of all other species through the adjacency matrix of the network. A common system of ranking is to order species in such a way that their neighbours form maximally nested sets, a problem called nested maximization problem (NMP). Here we show that the NMP can be formulated as an instance of the Quadratic Assignment Problem, one of the most important combinatorial optimization problem widely studied in computer science, economics, and operations research. We tackle the problem by Statistical Physics techniques: we derive a set of self-consistent nonlinear equations whose fixed point represents the optimal rankings of species in an arbitrary bipartite mutualistic network, which generalize the Fitness-Complexity equations widely used in the field of economic complexity. Furthermore, we present an efficient algorithm to solve the NMP that outperforms state-of-the-art network-based metrics and genetic algorithms. Eventually, our theoretical framework may be easily generalized to study the relationship between ranking and network structure beyond pairwise interactions, e.g. in higher-order networks.

cond-mat.stat-mech

Inferring comparative advantage via entropy maximization

We revise the procedure proposed by Balassa to infer comparative advantage, which is a standard tool, in Economics, to analyze specialization (of countries, regions, etc.). Balassa's approach compares the export of a product for each country with what would be expected from a benchmark based on the total volumes of countries and products flows. Based on results in the literature, we show that the implementation of Balassa's idea generates a bias: the prescription of the maximum likelihood used to calculate the parameters of the benchmark model conflicts with the model's definition. Moreover, Balassa's approach does not implement any statistical validation. Hence, we propose an alternative procedure to overcome such a limitation, based upon the framework of entropy maximisation and implementing a proper test of hypothesis: the `key products' of a country are, now, the ones whose production is significantly larger than expected, under a null-model constraining the same amount of information employed by Balassa's approach. What we found is that countries diversification is always observed, regardless of the strictness of the validation procedure. Besides, the ranking of countries' fitness is only partially affected by the details of the validation scheme employed for the analysis while large differences are found to affect the rankings of products Complexities. The routine for implementing the entropy-based filtering procedures employed here is freely available through the official Python Package Index PyPI.

cs.SI

Equivalence between the Fitness-Complexity and the Sinkhorn-Knopp algorithms

We uncover the connection between the Fitness-Complexity algorithm, developed in the economic complexity field, and the Sinkhorn-Knopp algorithm, widely used in diverse domains ranging from computer science and mathematics to economics. Despite minor formal differences between the two methods, both converge to the same fixed-point solution up to normalization. The discovered connection allows us to derive a rigorous interpretation of the Fitness and the Complexity metrics as the potentials of a suitable energy function. Under this interpretation, high-energy products are unfeasible for low-fitness countries, which explains why the algorithm is effective at displaying nested patterns in bipartite networks. We also show that the proposed interpretation reveals the scale invariance of the Fitness-Complexity algorithm, which has practical implications for the algorithm's implementation in different datasets. Further, analysis of empirical trade data under the new perspective reveals three categories of countries that might benefit from different development strategies.

econ.GN

Effective submodularity of influence maximization on temporal networks

We study influence maximization on temporal networks. This is a special setting where the influence function is not submodular, and there is no optimality guarantee for solutions achieved via greedy optimization. We perform an exhaustive analysis on both real and synthetic networks. We show that the influence function of randomly sampled sets of seeds often violates the necessary conditions for submodularity. However, when sets of seeds are selected according to the greedy optimization strategy, the influence function behaves effectively as a submodular function. Specifically, violations of the necessary conditions for submodularity are never observed in real networks, and only rarely in synthetic ones. The direct comparison with exact solutions obtained via brute-force search indicate that the greedy strategy provides approximate solutions that are well within the optimality gap guaranteed for strictly submodular functions. Greedy optimization appears therefore an effective strategy for the maximization of influence on temporal networks.

physics.soc-ph

A Bayesian approach to translators' reliability assessment

Translation Quality Assessment (TQA) is a process conducted by human translators and is widely used, both for estimating the performance of (increasingly used) Machine Translation, and for finding an agreement between translation providers and their customers. While translation scholars are aware of the importance of having a reliable way to conduct the TQA process, it seems that there is limited literature that tackles the issue of reliability with a quantitative approach. In this work, we consider the TQA as a complex process from the point of view of physics of complex systems and approach the reliability issue from the Bayesian paradigm. Using a dataset of translation quality evaluations (in the form of error annotations), produced entirely by the Professional Translation Service Provider Translated SRL, we compare two Bayesian models that parameterise the following features involved in the TQA process: the translation difficulty, the characteristics of the translators involved in producing the translation, and of those assessing its quality - the reviewers. We validate the models in an unsupervised setting and show that it is possible to get meaningful insights into translators even with just one review per translation; subsequently, we extract information like translators' skills and reviewers' strictness, as well as their consistency in their respective roles. Using this, we show that the reliability of reviewers cannot be taken for granted even in the case of expert translators: a translator's expertise can induce a cognitive bias when reviewing a translation produced by another translator. The most expert translators, however, are characterised by the highest level of consistency, both in translating and in assessing the translation quality.

cs.CL

Universality, criticality and complexity of information propagation in social media

Information avalanches in social media are typically studied in a similar fashion as avalanches of neuronal activity in the brain. Whereas a large body of literature reveals substantial agreement about the existence of a unique process characterizing neuronal activity across organisms, the dynamics of information in online social media is far less understood. Statistical laws of information avalanches are found in previous studies to be not robust across systems, and radically different processes are used to represent plausible driving mechanisms for information propagation. Here, we analyze almost 1 billion time-stamped events collected from a multitude of online platforms -- including Telegram, Twitter and Weibo -- over observation windows longer than 10 years to show that the propagation of information in social media is a universal and critical process. Universality arises from the observation of identical macroscopic patterns across platforms, irrespective of the details of the specific system at hand. Critical behavior is deduced from the power-law distributions, and corresponding hyperscaling relations, characterizing size and duration of avalanches of information. Neuronal activity may be modeled as a simple contagion process, where only a single exposure to activity may be sufficient for its diffusion. On the contrary, statistical testing on our data indicates that a mixture of simple and complex contagion, where involvement of an individual requires exposure from multiple acquaintances, characterizes the propagation of information in social media. We show that the complexity of the process is correlated with the semantic content of the information that is propagated. Conversational topics about music, movies and TV shows tend to propagate as simple contagion processes, whereas controversial discussions on political/societal themes obey the rules of complex contagion.

physics.soc-ph

Percolation theory of self-exciting temporal processes

We investigate how the properties of inhomogeneous patterns of activity, appearing in many natural and social phenomena, depend on the temporal resolution used to define individual bursts of activity. To this end, we consider time series of microscopic events produced by a self-exciting Hawkes process, and leverage a percolation framework to study the formation of macroscopic bursts of activity as a function of the resolution parameter. We find that the very same process may result in different distributions of avalanche size and duration, which are understood in terms of the competition between the 1D percolation and the branching process universality class. Pure regimes for the individual classes are observed at specific values of the resolution parameter corresponding to the critical points of the percolation diagram. A regime of crossover characterized by a mixture of the two universal behaviors is observed in a wide region of the diagram. The hybrid scaling appears to be a likely outcome for an analysis of the time series based on a reasonably chosen, but not precisely adjusted, value of the resolution parameter.

physics.soc-ph

Combinatorial approach to spreading processes on networks

Stochastic spreading models defined on complex network topologies are used to mimic the diffusion of diseases, information, and opinions in real-world systems. Existing theoretical approaches to the characterization of the models in terms of microscopic configurations rely on some approximation of independence among dynamical variables, thus introducing a systematic bias in the prediction of the ground-truth dynamics. Here, we develop a combinatorial framework based on the approximation that spreading may occur only along the shortest paths connecting pairs of nodes. The approximation overestimates dynamical correlations among node states and leads to biased predictions. Systematic bias is, however, pointing in the opposite direction of existing approximations. We show that the combination of the two biased approaches generates predictions of the ground-truth dynamics that are more accurate than the ones given by the two approximations if used in isolation. We further take advantage of the combinatorial approximation to characterize theoretical properties of some inference problems, and show that the reconstruction of microscopic configurations is very sensitive to both the place where and the time when partial knowledge of the system is acquired.

physics.soc-ph

Influence maximization on temporal networks

We consider the optimization problem of seeding a spreading process on a temporal network so that the expected size of the resulting outbreak is maximized. We frame the problem for a spreading process following the rules of the susceptible-infected-recovered model with temporal scale equal to the one characterizing the evolution of the network topology. We perform a systematic analysis based on a corpus of 12 real-world temporal networks and quantify the performance of solutions to the influence maximization problem obtained using different level of information about network topology and dynamics. We find that having perfect knowledge of the network topology but in a static and/or aggregated form is not helpful in solving the influence maximization problem effectively. Knowledge, even if partial, of the early stages of the network dynamics appears instead essential for the identification of quasioptimal sets of influential spreaders.

physics.soc-ph

A new and stable estimation method of country economic fitness and product complexity

We present a new metric estimating fitness of countries and complexity of products by exploiting a non-linear non-homogeneous map applied to the publicly available information on the goods exported by a country. The non homogeneous terms guarantee both convergence and stability. After a suitable rescaling of the relevant quantities, the non homogeneous terms are eventually set to zero so that this new metric is parameter free. This new map almost reproduces the results of the original homogeneous metrics already defined in literature and allows for an approximate analytic solution in case of actual binarized matrices based on the Revealed Comparative Advantage (RCA) indicator. This solution is connected with a new quantity describing the neighborhood of nodes in bipartite graphs, representing in this work the relations between countries and exported products. Moreover, we define the new indicator of country net-efficiency quantifying how a country efficiently invests in capabilities able to generate innovative complex high quality products. Eventually, we demonstrate analytically the local convergence of the algorithm involved.

econ.GN

Economic Complexity: "Buttarla in caciara" vs a constructive approach

This note is a contribution to the debate about the optimal algorithm for Economic Complexity that recently appeared on ArXiv [1, 2] . The authors of [2] eventually agree that the ECI+ algorithm [1] consists just in a renaming of the Fitness algorithm we introduced in 2012, as we explicitly showed in [3]. However, they omit any comment on the fact that their extensive numerical tests claimed to demonstrate that the same algorithm works well if they name it ECI+, but not if its name is Fitness. They should realize that this eliminates any credibility to their numerical methods and therefore also to their new analysis, in which they consider many algorithms [2]. Since by their own admission the best algorithm is the Fitness one, their new claim became that the search for the best algorithm is pointless and all algorithms are alike. This is exactly the opposite of what they claimed a few days ago and it does not deserve much comments. After these clarifications we also present a constructive analysis of the status of Economic Complexity, its algorithms, its successes and its perspectives. For us the discussion closes here, we will not reply to further comments.

econ.GN