SearcharxivSearch

arXiv subjects

Andrea Tacchella

Publications and source records attributed to Andrea Tacchella.

15 recordsLinked to original sources

Anticipating Innovation Using Large Language Models

Forecasting innovation, intended as the emergence of new technological combinations, is a fundamental challenge for science and policy. We show that forthcoming combinations leave an early trace in the collective language of patents, with predictive signals detectable even decades in advance. We show that signal is not attributable to any single inventor, but emerges as a collective shift in how technologies are described across thousands of patents. To this end, we introduce TechToken, a transformer-based model that treats technologies, classified by International Patent Classification codes, as words in its vocabulary, learning the language of technologies by embedding these codes during fine-tuning. We define context similarity between code embeddings as a measure of linguistic convergence and show that it accurately predicts first technological combinations. TechToken also improves general representation quality, outperforming state-of-the-art models across different patent-related tasks.

cs.CL

Product-level value chains from firm data: mapping trophic levels into economic growth

We reconstruct a product-level input-output network based on firm-level import-export data of Italian firms. We show that the network has a statistically significant, yet nuanced trophic structure, which is evident at the product level but is lost when the classification is coarse-grained. This detailed value chain allows us to characterize the trophic distance between inputs and outputs of single firms, and to derive a coherent picture at the sector level, finding that sectors such as weapons and vehicles are the ones with the largest increase in downstreamness between their inputs and their outputs. Our measure of downstreamness at the product level can be used to derive country-level indicators that characterize industrial strategies and capabilities and act as predictors of economic growth. With respect to the standard input/output analysis, we show that the fine-grained structure is qualitatively different from what can be observed using sector-level data. We finally prove that, even if we leverage exclusively data from Italian firms, the metrics that we derive are predictive at the country level and capture a significant description of the input-output relations of global value chains.

physics.soc-ph

Comparative Analysis of Technological Fitness and Coherence at different geographical scales

Debates over the trade-offs between specialization and diversification have long intrigued scholars and policymakers. Specialization can amplify an economy by concentrating on core strengths, while diversification reduces vulnerability by distributing investments across multiple sectors. In this paper, we use patent data and the framework of Economic Complexity to investigate how the degree of technological specialization and diversification affects economic development at different scales: metropolitan areas, regions and countries. We examine two Economic Complexity indicators. Technological Fitness assesses an economic player's ability to diversify and generate sophisticated technologies, while Technological Coherence quantifies the degree of specialization by measuring the similarity among technologies within an economic player's portfolio. Our results indicate that a high degree of Technological Coherence is associated with increased economic growth only at the metropolitan area level, while its impact turns negative at larger scales. In contrast, Technological Fitness shows a U-shaped relationship with a positive effect in metropolitan areas, a negative influence at the regional level, and again a positive effect at the national level. These findings underscore the complex interplay between technological specialization and diversification across geographical scales. Understanding these distinctions can inform policymakers and stakeholders in developing tailored strategies for technological advancement and economic growth.

physics.soc-ph

Vulnerabilities and capabilities in the EU Automotive industry: Leveraging Input-Output Analysis and Economic Complexity

This paper investigates the structural vulnerabilities and competitive dynamics of the EU27 automotive sector, with a focus on the complexity and the fragmentation of production processes across global value chains. Employing a mixed-methods approach, our analysis integrates input-output tables to quantify the sector's reliance on non-EU economic branches, alongside an economic complexity framework to assess the underlying productive capabilities of European countries in automotive-related industries. The findings indicate an increasing dependency on extra-EU suppliers, particularly China, for critical components such as lithium-ion batteries, which heightens supply chain risks. Currently, Eastern European countries-most notably Poland, Czechia, and Hungary-have enhanced their competitiveness in the production of automotive components, surpassing traditional leaders such as Germany. The paper advances the literature by providing a novel, granular list of 6-digit products within the automotive supply chain and offers new insights into the challenges posed by the ongoing electric mobility transition in the European Union, particularly in relation to electric accumulators.

econ.GN

Follow the money: a startup-based measure of AI exposure across occupations, industries and regions

The integration of artificial intelligence (AI) into the workplace is advancing rapidly, necessitating robust metrics to evaluate its tangible impact on the labour market. Existing measures of AI occupational exposure largely focus on AI's theoretical potential to substitute or complement human labour on the basis of technical feasibility, providing limited insight into actual adoption and offering inadequate guidance for policymakers. To address this gap, we introduce the AI Startup Exposure (AISE) index-a novel metric based on occupational descriptions from O*NET and AI applications developed by startups funded by the Y Combinator accelerator. Our findings indicate that while high-skilled professions are theoretically highly exposed according to conventional metrics, they are heterogeneously targeted by startups. Roles involving routine organizational tasks-such as data analysis and office management-display significant exposure, while occupations involving tasks that are less amenable to AI automation due to ethical or high-stakes, more than feasibility, considerations -- such as judges or surgeons -- present lower AISE scores. By focusing on venture-backed AI applications, our approach offers a nuanced perspective on how AI is reshaping the labour market. It challenges the conventional assumption that high-skilled jobs uniformly face high AI risks, highlighting instead the role of today's AI players' societal desirability-driven and market-oriented choices as critical determinants of AI exposure. Contrary to fears of widespread job displacement, our findings suggest that AI adoption will be gradual and shaped by social factors as much as by the technical feasibility of AI applications. This framework provides a dynamic, forward-looking tool for policymakers and stakeholders to monitor AI's evolving impact and navigate the changing labour landscape.

econ.GN

Which products activate a product? An explainable machine learning approach

Tree-based machine learning algorithms provide the most precise assessment of the feasibility for a country to export a target product given its export basket. However, the high number of parameters involved prevents a straightforward interpretation of the results and, in turn, the explainability of policy indications. In this paper, we propose a procedure to statistically validate the importance of the products used in the feasibility assessment. In this way, we are able to identify which products, called explainers, significantly increase the probability to export a target product in the near future. The explainers naturally identify a low dimensional representation, the Feature Importance Product Space, that enhances the interpretability of the recommendations and provides out-of-sample forecasts of the export baskets of countries. Interestingly, we detect a positive correlation between the complexity of a product and the complexity of its explainers.

econ.GN

A Bayesian approach to translators' reliability assessment

Translation Quality Assessment (TQA) is a process conducted by human translators and is widely used, both for estimating the performance of (increasingly used) Machine Translation, and for finding an agreement between translation providers and their customers. While translation scholars are aware of the importance of having a reliable way to conduct the TQA process, it seems that there is limited literature that tackles the issue of reliability with a quantitative approach. In this work, we consider the TQA as a complex process from the point of view of physics of complex systems and approach the reliability issue from the Bayesian paradigm. Using a dataset of translation quality evaluations (in the form of error annotations), produced entirely by the Professional Translation Service Provider Translated SRL, we compare two Bayesian models that parameterise the following features involved in the TQA process: the translation difficulty, the characteristics of the translators involved in producing the translation, and of those assessing its quality - the reviewers. We validate the models in an unsupervised setting and show that it is possible to get meaningful insights into translators even with just one review per translation; subsequently, we extract information like translators' skills and reviewers' strictness, as well as their consistency in their respective roles. Using this, we show that the reliability of reviewers cannot be taken for granted even in the case of expert translators: a translator's expertise can induce a cognitive bias when reviewing a translation produced by another translator. The most expert translators, however, are characterised by the highest level of consistency, both in translating and in assessing the translation quality.

cs.CL

Product Progression: a machine learning approach to forecasting industrial upgrading

Economic complexity methods, and in particular relatedness measures, lack a systematic evaluation and comparison framework. We argue that out-of-sample forecast exercises should play this role, and we compare various machine learning models to set the prediction benchmark. We find that the key object to forecast is the activation of new products, and that tree-based algorithms clearly overperform both the quite strong auto-correlation benchmark and the other supervised algorithms. Interestingly, we find that the best results are obtained in a cross-validation setting, when data about the predicted country was excluded from the training set. Our approach has direct policy implications, providing a quantitative and scientifically tested measure of the feasibility of introducing a new product in a given country.

cs.LG

Relatedness in the Era of Machine Learning

Relatedness is a quantification of how much two human activities are similar in terms of the inputs and contexts needed for their development. Under the idea that it is easier to move between related activities than towards unrelated ones, empirical approaches to quantify relatedness are currently used as predictive tools to inform policies and development strategies in governments, international organizations, and firms. Here we focus on countries' industries and we show that the standard, widespread approach of estimating Relatedness through the co-location of activities (e.g. Product Space) generates a measure of relatedness that performs worse than trivial auto-correlation prediction strategies. We argue that this is a consequence of the poor signal-to-noise ratio present in international trade data. In this paper we show two main findings. First, we find that a shift from two-products correlations (network-density based) to many-products correlations (decision trees) can dramatically improve the quality of forecasts with a corresponding reduction of the risk of wrong policy choices. Then, we propose a new methodology to empirically estimate Relatedness that we call Continuous Projection Space (CPS). CPS, which can be seen as a general network embedding technique, vastly outperforms all the co-location, network-based approaches, while retaining a similar interpretability in terms of pairwise distances.

physics.soc-ph

A new and stable estimation method of country economic fitness and product complexity

We present a new metric estimating fitness of countries and complexity of products by exploiting a non-linear non-homogeneous map applied to the publicly available information on the goods exported by a country. The non homogeneous terms guarantee both convergence and stability. After a suitable rescaling of the relevant quantities, the non homogeneous terms are eventually set to zero so that this new metric is parameter free. This new map almost reproduces the results of the original homogeneous metrics already defined in literature and allows for an approximate analytic solution in case of actual binarized matrices based on the Revealed Comparative Advantage (RCA) indicator. This solution is connected with a new quantity describing the neighborhood of nodes in bipartite graphs, representing in this work the relations between countries and exported products. Moreover, we define the new indicator of country net-efficiency quantifying how a country efficiently invests in capabilities able to generate innovative complex high quality products. Eventually, we demonstrate analytically the local convergence of the algorithm involved.

econ.GN

Economic Complexity: "Buttarla in caciara" vs a constructive approach

This note is a contribution to the debate about the optimal algorithm for Economic Complexity that recently appeared on ArXiv [1, 2] . The authors of [2] eventually agree that the ECI+ algorithm [1] consists just in a renaming of the Fitness algorithm we introduced in 2012, as we explicitly showed in [3]. However, they omit any comment on the fact that their extensive numerical tests claimed to demonstrate that the same algorithm works well if they name it ECI+, but not if its name is Fitness. They should realize that this eliminates any credibility to their numerical methods and therefore also to their new analysis, in which they consider many algorithms [2]. Since by their own admission the best algorithm is the Fitness one, their new claim became that the search for the best algorithm is pointless and all algorithms are alike. This is exactly the opposite of what they claimed a few days ago and it does not deserve much comments. After these clarifications we also present a constructive analysis of the status of Economic Complexity, its algorithms, its successes and its perspectives. For us the discussion closes here, we will not reply to further comments.

econ.GN

Why we like the ECI+ algorithm

Recently a measure for Economic Complexity named ECI+ has been proposed by Albeaik et al. We like the ECI+ algorithm because it is mathematically identical to the Fitness algorithm, the measure for Economic Complexity we introduced in 2012. We demonstrate that the mathematical structure of ECI+ is strictly equivalent to that of Fitness (up to normalization and rescaling). We then show how the claims of Albeaik et al. about the ability of Fitness to describe the Economic Complexity of a country are incorrect. Finally, we hypothesize how the wrong results reported by these authors could have been obtained by not iterating the algorithm.

econ.GN

The Build-Up of Diversity in Complex Ecosystems

Diversity is a fundamental feature of ecosystems, even when the concept of ecosystem is extended to sociology or economics. Diversity can be intended as the count of different items, animals, or, more generally, interactions. There are two classes of stylized facts that emerge when diversity is taken into account. The first are Diversity explosions: evolutionary radiations in biology, or the process of escaping 'Poverty Traps' in economics are two well known examples. The second is nestedness: entities with a very diverse set of interactions are the only ones that interact with more specialized ones. In a single sentence: specialists interact with generalists. Nestedness is observed in a variety of bipartite networks of interactions: Biogeographic, macroeconomic and mutualistic to name a few. This indicates that entities diversify following a pattern. Since they appear in such very different systems, these two stylized facts point out that the build up of diversity is driven by a fundamental probabilistic mechanism, and here we sketch its minimal features. We show how the contraction of a random tripartite network, which is maximally entropic in all its degree distributions but one, can reproduce stylized facts of real data with great accuracy which is qualitatively lost when that degree distribution is changed. We base our reasoning on the combinatoric picture that the nodes on one layer of these bipartite networks can be described as combinations of a number of fundamental building blocks. The stylized facts of diversity that we observe in real systems can be explained with an extreme heterogeneity (a scale-free distribution) in the number of meaningful combinations in which each building block is involved. We show that if the usefulness of the building blocks has a scale-free distribution, then maximally entropic baskets of building blocks will give rise to very rich behaviors.

physics.soc-ph

How the Taxonomy of Products Drives the Economic Development of Countries

We introduce an algorithm able to reconstruct the relevant network structure on which the time evolution of country-product bipartite networks takes place. The significant links are obtained by selecting the largest values of the projected matrix. We first perform a number of tests of this filtering procedure on synthetic cases and a toy model. Then we analyze the bipartite network constituted by countries and exported products, using two databases for a total of almost 50 years. It is then possible to build a hierarchically directed network, in which the taxonomy of products emerges in a natural way. We study the influence of the structure of this taxonomy network on countries' development; in particular, guided by an example taken from the industrialization of South Korea, we link the structure of the taxonomy network to the empirical temporal connections between product activations, finding that the most relevant edges for countries' development are the ones suggested by our network. These results suggest paths in the product space which are easier to achieve, and so can drive countries' policies in the industrialization process.

econ.GN

A network analysis of countries' export flows: firm grounds for the building blocks of the economy

In this paper we analyze the bipartite network of countries and products from UN data on country production. We define the country-country and product-product projected networks and introduce a novel method of filtering information based on elements' similarity. As a result we find that country clustering reveals unexpected socio-geographic links among the most competing countries. On the same footings the products clustering can be efficiently used for a bottom-up classification of produced goods. Furthermore we mathematically reformulate the "reflections method" introduced by Hidalgo and Hausmann as a fixpoint problem; such formulation highlights some conceptual weaknesses of the approach. To overcome such an issue, we introduce an alternative methodology (based on biased Markov chains) that allows to rank countries in a conceptually consistent way. Our analysis uncovers a strong non-linear interaction between the diversification of a country and the ubiquity of its products, thus suggesting the possible need of moving towards more efficient and direct non-linear fixpoint algorithms to rank countries and products in the global market.

physics.soc-ph