SearcharxivSearch

arXiv subjects

Andrea Zaccaria

Publications and source records attributed to Andrea Zaccaria.

At least 19 recordsLinked to original sources

Economic complexity at subnational level: A consistency analysis

Several network-based measures have been proposed to assess the economic complexity of countries. These measures have provided important insights into national economic development, and they are now widely applied at the subnational level as well. Here, we show that such applications lead to inconsistent results, in the sense that the estimated complexity of the same product appears to depend on methodological details such as the geographical scale of analysis. Building on these findings, we propose a measure of territorial economic complexity based on an exogenous and extensive computation. We show that these methodological choices yield estimates that are more consistent and more strongly aligned with standard economic indicators, such as GDP per capita and employment.

econ.GN

From macro to micro: Economic complexity indicators for firm growth

A rich theoretical and empirical literature investigated the link between export diversification and firm performance. Prior theoretical works hinted at the key role of capability accumulation in shaping production activities and performance, without however producing product-level indicators able to forecast corporate growth. Building on economic complexity theory and the corporate growth literature, this paper examines which characteristics of a firm's export basket predict future performance. We analyze a unique longitudinal dataset that covers export and financial data for 12,852 Italian firms. We find that firms exporting products typically exported by wealthier countries -- a proxy for greater product sophistication and market value -- tend to experience higher growth and profit per employee. Moreover, we find that diversification outside of a firm's core production area is positively associated with future growth, whereas diversification within the core is negatively associated. This is revealed by introducing novel measures of in-block and out-of-block diversification, based on algorithmically-detected production blocks. Our findings suggest that growth is driven not just by how many products a firm exports, but also by where these products lie within the production ecosystem, at both local and global scales.

econ.GN

A job-based assessment of economic complexity: from hidden to revealed

Economic complexity measures aim to quantify the capability content or endowment of industries and territories; however, capabilities are not observable, and therefore cannot be directly used in the computations. We estimate such endowments by quantifying the quality and diversity of the skills in the occupations required in specific industries. We refer to this job-based assessment as the hidden complexity, in contrast with the usual revealed complexity, which is computed from economic outputs such as exports or production. We show that our job-based measure of complexity is positively associated to wage levels and labor productivity growth, whereas the classic revealed measure is not. Finally, we discuss the application of these methods at the territorial level, showing their connection with economic growth.

econ.GN

Product-level value chains from firm data: mapping trophic levels into economic growth

We reconstruct a product-level input-output network based on firm-level import-export data of Italian firms. We show that the network has a statistically significant, yet nuanced trophic structure, which is evident at the product level but is lost when the classification is coarse-grained. This detailed value chain allows us to characterize the trophic distance between inputs and outputs of single firms, and to derive a coherent picture at the sector level, finding that sectors such as weapons and vehicles are the ones with the largest increase in downstreamness between their inputs and their outputs. Our measure of downstreamness at the product level can be used to derive country-level indicators that characterize industrial strategies and capabilities and act as predictors of economic growth. With respect to the standard input/output analysis, we show that the fine-grained structure is qualitatively different from what can be observed using sector-level data. We finally prove that, even if we leverage exclusively data from Italian firms, the metrics that we derive are predictive at the country level and capture a significant description of the input-output relations of global value chains.

physics.soc-ph

Follow the money: a startup-based measure of AI exposure across occupations, industries and regions

The integration of artificial intelligence (AI) into the workplace is advancing rapidly, necessitating robust metrics to evaluate its tangible impact on the labour market. Existing measures of AI occupational exposure largely focus on AI's theoretical potential to substitute or complement human labour on the basis of technical feasibility, providing limited insight into actual adoption and offering inadequate guidance for policymakers. To address this gap, we introduce the AI Startup Exposure (AISE) index-a novel metric based on occupational descriptions from O*NET and AI applications developed by startups funded by the Y Combinator accelerator. Our findings indicate that while high-skilled professions are theoretically highly exposed according to conventional metrics, they are heterogeneously targeted by startups. Roles involving routine organizational tasks-such as data analysis and office management-display significant exposure, while occupations involving tasks that are less amenable to AI automation due to ethical or high-stakes, more than feasibility, considerations -- such as judges or surgeons -- present lower AISE scores. By focusing on venture-backed AI applications, our approach offers a nuanced perspective on how AI is reshaping the labour market. It challenges the conventional assumption that high-skilled jobs uniformly face high AI risks, highlighting instead the role of today's AI players' societal desirability-driven and market-oriented choices as critical determinants of AI exposure. Contrary to fears of widespread job displacement, our findings suggest that AI adoption will be gradual and shaped by social factors as much as by the technical feasibility of AI applications. This framework provides a dynamic, forward-looking tool for policymakers and stakeholders to monitor AI's evolving impact and navigate the changing labour landscape.

econ.GN

Uncovering key predictors of high-growth firms via explainable machine learning

Predicting high-growth firms has attracted increasing interest from the technological forecasting and machine learning communities. Most existing studies primarily utilize financial data for these predictions. However, research suggests that a firm's research and development activities and its network position within technological ecosystems may also serve as valuable predictors. To unpack the relative importance of diverse features, this paper analyzes financial and patent data from 5,071 firms, extracting three categories of features: financial features, technological features of granted patents, and network-based features derived from firms' connections to their primary technologies. By utilizing ensemble learning algorithms, we demonstrate that incorporating financial features with either technological, network-based features, or both, leads to more accurate high-growth firm predictions compared to using financial features alone. To delve deeper into the matter, we evaluate the predictive power of each individual feature within their respective categories using explainable artificial intelligence methods. Among non-financial features, the maximum economic value of a firm's granted patents and the number of patents related to a firms' primary technologies stand out for their importance. Furthermore, firm size is positively associated with high-growth probability up to a certain threshold size, after which the association plateaus. Conversely, the maximum economic value of a firm's granted patents is positively linked to high-growth probability only after a threshold value is exceeded. These findings elucidate the complex predictive role of various features in forecasting high-growth firms and could inform technological resource allocation as well as investment decisions.

physics.soc-ph

Machine learning-based similarity measure to forecast M&A from patent data

Defining and finalizing Mergers and Acquisitions (M&A) requires complex human skills, which makes it very hard to automatically find the best partner or predict which firms will make a deal. In this work, we propose the MASS algorithm, a specifically designed measure of similarity between companies and we apply it to patenting activity data to forecast M&A deals. MASS is based on an extreme simplification of tree-based machine learning algorithms and naturally incorporates intuitive criteria for deals; as such, it is fully interpretable and explainable. By applying MASS to the Zephyr and Crunchbase datasets, we show that it outperforms LightGCN, a "black box" graph convolutional network algorithm. When similar companies have disjoint patenting activities, on the contrary, LightGCN turns out to be the most effective algorithm. This study provides a simple and powerful tool to model and predict M&A deals, offering valuable insights to managers and practitioners for informed decision-making.

physics.soc-ph

Pattern-detection in the global automotive industry: a manufacturer-supplier-product network analysis

Production networks arise from supply and customer relations among firms. These systems are gaining growing attention as a consequence of disruptions due to natural or man-made disasters that happened in the last years, such as the Covid-19 pandemic or the Russia-Ukraine war. However, data constraints force the few, available studies to consider only country-specific production networks. In order to fully capture the cross-country structure of modern supply chains, here we focus on the global automotive industry as represented by the MarkLines Automotive dataset. After representing this data as a network of manufacturers, suppliers, and products, we perform a pattern-detection exercise using a statistically grounded validation technique based on the maximum entropy principle. We reveal the presence of a significantly large number of V-shaped and square-shaped motifs, indicating that manufacturing firms compete and are seldom engaged in a buyer-supplier relationship, while they typically have many suppliers in common. Interestingly, generalist and specialist suppliers coexist in the network. Additionally, we unveil the presence of geographical patterns, with manufacturers clustering around groups of suppliers; for instance, Chinese firms constitute a disconnected community, likely an effect of the protectionist policies promoted by the Chinese government. We also show the tendency of suppliers to organize their production by targeting specific functional modules of a vehicle. Besides shedding light on the self-organising principles shaping production networks, our findings open up the possibility of designing realistic generative models of supply chains, to be used for testing the resilience of the interconnected global economy.

physics.soc-ph

Sapling Similarity: a performing and interpretable memory-based tool for recommendation

Many bipartite networks describe systems where an edge represents a relation between a user and an item. Measuring the similarity between either users or items is the basis of memory-based collaborative filtering, a widely used method to build a recommender system with the purpose of proposing items to users. When the edges of the network are unweighted, the popular common neighbors-based approaches, allowing only positive similarity values, neglect the possibility and the effect of two users (or two items) being very dissimilar. Moreover, they underperform with respect to model-based (machine learning) approaches, although providing higher interpretability. Inspired by the functioning of Decision Trees, we propose a method to compute similarity that allows also negative values, the Sapling Similarity. The key idea is to look at how the information that a user is connected to an item influences our prior estimation of the probability that another user is connected to the same item: if it is reduced, then the similarity between the two users will be negative, otherwise, it will be positive. We show that, when used to build memory-based collaborative filtering, Sapling Similarity provides better recommendations than existing similarity metrics. Then we compare the Sapling Similarity Collaborative Filtering (SSCF, a hybrid of the item-based and the user-based) with state-of-the-art models using standard datasets. Even if SSCF depends on only one straightforward hyperparameter, it has comparable or higher recommending accuracy, and outperforms all other models on the Amazon-Book dataset, while retaining the high explainability of memory-based approaches.

cs.IR

Mapping job complexity and skills into wages

We use algorithmic and network-based tools to build and analyze the bipartite network connecting jobs with the skills they require. We quantify and represent the relatedness between jobs and skills by using statistically validated networks. Using the fitness and complexity algorithm, we compute a skill-based complexity of jobs. This quantity is positively correlated with the average salary, abstraction, and non-routinarity level of jobs. Furthermore, coherent jobs - defined as the ones requiring closely related skills - have, on average, lower wages. We find that salaries may not always reflect the intrinsic value of a job, but rather other wage-setting dynamics that may not be directly related to its skill composition. Our results provide valuable information for policymakers, employers, and individuals to better understand the dynamics of the labor market and make informed decisions about their careers.

econ.GN

Which products activate a product? An explainable machine learning approach

Tree-based machine learning algorithms provide the most precise assessment of the feasibility for a country to export a target product given its export basket. However, the high number of parameters involved prevents a straightforward interpretation of the results and, in turn, the explainability of policy indications. In this paper, we propose a procedure to statistically validate the importance of the products used in the feasibility assessment. In this way, we are able to identify which products, called explainers, significantly increase the probability to export a target product in the near future. The explainers naturally identify a low dimensional representation, the Feature Importance Product Space, that enhances the interpretability of the recommendations and provides out-of-sample forecasts of the export baskets of countries. Interestingly, we detect a positive correlation between the complexity of a product and the complexity of its explainers.

econ.GN

Prediction and visualization of Mergers and Acquisitions using Economic Complexity

Mergers and Acquisitions represent important forms of business deals, both because of the volumes involved in the transactions and because of the role of the innovation activity of companies. Nevertheless, Economic Complexity methods have not been applied to the study of this field. By considering the patent activity of about one thousand companies, we develop a method to predict future acquisitions by assuming that companies deal more frequently with technologically related ones. We address both the problem of predicting a pair of companies for a future deal and that of finding a target company given an acquirer. We compare different forecasting methodologies, including machine learning and network-based algorithms, showing that a simple angular distance with the addition of the industry sector information outperforms the other approaches. Finally, we present the Continuous Company Space, a two-dimensional representation of firms to visualize their technological proximity and possible deals. Companies and policymakers can use this approach to identify companies most likely to pursue deals or to explore possible innovation strategies.

physics.soc-ph

Machine learning to assess relatedness: the advantage of using firm-level data

The relatedness between a country or a firm and a product is a measure of the feasibility of that economic activity. As such, it is a driver for investments at a private and institutional level. Traditionally, relatedness is measured using networks derived by country-level co-occurrences of product pairs, that is counting how many countries export both. In this work, we compare networks and machine learning algorithms trained not only on country-level data, but also on firms, that is something not much studied due to the low availability of firm-level data. We quantitatively compare the different measures of relatedness, by using them to forecast the exports at the country and firm-level, assuming that more related products have a higher likelihood to be exported in the future. Our results show that relatedness is scale-dependent: the best assessments are obtained by using machine learning on the same typology of data one wants to predict. Moreover, we found that while relatedness measures based on country data are not suitable for firms, firm-level data are very informative also for the development of countries. In this sense, models built on firm data provide a better assessment of relatedness. We also discuss the effect of using parameter optimization and community detection algorithms to identify clusters of related companies and products, finding that a partition into a higher number of blocks decreases the computational time while maintaining a prediction performance well above the network-based benchmarks.

cs.LG

The trickle down from environmental innovation to productive complexity

We study the empirical relationship between green technologies and industrial production at very fine-grained levels by employing Economic Complexity techniques. Firstly, we use patent data on green technology domains as a proxy for competitive green innovation and data on exported products as a proxy for competitive industrial production. Secondly, with the aim of observing how green technological development trickles down into industrial production, we build a bipartite directed network linking single green technologies at time $t_1$ to single products at time $t_2 \ge t_1$ on the basis of their time-lagged co-occurrences in the technological and industrial specialization profiles of countries. Thirdly we filter the links in the network by employing a maximum entropy null-model. In particular, we find that the industrial sectors most connected to green technologies are related to the processing of raw materials, which we know to be crucial for the development of clean energy innovations. Furthermore, by looking at the evolution of the network over time, we observe that more complex green technological know-how requires more time to be transmitted to industrial production, and is also linked to more complex products.

econ.GN

A Bayesian approach to translators' reliability assessment

Translation Quality Assessment (TQA) is a process conducted by human translators and is widely used, both for estimating the performance of (increasingly used) Machine Translation, and for finding an agreement between translation providers and their customers. While translation scholars are aware of the importance of having a reliable way to conduct the TQA process, it seems that there is limited literature that tackles the issue of reliability with a quantitative approach. In this work, we consider the TQA as a complex process from the point of view of physics of complex systems and approach the reliability issue from the Bayesian paradigm. Using a dataset of translation quality evaluations (in the form of error annotations), produced entirely by the Professional Translation Service Provider Translated SRL, we compare two Bayesian models that parameterise the following features involved in the TQA process: the translation difficulty, the characteristics of the translators involved in producing the translation, and of those assessing its quality - the reviewers. We validate the models in an unsupervised setting and show that it is possible to get meaningful insights into translators even with just one review per translation; subsequently, we extract information like translators' skills and reviewers' strictness, as well as their consistency in their respective roles. Using this, we show that the reliability of reviewers cannot be taken for granted even in the case of expert translators: a translator's expertise can induce a cognitive bias when reviewing a translation produced by another translator. The most expert translators, however, are characterised by the highest level of consistency, both in translating and in assessing the translation quality.

cs.CL

Meta-validation of bipartite network projections

Monopartite projections of bipartite networks are useful tools for modeling indirect interactions in complex systems. The standard approach to identify significant links is statistical validation using a suitable null network model, such as the popular configuration model (CM) that constrains node degrees and randomizes everything else. However different CM formulations exist, depending on how the constraints are imposed and for which sets of nodes. Here we systematically investigate the application of these formulations in validating the same network, showing that they lead to different results even when the same significance threshold is used. Instead a much better agreement is obtained for the same density of validated links. We thus propose a meta-validation approach that allows to identify model-specific significance thresholds for which the signal is strongest, and at the same time to obtain results independent of the way in which the null hypothesis is formulated. We illustrate this procedure using data on scientific production of world countries.

physics.soc-ph

The different structure of economic ecosystems at the scales of companies and countries

A key element to understand complex systems is the relationship between the spatial scale of investigation and the structure of the interrelation among its elements. When it comes to economic systems, it is now well-known that the country-product bipartite network exhibits a nested structure, which is the foundation of different algorithms that have been used to scientifically investigate countries' development and forecast national economic growth. Changing the subject from countries to companies, a significantly different scenario emerges. Through the analysis of a unique dataset of Italian firms' exports and a worldwide dataset comprising countries' exports, here we find that, while a globally nested structure is observed at the country level, a local, in-block nested structure emerges at the level of firms. Remarkably, this in-block nestedness is statistically significant with respect to suitable null models and the algorithmic partitions of products into blocks have a high correspondence with exogenous product classifications. These findings lay a solid foundation for developing a scientific approach based on the physics of complex systems to the analysis of companies, which has been lacking until now.

physics.soc-ph

Quantifying the Unexpected: a scientific approach to Black Swans

Many natural and socio-economic systems are characterized by power-law distributions that make the occurrence of extreme events not negligible. Such events are sometimes referred to as Black Swans, but a quantitative definition of a Black Swan is still lacking. Here, by leveraging on the properties of Zipf-Mandelbrot law, we investigate the relations between such extreme events and the dynamics of the upper cutoff of the inherent distribution. This approach permits a quantification of extreme events and allows to classify them as White, Grey, or Black Swans. Our criterion is in accordance with some previous findings, but also allows us to spot new examples of Black Swans, such as Lionel Messi and the Turkish Airline Flight 981 disaster. The systematic and quantitative methodology we developed allows a scientific and immediate categorization of rare events, providing also new insight into the generative mechanism behind Black Swans.

physics.soc-ph