SearcharxivSearch

arXiv subjects

Massimo Riccaboni

Publications and source records attributed to Massimo Riccaboni.

At least 19 recordsLinked to original sources

Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI

This paper examines the tension between the benefits of generative artificial intelligence (AI) for scientific research and the unresolved governance questions that accompany its rapid adoption. Drawing on an academic roundtable held at the AI for Science and Innovation Workshop (Scuola IMT Alti Studi Lucca, April 2026) and on a fast-expanding empirical literature, it maps the disagreement within the research community across four stages of the research process: funding, research tasks, publication and peer review, and use and uptake. The empirical case for AI's productivity, augmentation, and democratization effects has strengthened. The picture changes once productivity is disaggregated: AI-assisted work shows measurable gains in publication volume and citation share, while the evidence on novelty, disruption, and breakthrough output remains ambiguous or negative. We argue that the divergence between private and social returns arises through three analytically distinct mechanisms, namely information asymmetry, negative externalities on a shared knowledge base, and depletion of research capacity, and that each calls for a different governance instrument. We propose Responsible Research with AI (RRAI), an extension of the Responsible Research and Innovation tradition organized around four principles that operate at different levels of the research system: disclosure, differentiation, narrative, and proportionality. RRAI builds on existing institutional scaffolding, including the EU AI Act, UNESCO, and the OECD, and aims to preserve AI's productivity gains while addressing systemic risks that individual researchers can neither observe nor manage on their own.

econ.GN

Licensing and Innovation Regimes in Pharmaceutical R&D

We study how licensing affects the allocation of innovation in pharmaceutical R&D. We develop a model in which projects differ in both quality and innovation regime, distinguishing between incremental and novel innovations. Information precision is higher for incremental projects and lower for novel ones, generating different equilibrium dynamics in the market for technology. The model predicts that licensing sustains positive selection and competitive return equalization for incremental innovation, while novel projects may exhibit weaker screening consistent with lemons-type frictions. Using product-level data and Double Machine Learning methods, we test these predictions across success probabilities and monetary returns. We find that licensing increases success probability overall, but return equalization holds primarily for incremental projects. For novel innovation, licensing does not exhibit the same equilibrium adjustment, suggesting residual market imperfections. Instrumenting for licensing using exogenous pipeline shocks confirms this pattern causally: the competitive risk-return trade-off is preserved for incremental 'rushed' licenses, but it breaks down for novel ones. Our results reconcile evidence on both competitive efficiency and information frictions in markets for technologies, showing that market performance depends systematically on the type of innovation being transacted.

econ.GN

A large-scale dataset of Android applications and their SDK dependencies

Mobile applications (apps) increasingly rely on third-party Software Development Kits (SDKs) to provide services such as advertising, analytics, authentication, crash reporting, and location services. These components form an important, but often hidden, layer of the mobile ecosystem. Here, we present a large-scale dataset linking Android apps to the third-party SDKs they integrate. The dataset was constructed by combining app package files (APKs) and metadata from AndroZoo with SDK detection rules provided by Exodus Privacy. We implemented a reproducible, static-analysis pipeline that downloads APKs, inspects their compiled code, and detects SDKs through code-level signature matching: the released dataset contains 334,719 unique app-version observations (associated with 99,722 Android apps) and 246 third-party SDKs - including app-level metadata such as categories, comments, file size, download counts, ratings and SDK-level metadata such as categories, code roots, code versions, names. The dataset supports the construction of a large-scale app-SDK dependency network and its projection onto both layers; in addition, SDKs are mapped to their corresponding operating companies, enabling provider-level analysis of technological concentration and upstream control in the mobile ecosystem. To support reproducible research, the pipeline and dataset are publicly released on GitHub and Zenodo; more broadly, the project provides a reusable research infrastructure for studying third-party technological dependencies, data-collection capabilities, and privacy infrastructures across Android apps.

cs.SE

Bridging Distant Ideas: the Impact of AI on R&D and Recombinant Innovation

We study how artificial intelligence (AI) affects firms' incentives to pursue incremental versus radical knowledge recombinations. We develop a model of recombinant innovation embedded in a Schumpeterian quality-ladder framework, in which innovation arises from recombining ideas across varying distances in a knowledge space. R&D consists of multiple tasks, a fraction of which can be performed by AI. AI facilitates access to distant knowledge domains, but at the same time it also increases the aggregate rate of creative destruction, shortening the monopoly duration that rewards radical innovations. Moreover, excessive reliance on AI may reduce the originality of research and lead to duplication of research efforts. We obtain three main results. First, higher AI productivity encourages more distant recombinations, if the direct facilitation effect is stronger than the indirect effect due to intensified competition from rivals. Second, the effect of increasing the share of AI-automated R&D tasks is non-monotonic: firms initially target more radical innovations, but beyond a threshold of human-AI complementarity, they shift the focus toward incremental innovations. Third, in the limiting case of full automation, the model predicts that optimal recombination distance collapses to zero, suggesting that fully AI-driven research would undermine the very knowledge creation that it seeks to accelerate.

econ.TH

Most Western African migrants remain local and travel short distances

Migration patterns are complex and context-dependent, with the distances migrants travel varying greatly depending on socio-economic and demographic factors. While global migration studies often focus on Western countries, there is a crucial gap in our understanding of migration dynamics within the African continent, particularly in West Africa. Using data from over 60,000 individuals from eight West African countries, this study examines the determinants of migration distance in the region. Our analysis reveals a bimodal distribution of migration distances: while most migrants travel locally within a hundred km, a smaller yet significant portion undertakes long-distance journeys, often exceeding 3,000 km. Socio-economic factors such as employment status, marital status and level of education play a decisive role in determining migration distances. Unemployed migrants, for instance, travel substantially farther (1,467 km on average) than their employed counterparts (295 km). Furthermore, we find that conflict-induced migration is particularly variable, with migrants fleeing violence often undertaking longer and riskier journeys. Our findings highlight the importance of considering both local and long-distance migration in policy decisions and support systems, as well as the need for a comprehensive understanding of migration in non-Western contexts. This study contributes to the broader discourse on human mobility by providing new insights into migration patterns in Western Africa, which in turn has implications for global migration research and policy development.

physics.soc-ph

Leveraging Knowledge Networks: Rethinking Technological Value Distribution in mRNA Vaccine Innovations

This study examines the roles of public and private sector actors in the development of mRNA vaccines, a breakthrough innovation in modern medicine. Using a dataset of 151 core patent families and 2,416 antecedent (cited) patents, we analyze the structure and dynamics of the mRNA vaccine knowledge network through network theory. Our findings highlight the central role of biotechnology firms, such as Moderna and BioNTech, alongside the crucial contributions of universities and public research organizations (PROs) in providing foundational knowledge.We develop a novel credit allocation framework, showing that universities, PROs, government and research centers account for at least 27% of the external technological knowledge base behind mRNA vaccine breakthroughs - representing a minimum threshold of their overall contribution. Our study offers new insights into pharmaceutical and biotechnology innovation dynamics, emphasizing how Moderna and BioNTech's mRNA technologies have benefited from academic institutions, with notable differences in their institutional knowledge sources.

physics.soc-ph

Compositional Growth Models

We review models of compositional growth, which were introduced to explain the growth statistics of various quantities ranging from firm sizes to GDP. In these models, entities are decomposed into units that grow independently. Thus, the growth rate of the entity is the addition of the growth rates of the composing units, with possibly heterogeneous weights. We review such models and show that they can be understood through a unifying theoretical framework, explaining the resulting growth rate distributions using mixtures of Gaussians.

econ.GN

Breaking New Ground, Reinforcing Old Gaps: Gender Disparities in Access to Emerging Research Frontiers

This study exploits COVID-19 as an exogenous shock in biomedical research to show how the emergence of an unexpected new research topic exacerbates gender bias in key authorship positions of scientific publications relevant to new research topics (e.g. Vaccines, Epidemiology). We determine author's gender based on the names listed on their scientific publications and analyze the changes in the composition of the scientific teams after the COVID-19 outbreak. Using a Difference-in-Differences approach, we find that although the share of female authorship has increased overall, women are less likely to be first or last authors (the most prestigious positions) on COVID-19-related research papers and more likely to be found in middle author positions. Stay-at-home mandates, the journal importance and funding opportunities do not fully account for the decline of women in key author positions. The main difference in first authorship is due to the composition of the team and the experience of the lead authors in COVID-19 related research. First authorship by women declined after teams of novices emerged, where lead authors have no prior experience in COVID-related research. Discretionality in first-author appointments for newcomers, combined with high pressure to publish quickly, may have led to discriminatory biases. Conversely, there may also be differences in risk-taking attitudes in doing research in unfamiliar domains. Monitoring gender inequality in scientific production is crucial for reducing gender inequalities and for implementing timely policies that ensure equal access to emerging research topics.

econ.GN

Machine Learning for Zombie Hunting: Predicting Distress from Firms' Accounts and Missing Values

In this contribution, we propose machine learning techniques to predict zombie firms. First, we derive the risk of failure by training and testing our algorithms on disclosed financial information and non-random missing values of 304,906 firms active in Italy from 2008 to 2017. Then, we spot the highest financial distress conditional on predictions that lies above a threshold for which a combination of false positive rate (false prediction of firm failure) and false negative rate (false prediction of active firms) is minimized. Therefore, we identify zombies as firms that persist in a state of financial distress, i.e., their forecasts fall into the risk category above the threshold for at least three consecutive years. For our purpose, we implement a gradient boosting algorithm (XGBoost) that exploits information about missing values. The inclusion of missing values in our predictive model is crucial because patterns of undisclosed accounts are correlated with firm failure. Finally, we show that our preferred machine learning algorithm outperforms (i) proxy models such as Z-scores and the Distance-to-Default, (ii) traditional econometric methods, and (iii) other widely used machine learning techniques. We provide evidence that zombies are on average less productive and smaller, and that they tend to increase in times of crisis. Finally, we argue that our application can help financial institutions and public authorities design evidence-based policies-e.g., optimal bankruptcy laws and information disclosure policies.

econ.EM

The echo chamber effect resounds on financial markets: a social media alert system for meme stocks

The short squeeze of Gamestop (GME) has revealed to the world how retail investors pooling through social media can severely impact financial markets. In this paper, we devise an early warning signal to detect suspicious users' social network activity, which might affect the financial market stability. We apply our approach to the subreddit r/WallStreetBets, selecting two meme stocks (GME and AMC) and two non-meme stocks (AAPL and MSFT) as case studies. The alert system is structured in two stpng; the first one is based on extraordinary activity on the social network, while the second aims at identifying whether the movement seeks to coordinate the users to a bulk action. We run an event study analysis to see the reaction of the financial markets when the alert system catches social network turmoil. A regression analysis witnesses the discrepancy between the meme and non-meme stocks in how the social networks might affect the trend on the financial market.

q-fin.TR

The Impact of Acquisitions in the Biotechnology Sector on R&D Productivity

This study examines the effects of acquisitions on the retention and R&D productivity of inventors in the biotech sector, using data from 15,318 inventors involved in 1,375 acquisitions between 1990 and 2010. We employ a staggered difference-in-differences approach and find that acquisitions lead to a 13.5% decrease in inventor retention and a 35% drop in citation-weighted patent productivity post-acquisition. The productivity decline is more severe for inventors who remain with the acquiring firm, particularly for those whose expertise is closely tied to the target company. However, older inventors and those whose expertise aligns with the acquiring company's existing R&D portfolio tend to retain higher productivity levels after the acquisition.

econ.GN

Hierarchical Clustering and Matrix Completion for the Reconstruction of World Input-Output Tables

World Input-Output (I/O) matrices provide the networks of within- and cross-country economic relations. In the context of I/O analysis, the methodology adopted by national statistical offices in data collection raises the issue of obtaining reliable data in a timely fashion and it makes the reconstruction of (part of) the I/O matrices of particular interest. In this work, we propose a method combining hierarchical clustering and Matrix Completion (MC) with a LASSO-like nuclear norm penalty, to impute missing entries of a partially unknown I/O matrix. Through simulations based on synthetic matrices we study the effectiveness of the proposed method to predict missing values from both previous years data and current data related to countries similar to the one for which current data are obscured. To show the usefulness of our method, an application based on World Input-Output Database (WIOD) tables - which are an example of industry-by-industry I/O tables - is provided. Strong similarities in structure between WIOD and other I/O tables are also found, which make the proposed approach easily generalizable to them.

stat.ML

Product recalls, market size and innovation in the pharmaceutical industry

The idea that research investments respond to market rewards is well established in the literature on markets for innovation (Schmookler, 1966; Acemoglu and Linn, 2004; Bryan and Williams, 2021). Empirical evidence tells us that a change in market size, such as the one measured by demographical shifts, is associated with an increase in the number of new drugs available (Acemoglu and Linn, 2004; Dubois et al., 2015). However, the debate about potential reverse causality is still open (Cerda et al., 2007). In this paper we analyze market size's effect on innovation as measured by active clinical trials. The idea is to exploit product recalls an innovative instrument tested to be sharp, strong, and unexpected. The work analyses the relationship between US market size and innovation at ATC-3 level through an original dataset and the two-step IV methodology proposed by Wooldridge et al. (2019). The results reveal a robust and significantly positive response of number of active trials to market size.

econ.GN

Assessing the Heterogeneous Impact of Economy-Wide Shocks: A Machine Learning Approach Applied to Colombian Firms

Our paper presents a methodology to study the heterogeneous effects of economy-wide shocks and applies it to the case of the impact of the COVID-19 crisis on exports. This methodology is applicable in scenarios where the pervasive nature of the shock hinders the identification of a control group unaffected by the shock, as well as the ex-ante definition of the intensity of the shock's exposure of each unit. In particular, our study investigates the effectiveness of various Machine Learning (ML) techniques in predicting firms' trade and, by building on recent developments in causal ML, uses these predictions to reconstruct the counterfactual distribution of firms' trade under different COVID-19 scenarios and to study treatment effect heterogeneity. Specifically, we focus on the probability of Colombian firms surviving in the export market under two different scenarios: a COVID-19 setting and a non-COVID-19 counterfactual situation. On average, we find that the COVID-19 shock decreased a firm's probability of surviving in the export market by about 20 percentage points in April 2020. We study the treatment effect heterogeneity by employing a classification analysis that compares the characteristics of the firms on the tails of the estimated distribution of the individual treatment effects.

econ.GN

COVID-19 and Unemployment Risk: Lessons for the Vaccination Campaign

Assessing the economic impact of COVID-19 pandemic and public health policies is essential for a rapid recovery. In this paper, we analyze the impact of mobility contraction on furloughed workers and excess deaths in Italy. We provide a link between the reduction of mobility and excess deaths, confirming that the first countrywide lockdown has been effective in curtailing the COVID-19 epidemics. Our analysis points out that a mobility contraction of 10% leads to a mortality reduction of 5% whereas it leads to an increase of 50% in full time equivalent furloughed workers. Based on our results, we propose a prioritizing policy for the most advanced stage of the COVID-19 vaccination campaign, considering the unemployment risk of the healthy active population. Keywords: COVID-19 mortality; Furlough schemes; Economic impact of lockdowns; Vaccination rollout: Unemployment risk

physics.soc-ph

The Impact of the COVID-19 Pandemic on Scientific Research in the Life Sciences

The COVID-19 outbreak has posed an unprecedented challenge to humanity and science. On the one side, public and private incentives have been put in place to promptly allocate resources toward research areas strictly related to the COVID-19 emergency. But on the flip side, research in many fields not directly related to the pandemic has lagged behind. In this paper, we assess the impact of COVID-19 on world scientific production in the life sciences. We investigate how the usage of medical subject headings (MeSH) has changed following the outbreak. We estimate through a difference-in-differences approach the impact of COVID-19 on scientific production through PubMed. We find that COVID-related research topics have risen to prominence, displaced clinical publications, diverted funds away from research areas not directly related to COVID-19 and that the number of publications on clinical trials in unrelated fields has contracted. Our results call for urgent targeted policy interventions to reactivate biomedical research in areas that have been neglected by the COVID-19 emergency.

econ.GN

Early warnings of COVID-19 outbreaks across Europe from social media?

We analyze data from Twitter to uncover early-warning signals of COVID-19 outbreaks in Europe in the winter season 2019-2020, before the first public announcements of local sources of infection were made. We show evidence that unexpected levels of concerns about cases of pneumonia were raised across a number of European countries. Whistleblowing came primarily from the geographical regions that eventually turned out to be the key breeding grounds for infections. These findings point to the urgency of setting up an integrated digital surveillance system in which social media can help geo-localize chains of contagion that would otherwise proliferate almost completely undetected.

econ.GN

A network approach to expertise retrieval based on path similarity and credit allocation

With the increasing availability of online scholarly databases, publication records can be easily extracted and analysed. Researchers can promptly keep abreast of others' scientific production and, in principle, can select new collaborators and build new research teams. A critical factor one should consider when contemplating new potential collaborations is the possibility of unambiguously defining the expertise of other researchers. While some organisations have established database systems to enable their members to manually produce a profile, maintaining such systems is time-consuming and costly. Therefore, there has been a growing interest in retrieving expertise through automated approaches. Indeed, the identification of researchers' expertise is of great value in many applications, such as identifying qualified experts to supervise new researchers, assigning manuscripts to reviewers, and forming a qualified team. Here, we propose a network-based approach to the construction of authors' expertise profiles. Using the MEDLINE corpus as an example, we show that our method can be applied to a number of widely used data sets and outperforms other methods traditionally used for expertise identification.

cs.SI