SearcharxivSearch

arXiv subjects

Andrea Mina

Publications and source records attributed to Andrea Mina.

6 recordsLinked to original sources

Large Language Models, Encoder Architectures and Hybrid Approaches for Patent Classification

Automated patent classification is essential for organizing technological knowledge and constructing indicators of technological change, specialization, and leadership. To identify the relative strengths and weaknesses of popular state- of-the-art approaches to this problem, we perform a controlled comparison of patent-specific encoders and open-weight local LLMs for hierarchical multi-label Cooperative Patent Classification (CPC). We find that the best-performing encoder (task-adapted BERT-for-Patents) outperforms the best-performing LLM (fine-tuned Qwen3.5-9B), while requiring one to two orders of magnitude less energy for inference. The two model families show complementary capabilities, which we leverage through a hybrid pipeline that routes patents with the highest encoder uncertainty to the LLM, yielding significant gains on this subset. We also find that across models, classification errors are especially pronounced in CPC categories that are cross-cutting or semantically broad - such as Section Y - and that they have substantial consequences for technology mapping and country and assignee rankings. The analysis covers predictive and hierarchical performance, computational cost and energy consumption, external validation on EPO patents, and the propagation of classification errors into downstream technological indicators, with implications for automated patent classification procedures used by patent offices, technology analysts, and scientometric researchers.

cs.CE

The hidden structure of innovation networks

Innovation emerges from complex collaboration patterns - among inventors, firms, or institutions. However, not much is known about the overall mesoscopic structure around which inventive activity self-organizes. Here, we tackle this problem by employing patent data to analyze both individual (\textit{co-inventorship}) and organization (\textit{co-ownership}) networks in three strategic domains (\textit{artificial intelligence}, \textit{biotechnology} and \textit{semiconductors}). We characterize the mesoscale structure (in terms of clusters) of each domain by comparing two alternative methods: a standard baseline - modularity maximization - and one based on the minimization of the Bayesian Information Criterion, within the Stochastic Block Model and its degree-corrected variant. We find that, across sectors, inventor networks are denser and more clustered than organization ones - consistently with the presence of small recurrent teams embedded into broader institutional hierarchies - whereas organization networks have a neater role-based structures, with few bridging firms coordinating the most peripheral ones; still, both are characterized by the presence of local core-periphery structures. We also find that the discovered meso-structures are connected to innovation output. In particular, Lorenz curves of forward citations show a pervasive inequality in technological influence: across sectors and methods, both inventor (especially) and organization networks consistently show high levels of concentration of citations in a few of the discovered clusters. Our results demonstrate that the baseline modularity-based method may not be capable of fully capturing the way collaborations drive the spreading of inventive impact across technological domains. This is due to the presence of local hierarchies that call for the more refined tools of Bayesian inference.

econ.GN

The anatomy of Green AI technologies: structure, evolution, and impact

Artificial intelligence (AI) is a key enabler of innovation against climate change. In this study, we investigate the intersection of AI and climate adaptation and mitigation technologies through patent analyses of a novel dataset of approximately 63 000 Green AI patents. We analyze patenting trends, corporate ownership of the technology, the geographical distributions of patents, their impact on follow-on inventions and their market value. We use topic modeling (BERTopic) to identify 16 major technological domains, track their evolution over time, and identify their relative impact. We uncover a clear shift from legacy domains such as combustion engines technology to emerging areas like data processing, microgrids, and agricultural water management. We find evidence of growing concentration in corporate patenting against a rapidly increasing number of patenting firms. Looking at the technological and economic impact of patents, while some Green AI domains combine technological impact and market value, others reflect weaker private incentives for innovation, despite their relevance for climate adaptation and mitigation strategies. This is where policy intervention might be required to foster the generation and use of new Green AI applications.

econ.GN

There are different shades of green: heterogeneous environmental innovations and their effects on firm performance

Using a firm-level dataset from the Spanish Technological Innovation Panel (2003-2016), this study explores the characteristics of environmentally innovative firms and quantifies the effects of pursuing different types of environmental innovation strategies (resource-saving, pollution-reducing, and regulation-driven innovations) on sales, employment, and productivity dynamics.

econ.GN

Venture Capital investments through the lens of Network and Functional Data Analysis

In this paper we characterize the performance of venture capital-backed firms based on their ability to attract investment. The aim of the study is to identify relevant predictors of success built from the network structure of firms' and investors' relations. Focusing on deal-level data for the health sector, we first create a bipartite network among firms and investors, and then apply functional data analysis (FDA) to derive progressively more refined indicators of success captured by a binary, a scalar and a functional outcome. More specifically, we use different network centrality measures to capture the role of early investments for the success of the firm. Our results, which are robust to different specifications, suggest that success has a strong positive association with centrality measures of the firm and of its large investors, and a weaker but still detectable association with centrality measures of small investors and features describing firms as knowledge bridges. Finally, based on our analyses, success is not associated with firms' and investors' spreading power (harmonic centrality), nor with the tightness of investors' community (clustering coefficient) and spreading ability (VoteRank).

stat.AP

Can you always reap what you sow? Network and functional data analysis of VC investments in health-tech companies

"Success" of firms in venture capital markets is hard to define, and its determinants are still poorly understood. We build a bipartite network of investors and firms in the healthcare sector, describing its structure and its communities. Then, we characterize "success" introducing progressively more refined definitions, and we find a positive association between such definitions and the centrality of a company. In particular, we are able to cluster funding trajectories of firms into two groups capturing different "success" regimes and to link the probability of belonging to one or the other to their network features (in particular their centrality and the one of their investors). We further investigate this positive association by introducing scalar as well as functional "success" outcomes, confirming our findings and their robustness.

cs.SI