SearcharxivSearch

arXiv subjects

Lorenzo Emer

Publications and source records attributed to Lorenzo Emer.

3 recordsLinked to original sources

Large Language Models, Encoder Architectures and Hybrid Approaches for Patent Classification

Automated patent classification is essential for organizing technological knowledge and constructing indicators of technological change, specialization, and leadership. To identify the relative strengths and weaknesses of popular state- of-the-art approaches to this problem, we perform a controlled comparison of patent-specific encoders and open-weight local LLMs for hierarchical multi-label Cooperative Patent Classification (CPC). We find that the best-performing encoder (task-adapted BERT-for-Patents) outperforms the best-performing LLM (fine-tuned Qwen3.5-9B), while requiring one to two orders of magnitude less energy for inference. The two model families show complementary capabilities, which we leverage through a hybrid pipeline that routes patents with the highest encoder uncertainty to the LLM, yielding significant gains on this subset. We also find that across models, classification errors are especially pronounced in CPC categories that are cross-cutting or semantically broad - such as Section Y - and that they have substantial consequences for technology mapping and country and assignee rankings. The analysis covers predictive and hierarchical performance, computational cost and energy consumption, external validation on EPO patents, and the propagation of classification errors into downstream technological indicators, with implications for automated patent classification procedures used by patent offices, technology analysts, and scientometric researchers.

cs.CE

The hidden structure of innovation networks

Innovation emerges from complex collaboration patterns - among inventors, firms, or institutions. However, not much is known about the overall mesoscopic structure around which inventive activity self-organizes. Here, we tackle this problem by employing patent data to analyze both individual (\textit{co-inventorship}) and organization (\textit{co-ownership}) networks in three strategic domains (\textit{artificial intelligence}, \textit{biotechnology} and \textit{semiconductors}). We characterize the mesoscale structure (in terms of clusters) of each domain by comparing two alternative methods: a standard baseline - modularity maximization - and one based on the minimization of the Bayesian Information Criterion, within the Stochastic Block Model and its degree-corrected variant. We find that, across sectors, inventor networks are denser and more clustered than organization ones - consistently with the presence of small recurrent teams embedded into broader institutional hierarchies - whereas organization networks have a neater role-based structures, with few bridging firms coordinating the most peripheral ones; still, both are characterized by the presence of local core-periphery structures. We also find that the discovered meso-structures are connected to innovation output. In particular, Lorenz curves of forward citations show a pervasive inequality in technological influence: across sectors and methods, both inventor (especially) and organization networks consistently show high levels of concentration of citations in a few of the discovered clusters. Our results demonstrate that the baseline modularity-based method may not be capable of fully capturing the way collaborations drive the spreading of inventive impact across technological domains. This is due to the presence of local hierarchies that call for the more refined tools of Bayesian inference.

econ.GN

The anatomy of Green AI technologies: structure, evolution, and impact

Artificial intelligence (AI) is a key enabler of innovation against climate change. In this study, we investigate the intersection of AI and climate adaptation and mitigation technologies through patent analyses of a novel dataset of approximately 63 000 Green AI patents. We analyze patenting trends, corporate ownership of the technology, the geographical distributions of patents, their impact on follow-on inventions and their market value. We use topic modeling (BERTopic) to identify 16 major technological domains, track their evolution over time, and identify their relative impact. We uncover a clear shift from legacy domains such as combustion engines technology to emerging areas like data processing, microgrids, and agricultural water management. We find evidence of growing concentration in corporate patenting against a rapidly increasing number of patenting firms. Looking at the technological and economic impact of patents, while some Green AI domains combine technological impact and market value, others reflect weaker private incentives for innovation, despite their relevance for climate adaptation and mitigation strategies. This is where policy intervention might be required to foster the generation and use of new Green AI applications.

econ.GN