SearcharxivSearch

arXiv subjects

Frank Neffke

Publications and source records attributed to Frank Neffke.

18 recordsLinked to original sources

The Changing Global Division of Labor in Software: Emergence and Diffusion of New Programming Skills across IT Hubs

With the rise of new industries, often new jobs emerge. Evolutionary Economic Geography and in particular Industry Life Cycle perspectives predict that these activities first emerge in a limited number of cities to then diffuse to other locations as job descriptions become more standardized. Here, we focus on a particularly important new industry: software development, an activity that is economically important, quickly changing, and has a pronounced spatial concentration in a small number of global IT hubs. We use an online database of over 60 million questions and answers about problems in software development that yields a longitudinal dataset of 237 software skills. By geo-locating 3 million posting users at regular intervals, we link these skills to cities worldwide. We find that, in spite of its digital nature, the software industry exhibits similar spatial regularities as previously observed in more traditional sectors. First, cities diversify into skills that are related to their existing ones. Second, new skills first emerge in cities with large and diversified software sectors, and later diffuse -- mostly unhindered by geographical distance -- to smaller cities specialized in closely related skills. We find suggestive but limited support for a windows of locational opportunity account: although even brand-new skills still emerge first in cities with strong prior specialization in related skills, concentrations of related activities impact less the emergence of new skills than the diffusion of existing ones.

cs.SI

Large cities lose their growth advantage as countries urbanize

The share of the world population living in cities with more than one million people rose from 11% in 1975 to 24% in 2025 (our estimates). Will this trend towards greater concentration in large cities continue or level off? We introduce two new city population datasets that use consistent city definitions across countries and over time. The first covers the world between 1975 and 2025, using satellite imagery. The second covers the U.S. between 1850 and 2020, using census microdata. We find that urban growth follows a characteristic life cycle. In the early stages of a country's urbanization process, large cities grow faster than smaller ones. At later stages, growth rates equalize across sizes. We use this life cycle to project future population concentration in large cities. Our projections suggest that 38% of the world population will be living in cities with more than one million people by 2100. This estimate is higher than the 33% implied by the well-known theory of proportional growth, but lower than the 42% obtained by extrapolating current trends.

physics.soc-ph

Who is using AI to code? Global diffusion and impact of generative AI

Generative coding tools promise big productivity gains, but uneven uptake could widen skill and income gaps. We train a neural classifier to spot AI-generated Python functions in over 30 million GitHub commits by 170,000 developers, tracking how fast -- and where -- these tools take hold. Today, AI writes an estimated 29% of Python functions in the US, a modest and shrinking lead over other countries. We estimate that quarterly output, measured in online code contributions, has increased by 3.6% because of this. Our evidence suggests that programmers using AI may also more readily expand into new domains of software development. However, experienced programmers capture nearly all of these productivity and exploration gains, widening rather than closing the skill gap.

cs.CY

Synthesis of innovation and obsolescence

Innovation and obsolescence describe the dynamics of ever-churning social and biological systems, from the development of economic markets to scientific and technological progress to biological evolution. They have been widely discussed, but in isolation, leading to fragmented modeling of their dynamics. This poses a problem for connecting and building on what we know about their shared mechanisms. Here we collectively propose a conceptual and mathematical framework to transcend field boundaries and to explore unifying theoretical frameworks and open challenges. We ring an optimistic note for weaving together disparate threads with key ideas from the wide and largely disconnected literature by focusing on the duality of innovation and obsolescence and by proposing a mathematical framework to unify the metaphors between constitutive elements.

physics.soc-ph

Green building blocks reveal the complex anatomy of climate change mitigation technologies

Achieving net-zero emissions requires rapid innovation, yet the necessary technological knowhow is scattered across industries and countries. Comparing functionally similar green and nongreen patents, we identify "Green Building Blocks" (GBBs): modular components that can be added to reduce existing technologies' carbon footprints. These GBBs depict the anatomy of the green transition as a network that connects problems -- nongreen technologies -- to GBBs that mitigate their climate-change impact. Node degrees in this network are highly unequal, showing that the scope for climate-change mitigating innovation varies substantially across domains. The network also helps predict which green technologies firms develop themselves, and which alliances they form to do so. This reveals a critical dependence on international collaboration: optimal innovation partners for 84% of US, 87% of German, and 92% of Chinese firms are foreign, providing quantitative evidence that rising economic nationalism threatens the pace of innovation required to meet global climate goals.

physics.soc-ph

Using digital traces to analyze software work: skills, careers and programming languages

Recent waves of technological transformation are reshaping work in uncertain and hard-to-predict ways. However, jobs at the forefront of the digitizing economy offer an early glimpse of these changes and leave rich activity traces. We exploit such traces in tens of millions of Question and Answer posts on Stack Overflow for the creation of a fine-grained taxonomy of software skills to analyze human capital in the global software industry. Constructing a software skill space that maps relations among these skills reveals that real-world software jobs demand highly coherent skill sets and that programmers learn through a process of related diversification. The latter process often leads to the acquisition of lower-value skills. However, when programmers use Python they preferentially target higher-value skills, offering a potential explanation for Python's successful rise as a dominant general purpose language.

econ.GN

The Coherence of US cities

Diversified economies are critical for cities to sustain their growth and development, but they are also costly because diversification often requires expanding a city's capability base. We analyze how cities manage this trade-off by measuring the coherence of the economic activities they support, defined as the technological distance between randomly sampled productive units in a city. We use this framework to study how the US urban system developed over almost two centuries, from 1850 to today. To do so, we rely on historical census data, covering over 600M individual records to describe the economic activities of cities between 1850 and 1940, and 8 million patent records as well as detailed occupational and industrial profiles of cities for more recent decades. Despite massive shifts in the economic geography of the U.S. over this 170-year period, average coherence in its urban system remains unchanged. Moreover, across different time periods, datasets and relatedness measures, coherence falls with city size at the exact same rate, pointing to constraints to diversification that are governed by a city's size in universal ways.

physics.soc-ph

Close to Home: Analyzing Urban Consumer Behavior and Consumption Space in Seoul

This study explores how the relatedness density of amenities influences consumer buying patterns, focusing on multi-purpose shopping preferences. Using Seoul's credit card data from 2018 to 2023, we find a clear preference for shopping at amenities close to consumers' residences, particularly for trips within a 2 km radius, where relatedness density significantly influences purchasing decisions. The COVID-19 pandemic initially reduced this effect at shorter distances but rebounded in 2023, suggesting a resilient return to pre-pandemic patterns, which vary over regions. Our findings highlight the resilience of local shopping preferences despite economic disruptions, underscoring the importance of amenity-relatedness in urban consumer behavior.

econ.GN

Colocation of skill related suppliers -- Revisiting coagglomeration using firm-to-firm network data

Strong local clusters help firms compete on global markets. One explanation for this is that firms benefit from locating close to their suppliers and customers. However, the emergence of global supply chains shows that physical proximity is not necessarily a prerequisite to successfully manage customer-supplier relations anymore. This raises the question when firms need to colocate in value chains and when they can coordinate over longer distances. We hypothesize that one important aspect is the extent to which supply chain partners exchange not just goods but also know-how. To test this, we build on an expanding literature that studies the drivers of industrial coagglomeration to analyze when supply chain connections lead firms to colocation. We exploit detailed micro-data for the Hungarian economy between 2015 and 2017, linking firm registries, employer-employee matched data and firm-to-firm transaction data from value-added tax records. This allows us to observe colocation, labor flows and value chain connections at the level of firms, as well as construct aggregated coagglomeration patterns, skill relatedness and input-output connections between pairs of industries. We show that supply chains are more likely to support coagglomeration when the industries involved are also skill related. That is, input-output and labor market channels reinforce each other, but supplier connections only matter for colocation when industries have similar labor requirements, suggesting that they employ similar types of know-how. We corroborate this finding by analyzing the interactions between firms, showing that supplier relations are more geographically constrained between companies that operate in skill related industries.

physics.soc-ph

Skill dependencies uncover nested human capital

Modern economies require increasingly diverse and specialized skills, many of which depend on the acquisition of other skills first. Here we analyse US survey data to reveal a nested structure within skill portfolios, where the direction of dependency is inferred from asymmetrical conditional probabilities-occupations require one skill conditional on another. This directional nature suggests that advanced, specific skills and knowledge are often built upon broader, fundamental ones. We examine 70 million job transitions to show that human capital development and career progression follow this structured pathway in which skills more aligned with the nested structure command higher wage premiums, require longer education and are less likely to be automated. These disparities are evident across genders and racial-ethnic groups, explaining long-term wage penalties. Finally, we find that this nested structure has become even more pronounced over the past two decades, indicating increased barriers to upward job mobility.

physics.soc-ph

Information consumption and size in firms

Social and biological collectives need to exchange information to persist and to function. This happens across internal networks, whose structure represents static channels through which information flows. Less studied is the quantity and variety of information transmitted. We characterize a part of the information flow, the information going into organizations, primarily business firms. We measure what firms read using a data set of hundreds of millions of records of news articles accessed by employees across millions of firms. We measure and relate quantitatively three essential aspects: reading volume, reading variety, and firm size. First we compare volume with firm size, showing that firms grow sublinearly with the volume of their reading. The scaling means that inequality in information volume exaggerates the classic Zipf's law inequality in firm size, pointing to an economy of scale in information consumption. Then, by connecting variety and volume, we show that the firms vary in their reading habits to a limited degree. Firms above a certain size become repetitive readers, consistent with the sudden onset of a coordination cost between teams, not individual employees. Finally, we relate information variety to size to show that large firms tend to increase investments in existing areas of interest instead of divesting from them to move to new areas. We argue that this reflects structural constraints in growth. The results indicate how information consumption reflects the role of internal structure, beyond individual employees, analogous to information processing in other social and biological systems.

physics.soc-ph

Evaluating the principle of relatedness: Estimation, drivers and implications for policy

A growing body of research documents that the size and growth of an industry in a place depends on how much related activity is found there. This fact is commonly referred to as the "principle of relatedness". However, there is no consensus on why we observe the principle of relatedness, how best to determine which industries are related or how this empirical regularity can help inform local industrial policy. We perform a structured search over tens of thousands of specifications to identify robust -- in terms of out-of-sample predictions -- ways to determine how well industries fit the local economies of US cities. To do so, we use data that allow us to derive relatedness from observing which industries co-occur in the portfolios of establishments, firms, cities and countries. Different portfolios yield different relatedness matrices, each of which help predict the size and growth of local industries. However, our specification search not only identifies ways to improve the performance of such predictions, but also reveals new facts about the principle of relatedness and important trade-offs between predictive performance and interpretability of relatedness patterns. We use these insights to deepen our theoretical understanding of what underlies path-dependent development in cities and expand existing policy frameworks that rely on inter-industry relatedness analysis.

physics.soc-ph

Bridging the short-term and long-term dynamics of economic structural change

Economic transformation -- change in what an economy produces -- is foundational to development and rising standards of living. Our understanding of this process has been propelled recently by two branches of work in the field of economic complexity, one studying how economies diversify, the other how the complexity of an economy is expressed in the makeup of its output. However, the connection between these branches is not well understood, nor how they relate to a classic understanding of structural transformation. Here, we present a simple dynamical modeling framework that unifies these areas of work, based on the widespread observation that economies diversify preferentially into activities that are related to ones they do already. We show how stylized facts of long-run structural change, as well as complexity metrics, can both emerge naturally from this one observation. However, complexity metrics take on new meanings, as descriptions of the long-term changes an economy experiences rather than measures of complexity per se. This suggests relatedness and complexity metrics are connected, in a hitherto overlooked way: Both describe structural change, on different time scales. Whereas relatedness probes transformation on short time scales, complexity metrics capture long-term change.

physics.soc-ph

What can the millions of random treatments in nonexperimental data reveal about causes?

We propose a new method to estimate causal effects from nonexperimental data. Each pair of sample units is first associated with a stochastic 'treatment' - differences in factors between units - and an effect - a resultant outcome difference. It is then proposed that all such pairs can be combined to provide more accurate estimates of causal effects in observational data, provided a statistical model connecting combinatorial properties of treatments to the accuracy and unbiasedness of their effects. The article introduces one such model and a Bayesian approach to combine the $O(n^2)$ pairwise observations typically available in nonexperimnetal data. This also leads to an interpretation of nonexperimental datasets as incomplete, or noisy, versions of ideal factorial experimental designs. This approach to causal effect estimation has several advantages: (1) it expands the number of observations, converting thousands of individuals into millions of observational treatments; (2) starting with treatments closest to the experimental ideal, it identifies noncausal variables that can be ignored in the future, making estimation easier in each subsequent iteration while departing minimally from experiment-like conditions; (3) it recovers individual causal effects in heterogeneous populations. We evaluate the method in simulations and the National Supported Work (NSW) program, an intensively studied program whose effects are known from randomized field experiments. We demonstrate that the proposed approach recovers causal effects in common NSW samples, as well as in arbitrary subpopulations and an order-of-magnitude larger supersample with the entire national program data, outperforming Statistical, Econometrics and Machine Learning estimators in all cases...

stat.ME

An information-theoretic approach to the analysis of location and co-location patterns

We propose a statistical framework to quantify location and co-location associations of economic activities using information-theoretic measures. We relate the resulting measures to existing measures of revealed comparative advantage, localization and specialization and show that they can all be seen as part of the same framework. Using a Bayesian approach, we provide measures of uncertainty of the estimated quantities. Furthermore, the information-theoretic approach can be readily extended to move beyond pairwise co-locations and instead capture multivariate associations. To illustrate the framework, we apply our measures to the co-location of occupations in US cities, showing the associations between different groups of occupations.

stat.AP

Network Backboning with Noisy Data

Networks are powerful instruments to study complex phenomena, but they become hard to analyze in data that contain noise. Network backbones provide a tool to extract the latent structure from noisy networks by pruning non-salient edges. We describe a new approach to extract such backbones. We assume that edge weights are drawn from a binomial distribution, and estimate the error-variance in edge weights using a Bayesian framework. Our approach uses a more realistic null model for the edge weight creation process than prior work. In particular, it simultaneously considers the propensity of nodes to send and receive connections, whereas previous approaches only considered nodes as emitters of edges. We test our model with real world networks of different types (flows, stocks, co-occurrences, directed, undirected) and show that our Noise-Corrected approach returns backbones that outperform other approaches on a number of criteria. Our approach is scalable, able to deal with networks with millions of edges.

physics.soc-ph

Exploring the Uncharted Export: an Analysis of Tourism-Related Foreign Expenditure with International Spend Data

Tourism is one of the most important economic activities in the world: for many countries it represents the single largest product in their export basket. However, it is a product difficult to chart: "exporters" of tourism do not ship it abroad, but they welcome importers inside the country. Current research uses social accounting matrices and general equilibrium models, but the standard industry classifications they use make it hard to identify which domestic industries cater to foreign visitors. In this paper, we make use of open source data and of anonymized and aggregated transaction data giving us insights about the spend behavior of foreigners inside two countries, Colombia and the Netherlands, to inform our research. With this data, we are able to describe what constitutes the tourism sector, and to map the most attractive destinations for visitors. In particular, we find that countries might observe different geographical tourists' patterns -- concentration versus decentralization --; we show the importance of distance, a country's reported wealth and cultural affinity in informing tourism; and we show the potential of combining open source data and anonymized and aggregated transaction data on foreign spend patterns in gaining insight as to the evolution of tourism from one year to another.

q-fin.GN

Report on the Poblacion Flotante of Bogota (D.C.)

In this document we describe the size of the Poblacion Flotante of Bogota (D.C.). The Poblacion Flotante is composed by people who live outside Bogota (D.C.), but who rely on the city for performing their job. We estimate the Poblacion Flotante impact relying on a new data source provided by telecommunications operators in Colombia, which enables us to estimate how many people commute daily from every municipality of Colombia to a specific area of Bogota (D.C.). We estimate that the size of the Poblacion Flotante could represent a 5.4% increase of Bogota (D.C.)'s population. During weekdays, the commuters tend to visit the city center more.

cs.SI