SearcharxivSearch

arXiv subjects

Ronaldo Menezes

Publications and source records attributed to Ronaldo Menezes.

At least 19 recordsLinked to original sources

A Shared IPTC Topic Space for Cross-Source Topic Modelling

Comparing topic attention across different media is hindered by a fundamental modelling problem: topic models fitted separately to each corpus produce corpus-specific topic spaces that cannot be aligned directly. This paper presents a reproducible framework that places corpora in a single shared topic space defined by a taxonomy. Discovered topics are obtained with guided BERTopic, scored against the ninety-four IPTC Media Topics' taxonomy topics (level-1) through weighted keyword and target centroids, and then collapsed upward to seventeen IPTC parent topics by a maximum-similarity rule. The framework was developed and selected on a controlled New York Times 2011 corpus through a narrowing sequence: a broad model screen, a focused mapping refinement, a strict finalist comparison, a target-construction ablation, and a threshold calibration. In this corpus, the guided family retained substantially stronger mapped coverage than a zero-shot benchmark under stricter assignment thresholds, a parent-enriched target construction improved both coverage and parent consistency, and coverage declined gradually rather than collapsing as the assignment threshold was tightened. The contribution is an externally anchored method for constructing a shared topic space that enables reproducible cross-source topic comparison.

cs.IR

Vector fields as a framework for modelling the mobility of commodities

Commodities, including livestock, flow through trade networks globally, with trajectories that can be effectively captured using mobility pattern modelling approaches similar to those used in human mobility studies. However, documenting these movements comprehensively presents significant challenges; it can be unrealistic, costly, and may conflict with data protection regulations. As a result, mobility datasets typically contain uncertainties due to sparsity and limitations in data collection. Origin-destination (OD) representations offer a powerful framework for modelling movement patterns and are widely adopted in mobility studies. However, these matrices possess inherent limitations: locations absent from the OD framework lack spatial information on potential mobility directions and intensities. This spatial incompleteness creates analytical gaps across different geographical scales, constraining our ability to characterise movement patterns in underrepresented areas. In this study, we introduce a vector-field-based method to address these data challenges, transforming OD data into vector fields capturing spatial flow patterns comprehensively enabling us to study mobility directions solidly. We use cattle trade data from Minas Gerais, Brazil, as our case study for commodity flows. This region's large livestock trading network makes it an ideal test case. Cattle movements are significant as they affect disease transmission, including foot-and-mouth disease. Accurately modelling these flows allows better surveillance and control strategies. Our vector-field approach reveals fundamental patterns in commodity mobility and can infer movement information for unrepresented locations. Our approach offers an alternative to traditional network-based models, enhancing our capacity to infer mobility patterns from incomplete datasets and advancing our understanding of large-scale commodity trades.

physics.soc-ph

Super-Linear Growth and Rising Inequality in Online Social Communities: Insights from Reddit

We study the effect of the number of users on the activity of communities within the online content sharing and discussion platform Reddit, called subreddits. We found that comment activity on Reddit has a heavy-tailed distribution, where a large fraction of the comments are made by a small set of users. Furthermore, as subreddits grow in size, this behavior becomes stronger, with activity (measured by the comments made in a subreddit) becoming even more centralised in a (relatively) smaller core of users. We verify that these changes are not explained by finite size nor by sampling effects. Instead, we observe a systematic change of the distribution with subreddit size. To quantify the centralisation and inequality of activity in a subreddit, we used the Gini coefficient. We found that as subreddits grow in users, so does the Gini coefficient, seemingly as a natural effect of the scaling. We found that the excess number of comments (the total number of comments minus the total number of users) follows a power law with exponent 1.27. For each subreddit we considered a snapshot of one month of data, as a compromise between statistical relevance and change in the system's dynamics. We show results over the whole year 2021 (with each subreddit having twelve snapshots, at most), nevertheless all results were consistent when using a single month or different years.

physics.soc-ph

Validating Urban Scaling Laws through Mobile Phone Data: A Continental-Scale Analysis of Brazil's Largest Cities

\abstract{Urban scaling theories posit that larger cities exhibit disproportionately higher levels of socioeconomic activity and human interactions. Yet, evidence from developing contexts (especially those marked by stark socioeconomic disparities) remains limited. To address this gap, we analyse a month-long dataset of 3.1~billion voice-call records from Brazil's 100 most populous cities, providing a continental-scale test of urban scaling laws. We measure interactions using two complementary proxies: the number of phone-based contacts (voice-call degrees) and the number of trips inferred from consecutive calls in distinct locations. Our findings reveal clear superlinear relationships in both metrics, indicating that larger urban centres exhibit intensified remote communication and physical mobility. We further observe that gross domestic product (GDP) also scales superlinearly with population, consistent with broader claims that economic output grows faster than city size. Conversely, the number of antennas required per user scales sublinearly, suggesting economies of scale in telecommunications infrastructure. Although the dataset covers a single provider, its widespread coverage in major cities supports the robustness of the results. We nonetheless discuss potential biases, including city-specific marketing campaigns and predominantly prepaid users, as well as the open question of whether higher interaction drives wealth or vice versa. Overall, this study enriches our understanding of urban scaling, emphasising how communication and mobility jointly shape the socioeconomic landscapes of rapidly growing cities.

physics.soc-ph

Systematic comparison of gender inequality in scientific rankings across disciplines

The participation of women in academia has increased in the last few decades across many fields (e.g., Computer Science, History, Medicine). However, this increase in the participation of women has not been the same at all career stages. Here, we study how gender participation within different fields is related to gender representation in top-ranking positions in productivity (number of papers), research impact (number of citations), and co-authorship networks (degree of connectivity). We analyzed over 80 million papers published from 1975 to 2020 in 19 academic fields. Our findings reveal that women remain a minority in all 19 fields, with physics, geology, and mathematics having the lowest percentage of papers authored by women at 14% and psychology having the largest percentage at 39%. Women are significantly underrepresented in top-ranking positions (top 10% or higher) across all fields and metrics (productivity, citations, and degree), indicating that it remains challenging for early researchers (especially women) to reach top-ranking positions, as our results reveal the rankings to be rigid over time. Finally, we show that in most fields, women and men with comparable productivity levels and career age tend to attain different levels of citations, where women tend to benefit more from co-authorships, while men tend to benefit more from productivity, especially in pSTEMs. Our findings highlight that while the participation of women has risen in some fields, they remain under-represented in top-ranking positions. Greater gender participation at entry levels often helps representation, but stronger interventions are still needed to achieve long-lasting careers for women and their participation in top-ranking positions.

cs.SI

The parenthood effect in urban mobility

We investigate how parenthood and marriage (two major life events) reshape urban mobility patterns, an aspect overlooked in traditional `average citizen' mobility models. Leveraging US census data, we analyse whether these life transitions create distinct urban experiences. Parenthood introduces new priorities including caregiving responsibilities, work-life balance adjustments, and access to family-friendly environments. Similarly, marriage introduces new dynamics including shared household decision-making, potential dual-income benefits, combined residential preferences, and shifts in social networks and lifestyle patterns. Our analysis demonstrates that cities vary significantly in how mobility can be accommodated by different household arrangements: some better accommodate either single individuals (Houston, Virginia Beach) or married people (Atlanta, Baltimore), whereas others favour parents (Cincinnati, Chicago). This classification becomes increasingly relevant for individuals and families as remote work expands relocation possibilities. We find that parents and married individuals face different mobility costs and amenity access patterns compared to their counterparts, with variations consistent across multiple null model tests. This research advances urban planning discourse by advocating for tailored design strategies addressing diverse demographic needs rather than one-size-fits-all approaches.

physics.soc-ph

Understanding the Structure and Resilience of the Brazilian Federal Road Network Through Network Science

Understanding how transportation networks work is important for improving connectivity, efficiency, and safety. In Brazil, where road transport is a significant portion of freight and passenger movement, network science can provide valuable insights into the structural properties of the infrastructure, thus helping decision makers responsible for proposing improvements to the system. This paper models the federal road network as weighted networks, with the intent to unveil its topological characteristics and identify key locations (cities) that play important roles for the country through 75,000 kilometres of roads. We start with a simple network to examine basic connectivity and topology, where weights are the distance of the road segment. We then incorporate other weights representing number of incidents, population, and number of cities in-between each segment. We then focus on community detection as a way to identify clusters of cities that form cohesive groups within a network. Our findings aim to bring clarity to the overall structure of federal roads in Brazil, thus providing actionable insights for improving infrastructure planning and prioritising resources to enhance network resilience.

physics.soc-ph

Geographical Isolation as a Driver of Political Violence in African Cities

Violence is commonly linked with large urban areas, and as a social phenomenon, it is presumed to scale super-linearly with population size. This study explores the hypothesis that smaller, isolated cities in Africa may experience a heightened intensity of violence against civilians. It aims to investigate the correlation between the risk of experiencing violence with a city's size and its geographical isolation. Over a 20-year period, the incidence of civilian casualties has been analysed to assess lethality in relation to varying degrees of isolation and city sizes. African cities are categorised by isolation (number of highway connections) and centrality (the estimated frequency of journeys). Findings suggest that violence against civilians exhibits a sub-linear pattern, with larger cities witnessing fewer casualties per 100,000 inhabitants. Remarkably, individuals in isolated cities face a quadrupled risk of a casualty compared to those in more connected cities.

physics.soc-ph

Environmental Insights: Democratizing Access to Ambient Air Pollution Data and Predictive Analytics with an Open-Source Python Package

Ambient air pollution is a pervasive issue with wide-ranging effects on human health, ecosystem vitality, and economic structures. Utilizing data on ambient air pollution concentrations, researchers can perform comprehensive analyses to uncover the multifaceted impacts of air pollution across society. To this end, we introduce Environmental Insights, an open-source Python package designed to democratize access to air pollution concentration data. This tool enables users to easily retrieve historical air pollution data and employ a Machine Learning model for forecasting potential future conditions. Moreover, Environmental Insights includes a suite of tools aimed at facilitating the dissemination of analytical findings and enhancing user engagement through dynamic visualizations. This comprehensive approach ensures that the package caters to the diverse needs of individuals looking to explore and understand air pollution trends and their implications.

physics.soc-ph

A Data-Driven Supervised Machine Learning Approach to Estimating Global Ambient Air Pollution Concentrations With Associated Prediction Intervals

Global ambient air pollution, a transboundary challenge, is typically addressed through interventions relying on data from spatially sparse and heterogeneously placed monitoring stations. These stations often encounter temporal data gaps due to issues such as power outages. In response, we have developed a scalable, data-driven, supervised machine learning framework. This model is designed to impute missing temporal and spatial measurements, thereby generating a comprehensive dataset for pollutants including NO$_2$, O$_3$, PM$_{10}$, PM$_{2.5}$, and SO$_2$. The dataset, with a fine granularity of 0.25$^{\circ}$ at hourly intervals and accompanied by prediction intervals for each estimate, caters to a wide range of stakeholders relying on outdoor air pollution data for downstream assessments. This enables more detailed studies. Additionally, the model's performance across various geographical locations is examined, providing insights and recommendations for strategic placement of future monitoring stations to further enhance the model's accuracy.

cs.LG

A Framework for Scalable Ambient Air Pollution Concentration Estimation

Ambient air pollution remains a critical issue in the United Kingdom, where data on air pollution concentrations form the foundation for interventions aimed at improving air quality. However, the current air pollution monitoring station network in the UK is characterized by spatial sparsity, heterogeneous placement, and frequent temporal data gaps, often due to issues such as power outages. We introduce a scalable data-driven supervised machine learning model framework designed to address temporal and spatial data gaps by filling missing measurements. This approach provides a comprehensive dataset for England throughout 2018 at a 1kmx1km hourly resolution. Leveraging machine learning techniques and real-world data from the sparsely distributed monitoring stations, we generate 355,827 synthetic monitoring stations across the study area, yielding data valued at approximately \pounds70 billion. Validation was conducted to assess the model's performance in forecasting, estimating missing locations, and capturing peak concentrations. The resulting dataset is of particular interest to a diverse range of stakeholders engaged in downstream assessments supported by outdoor air pollution concentration data for NO2, O3, PM10, PM2.5, and SO2. This resource empowers stakeholders to conduct studies at a higher resolution than was previously possible.

stat.AP

Dynamic predictability and spatio-temporal contexts in human mobility

Human travelling behaviours are markedly regular, to a large extent, predictable, and mostly driven by biological necessities (\eg sleeping, eating) and social constructs (\eg school schedules, synchronisation of labour). Not surprisingly, such predictability is influenced by an array of factors ranging in scale from individual (\eg preference, choices) and social (\eg household, groups) all the way to global scale (\eg mobility restrictions in a pandemic). In this work, we explore how spatio-temporal patterns in individual-level mobility, which we refer to as \emph{predictability states}, carry a large degree of information regarding the nature of the regularities in mobility. Our findings indicate the existence of contextual and activity signatures in predictability states, pointing towards the potential for more sophisticated, data-driven approaches to short-term, higher-order mobility predictions beyond frequentist/probabilistic methods.

physics.soc-ph

COVID-19 is linked to changes in the time-space dimension of human mobility

Socio-economic constructs and urban topology are crucial drivers of human mobility patterns. During the coronavirus disease 2019 pandemic, these patterns were reshaped in their components: the spatial dimension represented by the daily travelled distance, and the temporal dimension expressed as the synchronization time of commuting routines. Here, leveraging location-based data from de-identified mobile phone users, we observed that, during lockdowns restrictions, the decrease of spatial mobility is interwoven with the emergence of asynchronous mobility dynamics. The lifting of restriction in urban mobility allowed a faster recovery of the spatial dimension compared with the temporal one. Moreover, the recovery in mobility was different depending on urbanization levels and economic stratification. In rural and low-income areas, the spatial mobility dimension suffered a more considerable disruption when compared with urbanized and high-income areas. In contrast, the temporal dimension was more affected in urbanized and high-income areas than in rural and low-income areas.

physics.soc-ph

Does Transport Inequality Perpetuate Housing Insecurity?

With trends of urbanisation on the rise, providing adequate housing to individuals remains a complex issue to be addressed. Often, the slow output of relevant housing policies, coupled with quickly increasing housing costs, leaves individuals with the burden of finding housing that is affordable and safe. In this paper, we unveil how urban planning, not just housing policies, can prevent individuals from accessing better housing conditions. We begin by proposing a clustering approach to characterising levels of housing insecurity in a city, by considering multiple dimensions of housing. Then we define levels of transit efficiency in 20 US cities by comparing public transit journeys to car-based journeys. Finally, we use geospatial autocorrelation to highlight how commuting to areas associated with better housing conditions results in transit commute times of over 30 minutes in most cities, and commute times of over an hour in some cases. Ultimately, we show the role that public transportation plays in locking vulnerable demographics into a cycle of poverty, thus motivating a more holistic approach to addressing housing insecurity that extends beyond changing housing policies.

physics.soc-ph

The structure of segregation in co-authorship networks and its impact on scientific production

Co-authorship networks, where nodes represent authors and edges represent co-authorship relations, are key to understanding the production and diffusion of knowledge in academia. Social constructs, biases (implicit and explicit), and constraints (e.g. spatial, temporal) affect who works with whom and cause co-authorship networks to organise into tight communities with different levels of segregation. We aim to look at aspects of the co-authorship network structure that lead to segregation and its impact on scientific production. We measure segregation using the Spectral Segregation Index (SSI) and find 4 ordered segregation categories: completely segregated, highly segregated, moderately segregated and non-segregated communities. We direct our attention to the non-segregated and highly segregated communities, quantifying and comparing their structural topologies and k-core positions. When considering communities of both categories (controlling for size), our results show no differences in density and clustering but substantial variability in core position. Larger non-segregated communities are more likely to occupy cores near the network nucleus, while the highly segregated ones tend to be closer to the network periphery. Finally, we analyse differences in citations gained by researchers within communities showing different segregation categories. Researchers in highly segregated communities get more citations from their community members in middle cores and gain more citations per publication in middle/periphery cores. Those in non-segregated communities get more citations per publication in the nucleus. To our knowledge, this work is the first to characterise community segregation in co-authorship networks and investigate the relationship between community segregation and author citations.

cs.SI

Network Entropy as a Measure of Socioeconomic Segregation in Residential and Employment Landscapes

Cities create potential for individuals from different backgrounds to interact with one another. It is often the case, however, that urban infrastructure obfuscates this potential, creating dense pockets of affluence and poverty throughout a region. The spatial distribution of job opportunities, and how it intersects with the residential landscape, is one of many such obstacles. In this paper, we apply global and local measures of entropy to the commuting networks of 25 US cities to capture structural diversity in residential and work patterns. We identify significant relationships between the heterogeneity of commuting origins and destinations with levels of employment and residential segregation, respectively. Finally, by comparing the local entropy values of low and high-income networks, we highlight how disparities in entropy are indicative of both employment segregation and residential inhomogeneities. Ultimately, this work motivates the application of network entropy to understand segregation not just from a residential perspective, but an experiential one as well.

physics.soc-ph

Mobility and Transit Segregation in Urban Spaces

Segregation is a highly nuanced concept that researchers have worked to define and measure over the past several decades. Conventional approaches tend to estimate segregation based on residential patterns in a static manner. In this work, we analyse socioeconomic inequalities, assessing segregation in various dimensions of the urban experience. Moreover, we consider the pivotal role that transport plays in democratising access to opportunities. Using transport networks, amenity visitations, and census data, we develop a framework to approximate segregation, within the United States, for various dimensions of urban life. We find that neighbourhoods that are segregated in the residential domain, tend to exhibit similar levels of segregation in amenity visitation patterns and transit usage, albeit to a lesser extent. We identify inequalities embedded into transit service, which impose constraints on residents from segregated areas, limiting the neighbourhoods that they can access within an hour to areas that are similarly disadvantaged. By exploring socioeconomic segregation from a transit perspective, we underscore the importance of conceptualising segregation as a dynamic measure, while also highlighting how transport systems can contribute to a cycle of disadvantage.

physics.soc-ph

Estimating Ambient Air Pollution Using Structural Properties of Road Networks

In recent years, the world has become increasingly concerned with air pollution. Particularly in the global north, countries are implementing systems to monitor air pollution on a large scale to aid decision-making. Such efforts are essential but they have at least three shortcomings: (1) they are costly and are difficult to implement expediently; (2) they focus on urban areas, which is where most people live, but this choice is prone to inequalities; and (3) the process of estimating air pollution lacks transparency. In this paper, we demonstrate that we can estimate air pollution using open-source information about the structural properties of roads; we focus on England and Wales in the United Kingdom (UK) in this paper although the methods here described are not dependent on specific datasets. Our approach makes it possible to implement an inexpensive method of estimating air pollution concentrations to an accuracy level that can underpin policymakers' decisions while providing an estimate in all districts, not just urban areas, and in a process that is transparent and explainable. Impact Statement. We show that a linear regression model using a single structural property -- length of the track and unclassified road network within 0.36% of districts within England and Wales (in the UK) -- can accurately estimate which districts are the most polluted. The model presents a transparent and low-cost, yet effective, alternative to more expensive models such as the one currently used by DEFRA in the UK. The model has apparent practical uses for policymakers who want to pursue clean-air initiatives but lack the capital to invest in comprehensive monitoring networks. Its low implementation cost, accessible model design, and worldwide coverage of the dataset provide a basis for implementing systems to estimate air pollution concentrations in low-income countries.

physics.soc-ph