SearcharxivSearch

arXiv subjects

Kazuki Nakajima

Publications and source records attributed to Kazuki Nakajima.

15 recordsLinked to original sources

Declining Modularity of Intellectual Bases During the Emergence of Research Areas

Understanding how research areas emerge can help identify nascent areas early and inform research strategy, yet how the intellectual base of a field restructures as an area takes shape remains unclear. We hypothesize that the emergence of a research area is accompanied by the integration of largely separate knowledge communities, observable as a decline in the modularity of its co-citation network, which represents its intellectual base. We propose a framework that tracks this modularity over time, evaluates the statistical robustness of its changes, and identifies the papers highly associated with the decline. We applied it to three areas with different modes of growth: higher-order network science, superstring theory, and graph representation learning. In all three, modularity declined in correspondence with each area's emergence or transformation, and in superstring theory, the decline aligns with an independently documented transition. Further analysis of higher-order network science shows that its decline reflects a cross-disciplinary integration. In graph representation learning, the gradual decline is followed by a rise, which we interpret as a re-differentiation after the emergence period. Our results suggest that a decline in the modularity of a co-citation network can serve as a structural signature that retrospectively characterizes this integrative mode of emergence.

physics.soc-ph

Systemic Gendered Citation Imbalance in Computer Science: Evidence from Conferences and Journals

Gender imbalance persists across science, technology, engineering, and mathematics (STEM) fields, including computer science, where it appears in researcher demographics, productivity, recognition, hiring, and career progression. Given computer science's rapid expansion and global influence, addressing this imbalance is essential for broadening participation and fueling innovation. Although journal-oriented disciplines exhibit consistent gender imbalances in citation practices, it remains unclear whether similar patterns arise in the conference-centric culture of computer science. Here, we systematically investigate gender imbalance in citations of conference and journal papers in computer science. We find that papers for which a woman is listed as either first or last author receive fewer citations than expected, partly because of homophilic citation tendencies (i.e., authors tend to cite papers that share specific attributes). This imbalance is especially pronounced for conference papers--particularly those published at top-tier venues--relative to journals. Moreover, we find that the prominence of the first or last author and the structure of their local co-authorship networks are potential drivers of these imbalances. By exploring how conference-centric publishing practices can amplify systemic imbalances in computer science, our study offers insights that may inform efforts to foster more equitable representation in academia.

cs.DL

Researcher Population Pyramids: Tracking Demographic and Gender Trajectories Across Countries

The sustainability of the academic ecosystem relies on researcher demographics and gender balance, yet assessing these dynamics in a timely manner for policy is challenging. Here, we propose a researcher population pyramid framework for tracking demographic and gender trajectories across countries using publication data. We provide a timely snapshot of historical and present demographics and gender balance across 58 countries, revealing three contrasting patterns among research systems: Emerging systems (e.g., Arab countries) exhibit high researcher inflows with widening gender gaps in cumulative productivity; Mature systems (e.g., the United States) show modest inflows with narrowing gender gaps; and Rigid systems (e.g., Japan) lag in both. Furthermore, by simulating future scenarios, the framework makes potential trajectories visible. If 2023 demographic patterns persist, Arab countries' systems could resemble mature or even rigid ones by 2050. Our framework provides a robust diagnostic tool for policymakers worldwide to foster sustainable talent pipelines and gender equality in academia.

cs.DL

Learning Multi-Order Block Structure in Higher-Order Networks

Higher-order networks, naturally described as hypergraphs, are essential for modeling real-world systems involving interactions among three or more entities. Stochastic block models offer a principled framework for characterizing mesoscale organization, yet their extension to hypergraphs involves a trade-off between expressive power and computational complexity. A recent simplification, a single-order model, mitigates this complexity by assuming a single affinity pattern governs interactions of all orders. This universal assumption, however, may overlook order-dependent structural details. Here, we propose a framework that relaxes this assumption by introducing a multi-order block structure, in which different affinity patterns govern distinct subsets of interaction orders. Our framework is based on a multi-order stochastic block model and searches for the optimal partition of the set of interaction orders that maximizes out-of-sample hyperlink prediction performance. Analyzing a diverse range of real-world networks, we find that multi-order block structures are prevalent. Accounting for them not only yields better predictive performance over the single-order model but also uncovers sharper, more interpretable mesoscale organization. Our findings reveal that order-dependent mechanisms are a key feature of the mesoscale organization of real-world higher-order networks.

cs.SI

Akaike information criterion for segmented regression models

In segmented regression, when the regression function is continuous at the change-points that are the boundaries of the segments, it is also called joinpoint regression, and the analysis package developed by \cite{KimFFM00} has become a standard tool for analyzing trends in longitudinal data in the field of epidemiology. In addition, it is sometimes natural to expect the regression function to be discontinuous at the change-points, and in the field of epidemiology, this model is used in \cite{JiaZS22}, which is considered important due to the analysis of COVID-19 data. On the other hand, model selection is also indispensable in segmented regression, including the estimation of the number of change-points; however, it can be said that only BIC-type information criteria have been developed. In this paper, we derive an information criterion based on the original definition of AIC, aiming to minimize the divergence between the true structure and the estimated structure. Then, using the statistical asymptotic theory specific to the segmented regression, we confirm that the penalty for the change-point parameter is 6 in the discontinuous case. On the other hand, in the continuous case, we show that the penalty for the change-point parameter remains 2 despite the rapid change in the derivative coefficients. Through numerical experiments, we observe that our AIC tends to reduce the divergence compared to BIC. In addition, through analyzing the same real data as in \cite{JiaZS22}, we find that the selection between continuous and discontinuous using our AIC yields new insights and that our AIC and BIC may yield different results.

stat.ME

Sampling nodes and hyperedges via random walks on large hypergraphs

Hypergraphs provide a fundamental framework for representing complex systems involving interactions among three or more entities. As empirical hypergraphs grow in size, characterizing their structural properties becomes increasingly challenging due to computational complexity and, in some cases, restricted access to complete data, requiring efficient sampling methods. Random walks offer a practical approach to hypergraph sampling, as they rely solely on local neighborhood information from nodes and hyperedges. In this study, we investigate methods for simultaneously sampling nodes and hyperedges via random walks on large hypergraphs. First, we compare three existing random walks in the context of hypergraph sampling and identify an advantage of the so-called higher-order random walk. Second, by extending an established technique for graphs to the case of hypergraphs, we present a non-backtracking variant of the higher-order random walk. We derive theoretical results on estimators based on the non-backtracking higher-order random walk and validate them through numerical simulations on large empirical hypergraphs. Third, we apply the non-backtracking higher-order random walk to a large hypergraph of co-authorships indexed in the OpenAlex database, where full access to the data is not readily available. Despite the relatively small sample size, our estimates largely align with previous findings on author productivity, team size, and the prevalence of open-access publications. Our findings contribute to the development of analysis methods for large hypergraphs, offering insights into sampling strategies and estimation techniques applicable to real-world complex systems.

cs.SI

Public Perceptions of Fairness Metrics Across Borders

Which fairness metrics are appropriately applicable in your contexts? There may be instances of discordance regarding the perception of fairness, even when the outcomes comply with established fairness metrics. Several questionnaire-based surveys have been conducted to evaluate fairness metrics with human perceptions of fairness. However, these surveys were limited in scope, including only a few hundred participants within a single country. In this study, we conduct an international survey to evaluate public perceptions of various fairness metrics in decision-making scenarios. We collected responses from 1,000 participants in each of China, France, Japan, and the United States, amassing a total of 4,000 participants, to analyze the preferences of fairness metrics. Our survey consists of three distinct scenarios paired with four fairness metrics. This investigation explores the relationship between personal attributes and the choice of fairness metrics, uncovering a significant influence of national context on these preferences.

cs.AI

Inference and Visualization of Community Structure in Attributed Hypergraphs Using Mixed-Membership Stochastic Block Models

Hypergraphs represent complex systems involving interactions among more than two entities and allow the investigation of higher-order structure and dynamics in complex systems. Node attribute data, which often accompanies network data, can enhance the inference of community structure in complex systems. While mixed-membership stochastic block models have been employed to infer community structure in hypergraphs, they complicate the visualization and interpretation of inferred community structure by assuming that nodes may possess soft community memberships. In this study, we propose a framework, HyperNEO, that combines mixed-membership stochastic block models for hypergraphs with dimensionality reduction methods. Our approach generates a node layout that largely preserves the community memberships of nodes. We evaluate our framework on both synthetic and empirical hypergraphs with node attributes. We expect our framework will broaden the investigation and understanding of higher-order community structure in complex systems.

cs.SI

Quantifying gendered citation imbalance in computer science conferences

The number of citations received by papers often exhibits imbalances in terms of author attributes such as country of affiliation and gender. While recent studies have quantified citation imbalance in terms of the authors' gender in journal papers, the computer science discipline, where researchers frequently present their work at conferences, may exhibit unique patterns in gendered citation imbalance. Additionally, understanding how network properties in citations influence citation imbalances remains challenging due to a lack of suitable reference models. In this paper, we develop a family of reference models for citation networks and investigate gender imbalance in citations between papers published in computer science conferences. By deploying these reference models, we found that homophily in citations is strongly associated with gendered citation imbalance in computer science, whereas heterogeneity in the number of citations received per paper has a relatively minor association with it. Furthermore, we found that the gendered citation imbalance is most pronounced in papers published in the highest-ranked conferences, is present across different subfields, and extends to citation-based rankings of papers. Our study provides a framework for investigating associations between network properties and citation imbalances, aiming to enhance our understanding of the structure and dynamics of citations between research publications.

cs.SI

Quantifying gender imbalance in East Asian academia: Research career and citation practice

Gender imbalance in academia has been confirmed in terms of a variety of indicators, and its magnitude often varies from country to country. Europe and North America, which cover a large fraction of research workforce in the world, have been the main geographical regions for research on gender imbalance in academia. However, the academia in East Asia, which accounts for a substantial fraction of research, may be exposed to strong gender imbalance because Asia has been facing persistent and stronger gender imbalance in society at large than Europe and North America. Here we use publication data between 1950 and 2020 to analyze gender imbalance in academia in China, Japan, and South Korea in terms of the number of researchers, their career, and citation practice. We found that, compared to the average of the other countries, gender imbalance is larger in these three East Asian countries in terms of the number of researchers and their citation practice and additionally in Japan in terms of research career. Moreover, we found that Japan has been exposed to the larger gender imbalance than China and South Korea in terms of research career and citation practice.

cs.DL

Higher-order rich-club phenomenon in collaborative research grants

Modern scientific work, including writing papers and submitting research grant proposals, increasingly involves researchers from different institutions. In grant collaborations, it is known that institutions involved in many collaborations tend to densely collaborate with each other, forming rich clubs. Here we investigate higher-order rich-club phenomena in collaborative research grants among institutions and their associations with research productivity. Using publicly available data from the National Science Foundation in the US, we construct a bipartite network of institutions and collaborative grants, which distinguishes among the collaboration with different numbers of institutions. By extending the concept and algorithms of the rich club for dyadic networks to the case of bipartite networks, we find rich clubs both in the entire bipartite network and the bipartite subnetwork induced by the collaborative grants involving a given number of institutions up to five. We also find that the collaborative grants within rich clubs tend to be more productive in a per-dollar sense than the control. Our results highlight advantages of collaborative grants among the institutions in the rich clubs.

physics.soc-ph

Randomizing hypergraphs preserving degree correlation and local clustering

Many complex systems involve direct interactions among more than two entities and can be represented by hypergraphs, in which hyperedges encode higher-order interactions among an arbitrary number of nodes. To analyze structures and dynamics of given hypergraphs, a solid practice is to compare them with those for randomized hypergraphs that preserve some specific properties of the original hypergraphs. In the present study, we propose a family of such reference models for hypergraphs, called the hyper dK-series, by extending the so-called dK-series for dyadic networks to the case of hypergraphs. The hyper dK-series preserves up to the individual node's degree, node's degree correlation, node's redundancy coefficient, and/or the hyperedge's size depending on the parameter values. We also apply the hyper dK-series to numerical simulations of epidemic spreading and evolutionary game dynamics on empirical hypergraphs.

physics.soc-ph

Random Walk Sampling in Social Networks Involving Private Nodes

Analysis of social networks with limited data access is challenging for third parties. To address this challenge, a number of studies have developed algorithms that estimate properties of social networks via a simple random walk. However, most existing algorithms do not assume private nodes that do not publish their neighbors' data when they are queried in empirical social networks. Here we propose a practical framework for estimating properties via random walk-based sampling in social networks involving private nodes. First, we develop a sampling algorithm by extending a simple random walk to the case of social networks involving private nodes. Then, we propose estimators with reduced biases induced by private nodes for the network size, average degree, and density of the node label. Our results show that the proposed estimators reduce biases induced by private nodes in the existing estimators by up to 92.6% on social network datasets involving private nodes.

cs.SI

Social Graph Restoration via Random Walk Sampling

Analyzing social graphs with limited data access is challenging for third-party researchers. To address this challenge, a number of algorithms that estimate structural properties via a random walk have been developed. However, most existing algorithms are limited to the estimation of local structural properties. Here we propose a method for restoring the original social graph from the small sample obtained by a random walk. The proposed method generates a graph that preserves the estimates of local structural properties and the structure of the subgraph sampled by a random walk. We compare the proposed method with subgraph sampling using a crawling method and the existing method for generating a graph that structurally resembles the original graph via a random walk. Our experimental results show that the proposed method more accurately reproduces the local and global structural properties on average and the visual representation of the original graph than the compared methods. We expect that our method will lead to exhaustive analyses of social graphs with limited data access.

cs.SI

Estimating Properties of Social Networks via Random Walk considering Private Nodes

Accurately analyzing graph properties of social networks is a challenging task because of access limitations to the graph data. To address this challenge, several algorithms to obtain unbiased estimates of properties from few samples via a random walk have been studied. However, existing algorithms do not consider private nodes who hide their neighbors in real social networks, leading to some practical problems. Here we design random walk-based algorithms to accurately estimate properties without any problems caused by private nodes. First, we design a random walk-based sampling algorithm that comprises the neighbor selection to obtain samples having the Markov property and the calculation of weights for each sample to correct the sampling bias. Further, for two graph property estimators, we propose the weighting methods to reduce not only the sampling bias but also estimation errors due to private nodes. The proposed algorithms improve the estimation accuracy of the existing algorithms by up to 92.6% on real-world datasets.

cs.SI