Searcharxiv⌕ Search

arXiv subjects

Liubov Tupikina

Publications and source records attributed to Liubov Tupikina.

At least 19 recordsLinked to original sources

Two-dimensional Taxonomy for N-ary Knowledge Representation Learning Methods

Real-world knowledge can take various forms, including structured, semi-structured, and unstructured data. Among these, Knowledge Graphs (KGs) are structured representations that integrate heterogeneous data sources into structured representations. However, KGs typically reduce complex n-ary relations to simple triples, thereby losing higher-order relational details. In contrast, hypergraphs naturally represent n-ary relations with hyperedges that directly connect multiple entities. Recent advances have led to many hypergraph representation learning methods; however, they often overlook entity roles in hyperedges, limiting fine-grained semantic modelling. To address these issues, Knowledge Hypergraphs (KHGs) and Hyper-relational Knowledge Graphs (HKGs) combine the advantages of KGs and hypergraphs to better capture complex relational structures and role-specific semantics of real-world knowledge. This survey provides a structured review of representation learning methods for n-ary relational data, with a primary focus on static KHGs and HKGs. We propose a two-dimensional taxonomy: the first dimension categorises models based on their methodology, including translation-based models, tensor factorisation-based models, deep neural network-based models, logic rule-based models, and hyperedge expansion-based models. The second dimension classifies models according to their awareness of entity roles and positions in n-ary relations, dividing them into position-aware, role-aware, and aware-less approaches. Finally, we summarise benchmark datasets, training settings, negative sampling strategies, and benchmark-style performance comparisons, and outline open challenges to inspire future research.

cs.LG↗

Topological measures in weighted hypergraphs

Higher-order interactions introduce an additional structural dimension to complex networks, requiring consistent generalizations of classical topological measures. In hypergraphs, the definition of distance between nodes is not unique: beyond the conventional measure derived from clique projection, an alternative formulation that explicitly incorporates the sizes of hyperedges, those of their intersection and their weights has been recently proposed. Here, we generalize three distance-based topological measures, namely closeness centrality, betweenness centrality and node eccentricity, using this new hypergraph distance. Trough tractable illustrative examples, we demonstrate that the differences between results obtained with the two distances are systematic and arise from structurally meaningful features of the higher-order networks. Also, analyzing a series of real-world datasets, we show that hypergraphs can be divided into three distinct classes, corresponding to the possible dominance of specific orders of interaction over their general metric structure. This provides practical guidance on the possibility of limiting the analysis to only some specific interaction orders, reducing its complexity while maintaining the full information of the system.

physics.soc-ph↗

Thinking Before Constraining: A Unified Decoding Framework for Large Language Models

Natural generation allows Large Language Models (LLMs) to produce free-form responses with rich reasoning, yet the lack of structure makes outputs difficult to verify. Conversely, constrained decoding ensures standardized formats but can inadvertently restrict reasoning capabilities by imposing constraints too early in the generation process. We propose a hybrid approach, namely In-Writing, that combines free-form reasoning and structured generation in a single call. The model first performs unconstrained reasoning and only applies structured decoding after a trigger token is generated, explicitly decoupling reasoning from formatting. We establish that our trigger-token strategies are able to virtually eradicate premature triggering, a failure mode in which constrained decoding interrupts on-going reasoning. Evaluations across diverse datasets covering classification and reasoning tasks demonstrate that our approach outperforms the state-of-the-art by achieving accuracy gains of up to 27% over natural generation. Our code are available at: https://github.com/Nokia-Bell-Labs/InWriting.

cs.CL↗

Iterative Ricci-Foster Curvature Flow with GMM-Based Edge Pruning: A Novel Approach to Community Detection

Community detection in complex networks is a fundamental problem, open to new approaches in various scientific settings. We introduce a novel community detection method, based on Ricci flow on graphs. Our technique iteratively updates edge weights (their metric lengths) according to their (combinatorial) Foster version of Ricci curvature computed from effective resistance distance between the nodes. The latter computation is known to be done by pseudo-inverting the graph Laplacian matrix. At that, our approach is alternative to one based on Ollivier-Ricci geometric flow for community detection on graphs, significantly outperforming it in terms of computation time. In our proposed method, iterations of Foster-Ricci flow that highlight network regions of different curvature -- are followed by a Gaussian Mixture Model (GMM) separation heuristic. That allows to classify edges into ''strong'' (intra-community) and ''weak'' (inter-community) groups, followed by a systematic pruning of the former to isolate communities. We benchmark our algorithm on synthetic networks generated from the Stochastic Block Model (SBM), evaluating performance with the Adjusted Rand Index (ARI). Our results demonstrate that proposed framework robustly recovers the planted community structure of SBM-s, establishing Ricci-Foster Flow with GMM-clustering as a principled and computationally effective new tool for network analysis, tested against alternative Ricci-Ollivier flow coupled with spectral clustering.

cs.SI↗

Dissecting embedding method: learning higher-order structures from data

Active area of research in AI is the theory of manifold learning and finding lower-dimensional manifold representation on how we can learn geometry from data for providing better quality curated datasets. There are however various issues with these methods related to finding low-dimensional representation of the data, the so-called curse of dimensionality. Geometric deep learning methods for data learning often include set of assumptions on the geometry of the feature space. Some of these assumptions include pre-selected metrics on the feature space, usage of the underlying graph structure, which encodes the data points proximity. However, the later assumption of using a graph as the underlying discrete structure, encodes only the binary pairwise relations between data points, restricting ourselves from capturing more complex higher-order relationships, which are often often present in various systems. These assumptions together with data being discrete and finite can cause some generalisations, which are likely to create wrong interpretations of the data and models outputs. Hence overall this can cause wrong outputs of the embedding models themselves, while these models being quite and trained on large corpora of data, such as BERT, Yi and other similar models.The objective of our research is twofold, first, it is to develop the alternative framework to characterize the embedding methods dissecting their possible inconsistencies using combinatorial approach of higher-order structures which encode the embedded data. Second objective is to explore the assumption of the underlying structure of embeddings to be graphs, substituting it with the hypergraph and using the hypergraph theory to analyze this structure. We also demonstrate the embedding characterization on the usecase of the arXiv data.

cs.LG↗

Transport collapse in dynamically evolving networks

Transport in complex networks can describe a variety of natural and human-engineered processes including biological, societal and technological ones. However, how the properties of the source and drain nodes can affect transport subject to random failures, attacks or maintenance optimization in the network remain unknown. In this paper, the effects of both the distance between the source and drain nodes and of the degree of the source node on the time of transport collapse are studied in scale-free and lattice-based transport networks. These effects are numerically evaluated for two strategies, which employ either transport-based or random link removal. Scale-free networks with small distances are found to result in larger times of collapse. In lattice-based networks, both the dimension and boundary conditions are shown to have a major effect on the time of collapse. We also show that adding a direct link between the source and the drain increases the robustness of scale-free networks when subject to random link removals. Interestingly, the distribution of the times of collapse is then similar to the one of lattice-based networks.

cond-mat.stat-mech↗

Charting mobility patterns in the scientific knowledge landscape

From small steps to great leaps, metaphors of spatial mobility abound to describe discovery processes. Here, we ground these ideas in formal terms by systematically studying scientific knowledge mobility patterns. We use low-dimensional embedding techniques to create a knowledge space made up of 1.5 million articles from the fields of physics, computer science, and mathematics. By analyzing the publication histories of individual researchers, we discover patterns of knowledge mobility that closely resemble physical mobility. In aggregate, the trajectories form mobility flows that can be described by a gravity model, with jumps more likely to occur in areas of high density and less likely to occur over longer distances. We identify two types of researchers from their individual mobility patterns: interdisciplinary explorers who pioneer new fields, and exploiters who are more likely to stay within their specific areas of expertise. Our results suggest that spatial mobility analysis is a valuable tool for understanding knowledge evolution.

physics.soc-ph↗

iNaturalist citizen science community during City Nature Challenge: new computational approach for analysis of user activity

Analysing patterns of engagement among citizen science participants can provide important insights into the organisation and practice of individual citizen science projects. In particular, methods from statistics and network science can be used to understand different types of user behaviour and user interactions to help the further implementation and organization of community efforts. Using publicly available data from the iNaturalist community and their yearly City Nature Challenges (CNC) from 2017-2020 as an example; we showcase computational methods to explore the spatio-temporal evolution of this citizen science community that typically interacts in a hybrid offline-online way. In particular, we investigate the user types present in the community along with their interactions, finding significant differences in usage-behavior on both the level of engagement and the types of community tasks/roles and how they interact with the network of contributors. We expect that these computational analysis strategies will be useful to gain further understanding of other citizen science communities and projects.

cs.SI↗

Quantifying the rise and fall of scientific fields

Science advances by pushing the boundaries of the adjacent possible. While the global scientific enterprise grows at an exponential pace, at the mesoscopic level the exploration and exploitation of research ideas is reflected through the rise and fall of research fields. The empirical literature has largely studied such dynamics on a case-by-case basis, with a focus on explaining how and why communities of knowledge production evolve. Although fields rise and fall on different temporal and population scales, they are generally argued to pass through a common set of evolutionary stages. To understand the social processes that drive these stages beyond case studies, we need a way to quantify and compare different fields on the same terms. In this paper we develop techniques for identifying scale-invariant patterns in the evolution of scientific fields, and demonstrate their usefulness using 1.5 million preprints from the arXiv repository covering 175 research fields spanning Physics, Mathematics, Computer Science, Quantitative Biology and Quantitative Finance. We show that fields consistently follows a rise and fall pattern captured by a two parameters right-tailed Gumbel temporal distribution. We introduce a field-specific rescaled time and explore the generic properties shared by articles and authors at the creation, adoption, peak, and decay evolutionary phases. We find that the early phase of a field is characterized by the mixing of cognitively distant fields by small teams of interdisciplinary authors, while late phases exhibit the role of specialized, large teams building on the previous works in the field. This method provides foundations to quantitatively explore the generic patterns underlying the evolution of research fields in science, with general implications in innovation studies.

physics.soc-ph↗

Modelling railway delay propagation as diffusion-like spreading

Railway systems form an important means of transport across the world. However, congestions or disruptions may significantly decrease these systems' efficiencies, making predicting and understanding the resulting train delays a priority for railway organisations. Delays are studied in a wide variety of models, which usually simulate trains as discrete agents carrying delays. In contrast, in this paper, we define a novel model for studying delays, where they spread across the railway network via a diffusion-like process. This type of modelling has various advantages such as quick computation and ease of applying various statistical tools like spectral methods, but it also comes with limitations related to the directional and discrete nature of delays and the trains carrying them. We apply the model to the Belgian railways and study its performance in simulating the delay propagation in severely disrupted railway situations. In particular, we discuss the role of spatial aggregation by proposing to cluster the Belgian railway system into sets of stations and adapt the model accordingly. We find that such aggregation significantly increases the model's performance. For some particular situations, a non-trivial optimal level of spatial resolution is found on which the model performs best. Our results show the potential of this type of delay modelling to understand large-scale properties of railway systems.

physics.soc-ph↗

Pseudo-Darwinian evolution of physical flows in complex networks

The evolution of complex transport networks is investigated under three strategies of link removal: random, intentional attack and "Pseudo-Darwinian" strategy. At each evolution step and regarding the selected strategy, one removes either a randomly chosen link, or the link carrying the strongest flux, or the link with the weakest flux, respectively. We study how the network structure and the total flux between randomly chosen source and drain nodes evolve. We discover a universal power-law decrease of the total flux, followed by an abrupt transport collapse. The time of collapse is shown to be determined by the average number of links per node in the initial network, highlighting the importance of this network property for ensuring safe and robust transport against random failures, intentional attacks and maintenance cost optimizations.

physics.soc-ph↗

Continuous limits of Heterogeneous Continuous Time Random Walk model

Continuous time random Walk model has been versatile analytical formalism for studying and modeling diffusion processes in heterogeneous structures, such as disordered or porous media. We are studying the continuous limits of Heterogeneous Continuous Time Random Walk model, when a random walk is making jumps on a graph within different time-length. We apply the concept of a generalized master equation to study heterogeneous continuous-time random walks on networks. Depending on the interpretations of the waiting time distributions the generalized master equation gives different forms of continuous equations.

cond-mat.stat-mech↗

Morphological organization of point-to-point transport in complex networks

We investigate the structural organization of the point-to-point electric, diffusive or hydraulic transport in complex scale-free networks. The random choice of two nodes, a source and a drain, to which a potential difference is applied, selects two tree-like structures, one emerging from the source and the other converging to the drain. These trees merge into a large cluster of the remaining nodes that is found to be quasi-equipotential and thus presents almost no resistance to transport. Such a global "tree-cluster-tree" structure is universal and leads to a power law decay of the currents distribution. Its exponent, $-2$, is determined by the multiplicative decrease of currents at successive branching points of a tree and is found to be independent of the network connectivity degree and resistance distribution.

physics.soc-ph↗

Structural and temporal heterogeneities on networks

A heterogeneous continuous time random walk is an analytical formalism for studying and modeling diffusion processes in heterogeneous structures on microscopic and macroscopic scales. In this paper we study both analytically and numerically the effects of spatio-temporal heterogeneities onto the diffusive dynamics on different types of networks. We investigate how the distribution of the first passage time is affected by the global topological network properties and heterogeneities in the distributions of the travel times. In particular, we analyze transport properties of random networks and define network measures based on the first-passage characteristics. The heterogeneous continuous time random walk framework has potential applications in biology, social and urban science, search of optimal transport properties, analysis of the effects of heterogeneities or bursts in transportation networks.

cond-mat.stat-mech↗

Epidemic extinction in networks: Insights from the 12,110 smallest graphs

We investigate the expected time to extinction in the susceptible-infectious-susceptible (SIS) model of disease spreading. Rather than using stochastic simulations, or asymptotic calculations in network models, we solve the extinction time exactly for all connected graphs with three to eight vertices. This approach enables us to discover descriptive relations that would be impossible with stochastic simulations. It also helps us discovering graphs and configurations of S and I with anomalous behaviors with respect to disease spreading. We find that for large transmission rates the extinction time is independent of the configurations, just dependent on the graph. In this limit, the number of vertices and edges determine the extinction time very accurately (deviations primarily coming from the fluctuations in degrees). We find that the rankings of configurations with respect to extinction times at low and high transmission rates are correlated at low prevalences and negatively correlated for high prevalences. The most important structural factor determining this ranking is the degrees of the infectious vertices.

q-bio.PE↗

Heterogeneous continuous time random walks

We introduce a heterogeneous continuous time random walk (HCTRW) model as a versatile analytical formalism for studying and modeling diffusion processes in heterogeneous structures, such as porous or disordered media, multiscale or crowded environments, weighted graphs or networks. We derive the exact form of the propagator and investigate the effects of spatio-temporal heterogeneities onto the diffusive dynamics via the spectral properties of the generalized transition matrix. In particular, we show how the distribution of first passage times changes due to local and global heterogeneities of the medium. The HCTRW formalism offers a unified mathematical language to address various diffusion-reaction problems, with numerous applications in material sciences, physics, chemistry, biology, and social sciences.

cond-mat.stat-mech↗

Dynamics on networks. Case of Heterogeneous Opinion Status Model

Here we developed a new conceptual, stochastic Heterogeneous Opinion-Status model (HOpS model), which is adaptive network model. The HOpS model admits to identify the main attributes of dynamics on networks and to study analytically the relation between topological network properties and processes taking place on a network. Another key point of the HOpS model is the possibility to study network dynamics via the novel parameter of heterogeneity. We show that not only clear topological network properties, such as node degree, but also, the nodes' status distribution (the factor of network heterogeneity) play an important role in so-called opinion spreading and information diffusion on a network. This model can be potentially used for studying the co-evolution of globally aggregated or averaged key observables of the earth system. These include natural variables such as atmospheric, oceanic and land carbon stocks, as well as socio-economic quantities such as global human population, economic production or wellbeing.

physics.soc-ph↗

Correlation networks from flows. The case of forced and time-dependent advection-diffusion dynamics

Complex network theory provides an elegant and powerful framework to statistically investigate different types of systems such as society, brain or the structure of local and long-range dynamical interrelationships in the climate system. Network links in climate networks typically imply information, mass or energy exchange. However, the specific connection between oceanic or atmospheric flows and the climate network's structure is still unclear. We propose a theoretical approach for verifying relations between the correlation matrix and the climate network measures, generalizing previous studies and overcoming the restriction to stationary flows. Our methods are developed for correlations of a scalar quantity (temperature, for example) which satisfies an advection-diffusion dynamics in the presence of forcing and dissipation. Our approach reveals that correlation networks are not sensitive to steady sources and sinks and the profound impact of the signal decay rate on the network topology. We illustrate our results with calculations of degree and clustering for a meandering flow resembling a geophysical ocean jet.

physics.flu-dyn↗