SearcharxivSearch

arXiv subjects

Sanjukta Krishnagopal

Publications and source records attributed to Sanjukta Krishnagopal.

At least 19 recordsLinked to original sources

A Network-Based Measure of Cosponsorship Influence on Bill Passing in the United States House of Representatives

Each year, the United States Congress considers thousands of legislative proposals to select bills to present to the US President to sign into law. Naturally, the decision processes of members of Congress are subject to peer influence. In this paper, we examine the effect on bill passage of accrued influence between US Congress members in the US House of Representatives. We explore how the influence of a bill's cosponsors affects the bill's outcome (specifically, whether or not it passes in the House). We define a notion of influence by analyzing the structure of a network that we construct using cosponsorship dynamics. We award `influence' between a pair of Congress members when they cosponsor a bill that achieves some amount of legislative success. We find that properties of the bill cosponsorship network can be a useful signal to examine influence in Congress; they help explain why some bills pass and others fail. We compare our measure of influence to off-the-shelf centrality measures and conclude that our influence measure is more indicative of bill passage.

cs.SI

Sparse Contextual Coupling Reshapes Diffusion Geometry in Multilayer Hypergraphs

Many complex systems combine dense background structure with sparse contextual information. We introduce a diffusion-based framework for analyzing how sparse condition-specific layers reshape diffusion geometry in multilayer hypergraphs. Each layer is represented as a weighted hypergraph, layers are coupled through shared entities, and random walks on the coupled system induce multiscale diffusion distances between nodes. We apply the framework to disease-conditioned gene networks by coupling a dense MSigDB functional gene-set layer to sparse disease-specific DGIdb drug-gene hypergraphs, with disease-associated drugs selected from DDDB and HumanNet-GSP used to define external gene weights. Across Bipolar Disorder, Schizophrenia, Leukemia, and Breast Cancer, the disease-specific layer contains less than 2 percent of genes in the coupled system, yet substantially changes diffusion distances and community structure. Centrality analysis suggests that this disproportionate effect arises because DGIdb-associated genes occupy influential positions in the MSigDB-derived functional network. The resulting diffusion-derived communities are stable under subsampling and show coherent post hoc functional enrichment, including signaling and neurotransmission categories in neuropsychiatric diseases and immune, translational, and metabolic categories in cancer-associated diseases. Community-level comparisons further reveal disease similarities not reducible to direct DGIdb gene overlap, including a Breast Cancer-Schizophrenia relationship consistent with recent biomedical evidence. These results show that sparse contextual layers can induce interpretable nonlocal changes in higher-order network geometry.

physics.soc-ph

Topological Neural Tangent Kernel

Graph neural tangent kernels give a principled infinite-width theory for graph neural networks, but inherit a basic limitation of graph models: they see only pairwise structure. Many relational systems contain higher-order interactions that are more naturally represented by simplicial complexes. We introduce the Topological Neural Tangent Kernel (TopoNTK), an infinite-width kernel for simplicial message passing on edge features. TopoNTK combines lower Hodge interactions, capturing graph-like coupling through shared vertices, with upper Hodge interactions, capturing coupling through filled simplices. This makes the kernel sensitive to topology invisible to graph kernels, allowing complexes with the same graph but different filled simplices to induce different kernels. Beyond expressivity, the Hodge structure gives the kernel an interpretable learning geometry. Edge signals decompose into gradient-like, harmonic, and local circulation components, and the spectrum of the TopoNTK determines how quickly each component is learned. This yields a topological form of spectral bias: components aligned with large-eigenvalue modes are learned quickly, while global harmonic modes, retained through the residual channel, often lie at smaller eigenvalues and are learned more slowly. We prove expressivity, Hodge-alignment, spectral learning, and stability properties, and validate them on synthetic simplicial tasks and DBLP higher-order link prediction. The results show that topology is not merely extra structure; it can provide coordinates that make relational learning more faithful, interpretable, and effective.

cs.LG

Spectral Graph Sparsification Preserves Representation Geometry in Graph Neural Networks

Spectral graph sparsification is a classical tool for reducing graph complexity while preserving Laplacian quadratic forms. In graph neural networks (GNNs), sparsification is often used to accelerate computation while maintaining predictive performance. In this work, we study a complementary representation-level question: does sparsification preserve the geometry of learned embeddings? For polynomial-filter GNNs, we prove that any $ε$-spectral sparsifier induces $O(ε)$ perturbations in polynomial graph filters, multilayer hidden representations, and their Gram matrices. These guarantees imply stability of squared pairwise distances, class means, and covariance structure in embedding space. We further establish finite-time training stability: under smoothness and boundedness assumptions, gradient descent on dense and sparsified graphs produces weight trajectories whose separation grows at most proportionally to the sparsification distortion. Empirically, effective-resistance sparsification validates the predicted perturbation chain on synthetic graphs and preserves hidden representation geometry on real datasets. In our experiments, the gram matrix and training dynamics show low divergence even under substantial sparsification, consistent with the predicted stability under spectral sparsification. Hidden Gram preservation strongly predicts neighborhood preservation and class-centroid stability across FashionMNIST, Cora, and Paul15. Together, these results show that spectral sparsification preserves not only graph operators, but also the representation geometry that supports downstream use of GNN embeddings for interpretability.

cs.LG

Fractal dimensions of complex networks: advocating for a topological approach

Topological Data Analysis (TDA) uses insights from topology to create representations of data able to capture global and local geometric and topological properties. Its methods have successfully been used to develop estimations of fractal dimensions for metric spaces that have been shown to outperform existing techniques. In a parallel line of work, networks are ubiquitously used to model a variety of complex systems. Higher-order interactions, i.e., simultaneous interactions between more than two nodes, are wide-spread in social and biological systems, and simplicial complexes, used in TDA, can capture important structural and topological properties of networks by modelling such higher-order interactions. In this position paper, we advocate for methods from TDA to be used to estimate fractal dimensions of complex networks, we discuss the possible advantages of such an approach and outline some of the challenges to be addressed.

math.AT

Beyond Attention: Learning Spatio-Temporal Dynamics with Emergent Interpretable Topologies

Spatio-temporal forecasting is critical in applications such as traffic prediction, energy demand modeling, and weather monitoring. While Graph Attention Networks (GATs) are popular for modeling spatial dependencies, they rely on predefined adjacency structures and dynamic attention scores, introducing inductive biases and computational overhead that can obscure interpretability. We propose InterGAT, a simplified alternative to GAT that replaces masked attention with a fully learnable, symmetric node interaction matrix, capturing latent spatial relationships without relying on fixed graph topologies. Our framework, InterGAT-GRU, which incorporates a GRU-based temporal decoder, outperforms the baseline GAT-GRU in forecasting accuracy, achieving at least a 21% improvement on the SZ-Taxi dataset and a 6% improvement on the Los-Loop dataset across all forecasting horizons (15 to 60 minutes). Additionally, we observed reduction in training time by 60-70% compared to GAT-GRU baseline. Crucially, the learned interaction matrix reveals interpretable structure: it recovers sparse, topology-aware attention patterns that align with community structure. Spectral and clustering analyses show that the model captures both localized and global dynamics, offering insights into the functional topology driving predictions. This highlights how structure learning can simultaneously support prediction, computational efficiency, and topological interpretabil-ity in dynamic graph-based domains.

cs.LG

Bounded-Confidence Models of Opinion Dynamics with Neighborhood Effects

We generalize bounded-confidence models (BCMs) of opinion dynamics by incorporating neighborhood effects. In a BCM, interacting agents influence each other through dyadic influence if their opinions are sufficiently similar to each other. In our "neighborhood BCMs" (NBCMs), interacting agents are influenced both by each other's opinions and by the opinions of the agents in each other's neighborhoods. Our NBCMs thus include both the usual dyadic influence between agents and a "transitive influence", which encodes the influence of an agent's neighbors, when determining whether or not an interaction changes the opinions of agents. In this transitive influence, an individual's opinion is influenced by a neighbor when, on average, the opinions of the neighbor's neighbors are sufficiently similar to its own opinion. We formulate both neighborhood Deffuant--Weisbuch (NDW) and neighborhood Hegselmann--Krause (NHK) BCMs. We build further on our NBCMs by introducing a neighborhood-based network adaptation in which a network coevolves with agent opinions by changing its structure through "transitive homophily". In this network evolution, an agent breaks a tie to one of its neighbors and then rewires that tie to a new agent, with a preference for agents with a mean neighbor opinion that is closer to its own opinion. Using numerical simulations on a variety of types of networks, we explore how the qualitative opinion dynamics and network properties of our adaptive NDW model change as we adjust the relative proportions of dyadic and transitive influence. In our numerical experiments, we find that incorporating neighborhood effects into the opinion dynamics and the network-adaptation rewiring strategy tends to reduce the spectral gap and degree assortativity of networks. (This is a shortened version of the paper's abstract.)

physics.soc-ph

Hate speech and hate crimes: a data-driven study of evolving discourse around marginalized groups

This study explores the dynamic relationship between online discourse, as observed in tweets, and physical hate crimes, focusing on marginalized groups. Leveraging natural language processing techniques, including keyword extraction and topic modeling, we analyze the evolution of online discourse after events affecting these groups. Examining sentiment and polarizing tweets, we establish correlations with hate crimes in Black and LGBTQ+ communities. Using a knowledge graph, we connect tweets, users, topics, and hate crimes, enabling network analyses. Our findings reveal divergent patterns in the evolution of user communities for Black and LGBTQ+ groups, with notable differences in sentiment among influential users. This analysis sheds light on distinctive online discourse patterns and emphasizes the need to monitor hate speech to prevent hate crimes, especially following significant events impacting marginalized communities.

cs.SI

Topology and dynamics of higher-order multiplex networks

Higher-order networks are gaining significant scientific attention due to their ability to encode the many-body interactions present in complex systems. However, higher-order networks have the limitation that they only capture many-body interactions of the same type. To address this limitation, we present a mathematical framework that determines the topology of higher-order multiplex networks and illustrates the interplay between their topology and dynamics. Specifically, we examine the diffusion of topological signals associated not only to the nodes but also to the links and to the higher-dimensional simplices of multiplex simplicial complexes. We leverage on the ubiquitous presence of the overlap of the simplices to couple the dynamics among multiplex layers, introducing a definition of multiplex Hodge Laplacians and Dirac operators. We show that the spectral properties of these operators determine higher-order diffusion on higher-order multiplex networks and encode their multiplex Betti numbers. Our numerical investigation of the spectral properties of synthetic and real (connectome, microbiome) multiplex simplicial complexes indicates that the coupling between the layers can either speed up or slow down the higher-order diffusion of topological signals. This mathematical framework is very general and can be applied to study generic higher-order systems with interactions of multiple types. In particular, these results might find applications in brain networks which are understood to be both multilayer and higher-order.

nlin.AO

Graph Neural Tangent Kernel: Convergence on Large Graphs

Graph neural networks (GNNs) achieve remarkable performance in graph machine learning tasks but can be hard to train on large-graph data, where their learning dynamics are not well understood. We investigate the training dynamics of large-graph GNNs using graph neural tangent kernels (GNTKs) and graphons. In the limit of large width, optimization of an overparametrized NN is equivalent to kernel regression on the NTK. Here, we investigate how the GNTK evolves as another independent dimension is varied: the graph size. We use graphons to define limit objects -- graphon NNs for GNNs, and graphon NTKs for GNTKs -- , and prove that, on a sequence of graphs, the GNTKs converge to the graphon NTK. We further prove that the spectrum of the GNTK, which is related to the directions of fastest learning which becomes relevant during early stopping, converges to the spectrum of the graphon NTK. This implies that in the large-graph limit, the GNTK fitted on a graph of moderate size can be used to solve the same task on the large graph, and to infer the learning dynamics of the large-graph GNN. These results are verified empirically on node regression and classification tasks.

cs.LG

Local Dirac Synchronization on Networks

We propose Local Dirac Synchronization which uses the Dirac operator to capture the dynamics of coupled nodes and link signals on an arbitrary network. In Local Dirac Synchronization, the harmonic modes of the dynamics oscillate freely while the other modes are interacting non-linearly, leading to a collectively synchronized state when the coupling constant of the model is increased. Local Dirac Synchronization is characterized by discontinuous transitions and the emergence of a rhythmic coherent phase. In this rhythmic phase, one of the two complex order parameters oscillates in the complex plane at a slow frequency (called emergent frequency) in the frame in which the intrinsic frequencies have zero average. Our theoretical results obtained within the annealed approximation are validated by extensive numerical results on fully connected networks and sparse Poisson and scale-free networks. Local Dirac Synchronization on both random and real networks, such as the connectome of Caenorhabditis Elegans, reveals the interplay between topology (Betti numbers and harmonic modes) and non-linear dynamics. This unveils how topology might play a role in the onset of brain rhythms.

cond-mat.dis-nn

The collective vs individual nature of mountaineering: a network and simplicial approach

Mountaineering is a sport of contrary forces: teamwork plays a large role in mental fortitude and skills, but the actual act of climbing, and indeed survival, is largely individualistic. This work studies the effects of the structure and topology of relationships within climbers on the level of cooperation and success. It does so using simplicial complexes, where relationships between climbers are captured through simplexes that correspond to joint previous expeditions with dimension given by the number of climbers minus one and weight given by the number of occurrences of the simplex. First, this analysis establishes the importance of relationships and shows that chances of failure to summit reduce drastically when climbing with repeat partners. From a climber-centric perspective, climbers that belong to simplexes with large dimension were more likely to be successful, across experience levels. From an expedition-centric perspective, the distribution of relationships within a group is explored to identify collective human behavior: from polarized to cooperative. Expeditions containing simplices with large dimension, and usually low weight, i.e., a large number of people had a small number of previous joint expeditions, tended to be more cooperative, with more homogeneity in success amongst climbers. On the other hand, the existence of small, usually strong, subgroups lead to a polarized style where climbers that were not a part of the subgroup were less likely to succeed. Lastly, this work examines the effects of individual features and expedition-wide factors that may play different roles in individualistic and cooperative expeditions. Centrality indicates that individual traits of youth and oxygen use while ascending are strong drivers of success. Of expedition-wide factors, the expedition size and number of expedition days are found to be strongly correlated with success rate.

physics.soc-ph

Spectral Detection of Simplicial Communities via Hodge Laplacians

Despite being a source of rich information, graphs are limited to pairwise interactions. However, several real-world networks such as social networks, neuronal networks, etc., involve interactions between more than two nodes. Simplicial complexes provide a powerful mathematical framework to model such higher-order interactions. It is well known that the spectrum of the graph Laplacian is indicative of community structure, and this relation is exploited by spectral clustering algorithms. Here we propose that the spectrum of the Hodge Laplacian, a higher-order Laplacian defined on simplicial complexes, encodes simplicial communities. We formulate an algorithm to extract simplicial communities (of arbitrary dimension). We apply this algorithm to simplicial complex benchmarks and to real higher-order network data including social networks and networks extracted using language or text processing tools. However, datasets of simplicial complexes are scarce, and for the vast majority of datasets that may involve higher-order interactions, only the set of pairwise interactions are available. Hence, we use known properties of the data to infer the most likely higher-order interactions. In other words, we introduce an inference method to predict the most likely simplicial complex given the community structure of its network skeleton. This method identifies as most likely the higher-order interactions inducing simplicial communities that maximize the adjusted mutual information measured with respect to ground-truth community structure. Finally, we consider higher-order networks constructed through thresholding the edge weights of collaboration networks (encoding only pairwise interactions) and provide an example of persistent simplicial communities that are sustained over a wide range of the threshold.

physics.data-an

Encoding priors in the brain: a reinforcement learning model for mouse decision making

In two-alternative forced choice tasks, prior knowledge can improve performance, especially when operating near the psychophysical threshold. For instance, if subjects know that one choice is much more likely than the other, they can make that choice when evidence is weak. A common hypothesis for these kinds of tasks is that the prior is stored in neural activity. Here we propose a different hypothesis: the prior is stored in synaptic strengths. We study the International Brain Laboratory task, in which a grating appears on either the right or left side of a screen, and a mouse has to move a wheel to bring the grating to the center. The grating is often low in contrast which makes the task relatively difficult, and the prior probability that the grating appears on the right is either 80% or 20%, in (unsignaled) blocks of about 50 trials. We model this as a reinforcement learning task, using a feedforward neural network to map states to actions, and adjust the weights of the network to maximize reward, learning via policy gradient. Our model uses an internal state that stores an estimate of the grating and confidence, and follows Bayesian updates, and can switch between engaged and disengaged states to mimic animal behavior. This model reproduces the main experimental finding - that the psychometric curve with respect to contrast shifts after a block switch in about 10 trials. Also, as seen in the experiments, in our model the difference in neuronal activity in the right and left blocks is small - it is virtually impossible to decode block structure from activity on single trials if noise is about 2%. The hypothesis that priors are stored in weights is difficult to test, but the technology to do so should be available in the not so distant future.

q-bio.NC

Encoded Prior Sliced Wasserstein AutoEncoder for learning latent manifold representations

While variational autoencoders have been successful in several tasks, the use of conventional priors are limited in their ability to encode the underlying structure of input data. We introduce an Encoded Prior Sliced Wasserstein AutoEncoder wherein an additional prior-encoder network learns an embedding of the data manifold which preserves topological and geometric properties of the data, thus improving the structure of latent space. The autoencoder and prior-encoder networks are iteratively trained using the Sliced Wasserstein distance. The effectiveness of the learned manifold encoding is explored by traversing latent space through interpolations along geodesics which generate samples that lie on the data manifold and hence are more realistic compared to Euclidean interpolation. To this end, we introduce a graph-based algorithm for exploring the data manifold and interpolating along network-geodesics in latent space by maximizing the density of samples along the path while minimizing total energy. We use the 3D-spiral data to show that the prior encodes the geometry underlying the data unlike conventional autoencoders, and to demonstrate the exploration of the embedded data manifold through the network algorithm. We apply our framework to benchmarked image datasets to demonstrate the advantages of learning data representations in outlier generation, latent structure, and geodesic interpolation.

cs.LG

Stroke recovery phenotyping through network trajectory approaches and graph neural networks

Stroke is a leading cause of neurological injury characterized by impairments in multiple neurological domains including cognition, language, sensory and motor functions. Clinical recovery in these domains is tracked using a wide range of measures that may be continuous, ordinal, interval or categorical in nature, which presents challenges for standard multivariate regression approaches. This has hindered stroke researchers' ability to achieve an integrated picture of the complex time-evolving interactions amongst symptoms. Here we use tools from network science and machine learning that are particularly well-suited to extracting underlying patterns in such data, and may assist in prediction of recovery patterns. To demonstrate the utility of this approach, we analyzed data from the NINDS tPA trial using the Trajectory Profile Clustering (TPC) method to identify distinct stroke recovery patterns for 11 different neurological domains at 5 discrete time points. Our analysis identified 3 distinct stroke trajectory profiles that align with clinically relevant stroke syndromes, characterized both by distinct clusters of symptoms, as well as differing degrees of symptom severity. We then validated our approach using graph neural networks to determine how well our model performed predictively for stratifying patients into these trajectory profiles at early vs. later time points post-stroke. We demonstrate that trajectory profile clustering is an effective method for identifying clinically relevant recovery subtypes in multidimensional longitudinal datasets, and for early prediction of symptom progression subtypes in individual patients. This paper is the first work introducing network trajectory approaches for stroke recovery phenotyping, and is aimed at enhancing the translation of such novel computational approaches for practical clinical application.

q-bio.QM

Success at high peaks: a multiscale approach combining individual and expedition-wide factors

This work presents a network-based data-driven study of the combination of factors that contribute to success in mountaineering. It simultaneously examines the effects of individual factors such as age, gender, experience etc., as well as expedition-wide factors such as number of camps, ratio of sherpas to paying climbers etc. Specifically, it combines the two perspectives into a multiscale network, i.e., a network of individual climber features within each expedition at the finer scale, and an expedition similarity network on the coarser scale. The latter is represented as a multiplex network where layers encode different factors. The analysis reveals that chances of failure to summit due to fatigue, altitude or logistical problems, drastically reduce when climbing with repeat partners, especially for experienced climbers. Additionally, node-centrality indicates that individual traits of youth and oxygen use are the strongest drivers of success. Further, the learning of network projections enables computation of correlations between intra-expedition networks and corresponding expedition success rates. Of expedition-wide factors, the expedition size and length layers are found to be strongly correlated with success rate. Lastly, community detection on the expedition-similarity network reveals distinct communities where a difference in success rates naturally emerges amongst the communities.

cs.SI

Multi-layer Trajectory Clustering: A Network Algorithm for Disease Subtyping

Many diseases display heterogeneity in clinical features and their progression, indicative of the existence of disease subtypes. Extracting patterns of disease variable progression for subtypes has tremendous application in medicine, for example, in early prognosis and personalized medical therapy. This work present a novel, data-driven, network-based Trajectory Clustering (TC) algorithm for identifying Parkinson's subtypes based on disease trajectory. Modeling patient-variable interactions as a bipartite network, TC first extracts communities of co-expressing disease variables at different stages of progression. Then, it identifies Parkinson's subtypes by clustering similar patient trajectories that are characterized by severity of disease variables through a multi-layer network. Determination of trajectory similarity accounts for direct overlaps between trajectories as well as second-order similarities, i.e., common overlap with a third set of trajectories. This work clusters trajectories across two types of layers: (a) temporal, and (b) ranges of independent outcome variable (representative of disease severity), both of which yield four distinct subtypes. The former subtypes exhibit differences in progression of disease domains (Cognitive, Mental Health etc.), whereas the latter subtypes exhibit different degrees of progression, i.e., some remain mild, whereas others show significant deterioration after 5 years. The TC approach is validated through statistical analyses and consistency of the identified subtypes with medical literature. This generalizable and robust method can easily be extended to other progressive multi-variate disease datasets, and can effectively assist in targeted subtype-specific treatment in the field of personalized medicine.

q-bio.QM