SearcharxivSearch

arXiv subjects

Siddharth Patwardhan

Publications and source records attributed to Siddharth Patwardhan.

17 recordsLinked to original sources

Crowding controls the scaling of bus frequency with demand

Cities must allocate limited resources to maintain mobility, with uncertainties about the resulting state of the system. Analyzing roughly 3,000 bus routes with more than 4 billion yearly riders across 19 metropolitan areas worldwide, we uncover a robust scaling law of the form $f \sim (d/t)^α$ with exponent $α\in [1/2,\,2/3]$, linking the service frequency $f$ to passenger demand $d$ and route duration $t$. We show that this scaling emerges from a simple optimization principle: cities implicitly minimize total passenger waiting time under a fixed operational budget when both schedule frequency and crowding are taken into account. This mechanism produces two universal regimes: a frequency-dominated regime with $α= 1/2$ when crowding is negligible, and a capacity-dominated regime with $α= 2/3$ when most routes are overloaded. Intermediate exponents arise when only part of the network operates near capacity. Furthermore, we find that the benefits of additional investment are highly uneven across systems. For instance, our model suggests that a $20\%$ budget increase yields nearly a 5-minute reduction in daily waiting time per passenger in Boston, compared to only about 1 minute in Paris. These findings place urban transit within a broader class of constrained capacity-allocation problems, while highlighting a distinct regime in which prescribed route demands shape the allocation of limited service resources. The resulting scaling laws show how simple optimization principles can generate systematic exponents in complex transport systems, beyond the dissipation-based frameworks usually considered in physical and biological flow networks.

physics.soc-ph

Multilayer network science: theory, methods, and applications

Multilayer network science has emerged as a central framework for analysing interconnected and interdependent complex systems. Its relevance has grown substantially with the increasing availability of rich, heterogeneous data, which makes it possible to uncover and exploit the inherently multilayered organisation of many real-world networks. In this review, we summarise recent developments in the field. On the theoretical and methodological front, we outline core concepts and survey advances in community detection, dynamical processes, temporal networks, higher-order interactions, and machine-learning-based approaches. On the application side, we discuss progress across diverse domains, including interdependent infrastructures, spreading dynamics, computational social science, economic and financial systems, ecological and climate networks, science-of-science studies, network medicine, and network neuroscience. We conclude with a forward-looking perspective, emphasizing the need for standardised datasets and software, deeper integration of temporal and higher-order structures, and a transition toward genuinely predictive models of complex systems.

physics.soc-ph

Epidemic spreading in group-structured populations

Individuals involved in common group activities/settings -- e.g., college students that are enrolled in the same class and/or live in the same dorm -- are exposed to recurrent contacts of physical proximity. These contacts are known to mediate the spread of an infectious disease, however, it is not obvious how the properties of the spreading process are determined by the structure of and the interrelation among the group settings that are at the root of those recurrent interactions. Here, we show that reshaping the organization of groups within a population can be used as an effective strategy to decrease the severity of an epidemic. Specifically, we show that when group structures are sufficiently correlated -- e.g., the likelihood for two students living in the same dorm to attend the same class is sufficiently high -- outbreaks are longer but milder than for uncorrelated group structures. Also, we show that the effectiveness of interventions for disease containment increases as the correlation among group structures increases. We demonstrate the practical relevance of our findings by taking advantage of data about housing and attendance of students at the Indiana University campus in Bloomington. By appropriately optimizing the assignment of students to dorms based on their enrollment, we are able to observe a two- to five-fold reduction in the severity of simulated epidemic processes.

physics.soc-ph

Reconstruction of multiplex networks via graph embeddings

Multiplex networks are collections of networks with identical nodes but distinct layers of edges. They are genuine representations for a large variety of real systems whose elements interact in multiple fashions or flavors. However, multiplex networks are not always simple to observe in the real world; often, only partial information on the layer structure of the networks is available, whereas the remaining information is in the form of aggregated, single-layer networks. Recent works have proposed solutions to the problem of reconstructing the hidden multiplexity of single-layer networks using tools proper of network science. Here, we develop a machine learning framework that takes advantage of graph embeddings, i.e., representations of networks in geometric space. We validate the framework in systematic experiments aimed at the reconstruction of synthetic and real-world multiplex networks, providing evidence that our proposed framework not only accomplishes its intended task, but often outperforms existing reconstruction techniques.

physics.soc-ph

Symmetry breaking in optimal transport networks

Despite its importance for practical applications, not much is known about the optimal shape of a network that connects in an efficient way a set of points. This problem can be formulated in terms of a multiplex network with a fast layer embedded in a slow one. To connect a pair of points, one can then use either the fast or slow layer, or both, with a switching cost when going from one layer to the other. We consider here distributions of points in spaces of arbitrary dimension d and search for the fast-layer network of given size that minimizes the average time to reach a central node. We discuss the d = 1 case analytically and the d > 1 case numerically, and show the existence of transitions when we vary the network size, the switching cost and/or the relative speed of the two layers. Surprisingly, there is a transition characterized by a symmetry breaking indicating that it is sometimes better to avoid serving a whole area in order to save on switching costs, at the expense of using more the slow layer. Our findings underscore the importance of considering switching costs while studying optimal network structures, as small variations of the cost can lead to strikingly dissimilar results. Finally, we discuss real-world subways and their efficiency for the cities of Atlanta, Boston, and Toronto. We find that real subways are farther away from the optimal shapes as traffic congestion increases.

physics.soc-ph

EELBERT: Tiny Models through Dynamic Embeddings

We introduce EELBERT, an approach for compression of transformer-based models (e.g., BERT), with minimal impact on the accuracy of downstream tasks. This is achieved by replacing the input embedding layer of the model with dynamic, i.e. on-the-fly, embedding computations. Since the input embedding layer accounts for a significant fraction of the model size, especially for the smaller BERT variants, replacing this layer with an embedding computation function helps us reduce the model size significantly. Empirical evaluation on the GLUE benchmark shows that our BERT variants (EELBERT) suffer minimal regression compared to the traditional BERT models. Through this approach, we are able to develop our smallest model UNO-EELBERT, which achieves a GLUE score within 4% of fully trained BERT-tiny, while being 15x smaller (1.2 MB) in size.

cs.CL

Machine Learning as an Accurate Predictor for Percolation Threshold of Diverse Networks

The percolation threshold is an important measure to determine the inherent rigidity of large networks. Predictors of the percolation threshold for large networks are computationally intense to run, hence it is a necessity to develop predictors of the percolation threshold of networks, that do not rely on numerical simulations. We demonstrate the efficacy of five machine learning-based regression techniques for the accurate prediction of the percolation threshold. The dataset generated to train the machine learning models contains a total of 777 real and synthetic networks. It consists of 5 statistical and structural properties of networks as features and the numerically computed percolation threshold as the output attribute. We establish that the machine learning models outperform three existing empirical estimators of bond percolation threshold, and extend this experiment to predict site and explosive percolation. Further, we compared the performance of our models in predicting the percolation threshold using RMSE values. The gradient boosting regressor, multilayer perceptron and random forests regression models achieve the least RMSE values among considered models.

physics.soc-ph

Multiplex reconstruction with partial information

A multiplex is a collection of network layers, each representing a specific type of edges. This appears to be a genuine representation for many real-world systems. However, due to a variety of potential factors, such as limited budget and equipment, or physical impossibility, multiplex data can be difficult to observe directly. Often, only partial information on the layer structure of the system is available, whereas the remaining information is in the form of a single-layer network. In this work, we face the problem of reconstructing the hidden multiplex structure of an aggregated network from partial information. We propose an algorithm that leverages the layer-wise community structure that can be learned from partial observations to reconstruct the ground-truth topology of the unobserved part of the multiplex. The algorithm is characterized by a computational time that grows linearly with the network size. We perform a systematic study of reconstruction problems for both synthetic and real-world multiplex networks. We show that the ability of the proposed method to solve the reconstruction problem is affected by the heterogeneity of the individual layers and the similarity among the layers. On real-world networks, we observe that the accuracy of the reconstruction saturates quickly as the amount of available information increases. In genetic interaction and scientific collaboration multiplexes for example, we find that 10% of ground-truth information yields 70% accuracy, while 30% information allows for more than 90% accuracy.

physics.soc-ph

Languages You Know Influence Those You Learn: Impact of Language Characteristics on Multi-Lingual Text-to-Text Transfer

Multi-lingual language models (LM), such as mBERT, XLM-R, mT5, mBART, have been remarkably successful in enabling natural language tasks in low-resource languages through cross-lingual transfer from high-resource ones. In this work, we try to better understand how such models, specifically mT5, transfer *any* linguistic and semantic knowledge across languages, even though no explicit cross-lingual signals are provided during pre-training. Rather, only unannotated texts from each language are presented to the model separately and independently of one another, and the model appears to implicitly learn cross-lingual connections. This raises several questions that motivate our study, such as: Are the cross-lingual connections between every language pair equally strong? What properties of source and target language impact the strength of cross-lingual transfer? Can we quantify the impact of those properties on the cross-lingual transfer? In our investigation, we analyze a pre-trained mT5 to discover the attributes of cross-lingual connections learned by the model. Through a statistical interpretation framework over 90 language pairs across three tasks, we show that transfer performance can be modeled by a few linguistic and data-derived features. These observations enable us to interpret cross-lingual understanding of the mT5 model. Through these observations, one can favorably choose the best source language for a task, and can anticipate its training data demands. A key finding of this work is that similarity of syntax, morphology and phonology are good predictors of cross-lingual transfer, significantly more than just the lexical similarity of languages. For a given language, we are able to predict zero-shot performance, that increases on a logarithmic scale with the number of few-shot target language data points.

cs.CL

Influence Maximization: Divide and Conquer

The problem of influence maximization, i.e., finding the set of nodes having maximal influence on a network, is of great importance for several applications. In the past two decades, many heuristic metrics to spot influencers have been proposed. Here, we introduce a framework to boost the performance of any such metric. The framework consists in dividing the network into sectors of influence, and then selecting the most influential nodes within these sectors. We explore three different methodologies to find sectors in a network: graph partitioning, graph hyperbolic embedding, and community structure. The framework is validated with a systematic analysis of real and synthetic networks. We show that the gain in performance generated by dividing a network into sectors before selecting the influential spreaders increases as the modularity and heterogeneity of the network increase. Also, we show that the division of the network into sectors can be efficiently performed in a time that scales linearly with the network size, thus making the framework applicable to large-scale influence maximization problems.

physics.soc-ph

Can Open Domain Question Answering Systems Answer Visual Knowledge Questions?

The task of Outside Knowledge Visual Question Answering (OKVQA) requires an automatic system to answer natural language questions about pictures and images using external knowledge. We observe that many visual questions, which contain deictic referential phrases referring to entities in the image, can be rewritten as "non-grounded" questions and can be answered by existing text-based question answering systems. This allows for the reuse of existing text-based Open Domain Question Answering (QA) Systems for visual question answering. In this work, we propose a potentially data-efficient approach that reuses existing systems for (a) image analysis, (b) question rewriting, and (c) text-based question answering to answer such visual questions. Given an image and a question pertaining to that image (a visual question), we first extract the entities present in the image using pre-trained object and scene classifiers. Using these detected entities, the visual questions can be rewritten so as to be answerable by open domain QA systems. We explore two rewriting strategies: (1) an unsupervised method using BERT for masking and rewriting, and (2) a weakly supervised approach that combines adaptive rewriting and reinforcement learning techniques to use the implicit feedback from the QA system. We test our strategies on the publicly available OKVQA dataset and obtain a competitive performance with state-of-the-art models while using only 10% of the training data.

cs.AI

Model Stability with Continuous Data Updates

In this paper, we study the "stability" of machine learning (ML) models within the context of larger, complex NLP systems with continuous training data updates. For this study, we propose a methodology for the assessment of model stability (which we refer to as jitter under various experimental conditions. We find that model design choices, including network architecture and input representation, have a critical impact on stability through experiments on four text classification tasks and two sequence labeling tasks. In classification tasks, non-RNN-based models are observed to be more stable than RNN-based ones, while the encoder-decoder model is less stable in sequence labeling tasks. Moreover, input representations based on pre-trained fastText embeddings contribute to more stability than other choices. We also show that two learning strategies -- ensemble models and incremental training -- have a significant influence on stability. We recommend ML model designers account for trade-offs in accuracy and jitter when making modeling choices.

cs.CL

Quantifying Dismantlement in Disconnected Networks

We propose a novel measure to quantify dismantlement of a fragmented network. The existing measure of dismantlement used to study problems like optimal percolation is usually the size of the largest component of the network. We modify the measure of uniformity used to prove the Szemeredi's Regularity Lemma to obtain the proposed measure. The proposed measure incorporates the notion that the measure of dismantlement increases as the number of disconnected components increase and decreases as the variance of sizes of these components increases.

physics.soc-ph

Annotating Electronic Medical Records for Question Answering

Our research is in the relatively unexplored area of question answering technologies for patient-specific questions over their electronic health records. A large dataset of human expert curated question and answer pairs is an important pre-requisite for developing, training and evaluating any question answering system that is powered by machine learning. In this paper, we describe a process for creating such a dataset of questions and answers. Our methodology is replicable, can be conducted by medical students as annotators, and results in high inter-annotator agreement (0.71 Cohen's kappa). Over the course of 11 months, 11 medical students followed our annotation methodology, resulting in a question answering dataset of 5696 questions over 71 patient records, of which 1747 questions have corresponding answers generated by the medical students.

cs.CL

The Role of Context Types and Dimensionality in Learning Word Embeddings

We provide the first extensive evaluation of how using different types of context to learn skip-gram word embeddings affects performance on a wide range of intrinsic and extrinsic NLP tasks. Our results suggest that while intrinsic tasks tend to exhibit a clear preference to particular types of contexts and higher dimensionality, more careful tuning is required for finding the optimal settings for most of the extrinsic tasks that we considered. Furthermore, for these extrinsic tasks, we find that once the benefit from increasing the embedding dimensionality is mostly exhausted, simple concatenation of word embeddings, learned with different context types, can yield further performance gains. As an additional contribution, we propose a new variant of the skip-gram model that learns word embeddings from weighted contexts of substitute words.

cs.CL

Addressing Limited Data for Textual Entailment Across Domains

We seek to address the lack of labeled data (and high cost of annotation) for textual entailment in some domains. To that end, we first create (for experimental purposes) an entailment dataset for the clinical domain, and a highly competitive supervised entailment system, ENT, that is effective (out of the box) on two domains. We then explore self-training and active learning strategies to address the lack of labeled data. With self-training, we successfully exploit unlabeled data to improve over ENT by 15% F-score on the newswire domain, and 13% F-score on clinical data. On the other hand, our active learning experiments demonstrate that we can match (and even beat) ENT using only 6.6% of the training data in the clinical domain, and only 5.8% of the training data in the newswire domain.

cs.CL

Efficient Controlled Quantum Secure Direct Communication Protocols

We study controlled quantum secure direct communication (CQSDC), a cryptographic scheme where a sender can send a secret bit-string to an intended recipient, without any secure classical channel, who can obtain the complete bit-string only with the permission of a controller. We report an efficient protocol to realize CQSDC using Cluster state and then go on to construct a (2-3)-CQSDC using Brown state, where a coalition of any two of the three controllers is required to retrieve the complete message. We argue both protocols to be unconditionally secure and analyze the efficiency of the protocols to show it to outperform the existing schemes while maintaining the same security specifications.

quant-ph