SearcharxivSearch

arXiv subjects

Ritwick Mishra

Publications and source records attributed to Ritwick Mishra.

7 recordsLinked to original sources

Cascading Impacts of the USA--China Trade War on Global Oilseed Supply Chain

Global supply chains are highly interconnected, making them vulnerable to cascading disruptions induced by trade policy shocks. Understanding how such disruptions propagate through production networks, and how mitigation mechanisms such as trade reallocation and production adjustment can alleviate their impacts, remains a central challenge. In this work, we develop a linear programming formulation of an Input-Output (IO) system that captures cascading supply-chain disruptions together with trade reallocation and production expansion. Our formulation yields a system-level equilibrium characterization that enables the joint analysis of disruption propagation and mitigation within a unified framework. We propose an efficient algorithm for computing approximate equilibrium solutions by minimizing total unmet demand in large IO systems. We apply our approach to tariff-induced disruptions in the global oilseeds supply chain arising from the U.S.-China trade war. Our results show that a localized 70% disruption to flows from the U.S. oilseeds sector to China leads to a 3.27% loss in global output, with China experiencing a disproportionate loss of 14.02%. As a counterfactual mitigation strategy, allowing a 20% reallocation from Brazil's oilseed sector to China significantly reduces global output losses to 1.36%, although pressure remains high on final-demand flows. We further investigate production expansion as an additional mitigation mechanism and show that it introduces tradeoffs between reducing global final-demand losses and protecting Brazil's domestic flows. Domestic reallocation disproportionately shifts losses toward smaller economies, while globally sourced expansion redistributes losses more broadly across the network.

econ.GN

Scalable K-clique Estimation with Differential Privacy

Counts of $k$-cliques are commonly used metrics in subgraph mining. Since graphs often have sensitive data, there also has been a lot of work on $k$-clique counts with differential privacy. However, these metrics have very high global sensitivity, and so more sophisticated techniques have been developed for counting $k$-cliques with privacy. Smooth sensitivity and ladder functions were developed for reducing the noise magnitude for private estimates of these metrics. However, these are computationally very inefficient to estimate. No polynomial time algorithms are known for smooth sensitivity of $k$-cliques for $k>3$, while the time complexity of ladder functions is lower bounded by the time for exact counts, which does not scale very well. In this paper, we develop a new highly scalable algorithm for estimating $k$-clique counts with differential privacy. Our algorithm adapts the ladder function to serve as a smooth upper bound on its local sensitivity, and utilizes the approximation sensitivity framework to calibrate noise with magnitude proportional to an approximation of the bound. This gives us a significant improvement in the running time. Experiments show that our method is several orders of magnitude faster than the ladder function based estimates of $k$-clique counts, while the accuracy is similar. Our algorithm is the first to scale to graphs with millions of edges, and for larger $k$, for which the ladder function algorithm doesn't complete.

cs.DS

Estimating the cascading global impacts of gas disruptions in Qatar

This study examines the global impacts of a localized disruption in Qatar's gas sector using a multi-regional input-output framework and scenario-based analysis. While the direct impacts of this disruption on importing countries are clear, indirect and cascading impacts are not well understood. We use a Multiregional input-output (MRIO) model to assess the impact of this disruption and to determine whether trade reallocations and increased production can mitigate its effects. Our analysis shows that this disruption leads to significant gas supply losses in Asia and Europe, with the largest aggregate impacts observed in India, China, and South Korea. Allowing for trade reallocation partially mitigates these losses. Further expansion of production capacity among major gas-producing countries improves supply conditions and leads to broader output gains; however, these benefits remain concentrated in a few large economies. Even significant increases in production among top producers offer limited relief to economies such as India and Pakistan. Overall, the results highlight the uneven distribution of both vulnerabilities and recovery potential within global supply chains. While adaptive mechanisms such as trade reallocation and production expansion can alleviate the effects of supply shocks, their effectiveness is limited and heterogeneous. The findings underscore the importance of network structure in shaping shock propagation and resilience, offering insights for managing systemic risks in an interconnected global economy.

physics.soc-ph

Reconstructing Network Outbreaks under Group Surveillance

A key public health problem during an outbreak is to reconstruct the disease cascade from a partial set of confirmed infections. This has been studied extensively under the Maximum Likelihood Estimation (MLE) formulation, which reduces the problem to finding some type of Steiner subgraph on a network. Group surveillance like wastewater or aerosol monitoring is a form of mass/pooled testing where samples from multiple individuals are pooled together and tested once for all. While a single negative test clears multiple individuals, a positive test does not reveal the infected individuals in the test pool. We introduce the POOLCASCADEMLE problem in the setting of a network propagation process, where the goal is to find a MLE cascade subgraph which is consistent with the pooled test outcomes. Previous work on reconstruction assumes that the test results are of individuals, i.e., pools of size one, and requires a consistent cascade to connect the positive testing nodes. In POOLCASCADEMLE, a consistent cascade must choose at least one node in each positive pool, adding another combinatorial layer. We show that, under the Independent Cascade (IC) model, POOLCASCADEMLE is NP-hard, and present an approximation algorithm based on a reduction to the Group Steiner Tree problem. We also consider a one-hop version of this problem, in which the disease can spread for one time step after being seeded. We show that even this restricted version is NP-hard, and develop a method using linear programming relaxation and rounding. We evaluate the performance of our methods on real and synthetic contact networks, in terms of missing infection recovery and prevalence estimation. We find that our approach outperforms meaningful baselines which correspond to pools of size one and use state-of-the-art methods.

cs.SI

Information Theoretic Optimal Surveillance for Epidemic Prevalence in Networks

Estimating the true prevalence of an epidemic outbreak is a key public health problem. This is challenging because surveillance is usually resource intensive and biased. In the network setting, prior work on cost sensitive disease surveillance has focused on choosing a subset of individuals (or nodes) to minimize objectives such as probability of outbreak detection. Such methods do not give insights into the outbreak size distribution which, despite being complex and multi-modal, is very useful in public health planning. We introduce TESTPREV, a problem of choosing a subset of nodes which maximizes the mutual information with disease prevalence, which directly provides information about the outbreak size distribution. We show that, under the independent cascade (IC) model, solutions computed by all prior disease surveillance approaches are highly sub-optimal for TESTPREV in general. We also show that TESTPREV is hard to even approximate. While this mutual information objective is computationally challenging for general networks, we show that it can be computed efficiently for various network classes. We present a greedy strategy, called GREEDYMI, that uses estimates of mutual information from cascade simulations and thus can be applied on any network and disease model. We find that GREEDYMI does better than natural baselines in terms of maximizing the mutual information as well as reducing the expected variance in outbreak size, under the IC model.

physics.soc-ph

A Look into Causal Effects under Entangled Treatment in Graphs: Investigating the Impact of Contact on MRSA Infection

Methicillin-resistant Staphylococcus aureus (MRSA) is a type of bacteria resistant to certain antibiotics, making it difficult to prevent MRSA infections. Among decades of efforts to conquer infectious diseases caused by MRSA, many studies have been proposed to estimate the causal effects of close contact (treatment) on MRSA infection (outcome) from observational data. In this problem, the treatment assignment mechanism plays a key role as it determines the patterns of missing counterfactuals -- the fundamental challenge of causal effect estimation. Most existing observational studies for causal effect learning assume that the treatment is assigned individually for each unit. However, on many occasions, the treatments are pairwisely assigned for units that are connected in graphs, i.e., the treatments of different units are entangled. Neglecting the entangled treatments can impede the causal effect estimation. In this paper, we study the problem of causal effect estimation with treatment entangled in a graph. Despite a few explorations for entangled treatments, this problem still remains challenging due to the following challenges: (1) the entanglement brings difficulties in modeling and leveraging the unknown treatment assignment mechanism; (2) there may exist hidden confounders which lead to confounding biases in causal effect estimation; (3) the observational data is often time-varying. To tackle these challenges, we propose a novel method NEAT, which explicitly leverages the graph structure to model the treatment assignment mechanism, and mitigates confounding biases based on the treatment assignment modeling. We also extend our method into a dynamic setting to handle time-varying observational data. Experiments on both synthetic datasets and a real-world MRSA dataset validate the effectiveness of the proposed method, and provide insights for future applications.

cs.LG

Data Selection for Fine-tuning Large Language Models Using Transferred Shapley Values

Although Shapley values have been shown to be highly effective for identifying harmful training instances, dataset size and model complexity constraints limit the ability to apply Shapley-based data valuation to fine-tuning large pre-trained language models. To address this, we propose TS-DShapley, an algorithm that reduces computational cost of Shapley-based data valuation through: 1) an efficient sampling-based method that aggregates Shapley values computed from subsets for valuation of the entire training set, and 2) a value transfer method that leverages value information extracted from a simple classifier trained using representations from the target language model. Our experiments applying TS-DShapley to select data for fine-tuning BERT-based language models on benchmark natural language understanding (NLU) datasets show that TS-DShapley outperforms existing data selection methods. Further, TS-DShapley can filter fine-tuning data to increase language model performance compared to training with the full fine-tuning dataset.

cs.CL