Searcharxiv⌕ Search

arXiv subjects

M. Amin Rahimian

Publications and source records attributed to M. Amin Rahimian.

At least 19 recordsLinked to original sources

Privacy-Aware Sequential Learning

Sequential learning often relies on individuals reporting private information, with early reports shaping later beliefs and collective decisions. When such reports are sensitive, privacy protection creates a fundamental trade-off in the design of sequential feedback systems. We study this trade-off as a platform design problem: before reporting begins, the platform commits to a local privacy-preserving reporting protocol, seeking to improve final-decision accuracy while limiting the time required to accumulate sufficient evidence. Surprisingly, privacy protection can accelerate learning. With continuous Gaussian signals, smooth randomized response under metric differential privacy preserves asymptotic learning and yields a public log-likelihood ratio growing at rate $Θ_{\varepsilon}(\log n)$, faster than the nonprivate $Θ(\sqrt{\log n})$ benchmark, substantially reducing high-confidence stopping times. With heterogeneous privacy parameters, learning can be faster still; when privacy budgets are uniformly distributed on $[0,1]$, public belief grows at rate $Θ(n^{1/4})$. With binary signals, randomized response generates a nonmonotone relationship between privacy and decision accuracy because privacy affects both report informativeness and cascade thresholds. Stopping time can also be nonmonotone because stronger privacy reduces report informativeness while encouraging participation. Overall, privacy is not merely a constraint on information release, but a platform design lever shaping participation, stopping, and collective learning.

econ.TH↗

Socio-Spatial Patterns of Suicide Mortality in the United States

Suicide causes more than 49,000 deaths annually in the United States, 55% involving firearms. Suicide mortality varies substantially across US counties, but the role of large-scale social networks remains underexplored. We combine county-level suicide mortality data for 2010-2022 with the Facebook Social Connectedness Index (SCI) and measures of exposure to Extreme Risk Protection Orders (ERPOs). In population-weighted two-way fixed effects models adjusting for sociodemographic and economic characteristics, COVID-19, drug-overdose and alcohol-related mortality, poor mental health days, state firearm policies, and geographical proximity, a one-standard-deviation increase in the SCI-weighted suicide mortality rate of socially connected counties was associated with 2.47 additional deaths per 100,000 residents in the focal county (95% CI [0.97, 4.00]). A one-standard-deviation increase in ERPO social exposure was associated with 0.266 fewer deaths per 100,000 (coefficient -0.266; 95% CI [-0.403, -0.128]) after adjustment for geographical proximity and state-by-year fixed effects. The social-proximity association persisted under alternative neighborhood definitions and spatial-error models, and subgroup analyses showed heterogeneity by age. These findings show that both suicide mortality and the lower mortality associated with exposure to firearm-restriction policies are patterned along inter-county social ties beyond what is explained by geographical proximity alone.

stat.AP↗

County-Level Heterogeneity in Opioid Harm Reduction and Treatment Effects: A Simulation Modeling Analysis

Opioid overdose deaths remain a severe public health crisis in the US, with heterogeneous burden across counties that differ in epidemic trajectory, baseline resources, and local context. While harm reduction through naloxone distribution and buprenorphine treatment are both evidence-based strategies, limited information on county-level effects hinders the ability of policymakers to prioritize resources across counties. We developed a simulation model of opioid use disorder (OUD), calibrated separately to six Pennsylvania counties spanning large urban (Allegheny, Philadelphia), intermediate-sized (Erie, Dauphin), and rural (Clearfield, Columbia) settings. We projected county-specific overdose mortality trajectories under three levels of increase in buprenorphine dispensing and naloxone distribution (10%, 20%, and 30% above each county's baseline), over a five-year horizon from 2025 to 2029. A 30% increase in naloxone distribution above observed county baseline levels was projected to reduce 2029 overdose deaths by approximately 70% (95% uncertainty interval, UI: 55-81%) in Allegheny County, 11% (95% UI:4-17%) in Erie County, and 28% (95% UI:2-64%) in Clearfield County. Projected reductions in overdose deaths from increasing buprenorphine were consistently smaller (10%-23%), except that in Erie buprenorphine produced larger projected reduction by 20% vs 11% for naloxone. Heterogeneity in naloxone responsiveness was strongly associated with each county's historical naloxone dispensing variability. The same proportional increase in naloxone distribution yields substantially different projected mortality reductions across counties depending on each county's baseline distribution history, a pattern invisible from mortality statistics alone. County-level context is important for informing harm reduction and treatment prioritization at the county level.

stat.AP↗

Differentially Private Distributed Inference for Multicenter Clinical Studies

Extracting reliable conclusions from data distributed across institutions is a core problem in healthcare: pooling patient records would improve inference, but privacy regulations and the lack of a trusted central authority frequently delay multicenter studies. We develop a framework for differentially private distributed inference in which institutions repeatedly exchange log belief-ratio statistics subject to differential privacy (DP). With arithmetic and geometric averaging of beliefs, we control the false-negative and false-positive rates as functions of the privacy budget, communication rounds, and statistical separation between hypotheses, exposing a three-way trade-off among accuracy, communication, and privacy. We derive finite-sample bounds on the Type I and Type II error probabilities that allow distributed hypothesis testing at a target significance level, and show that the Laplace mechanism minimizes convergence time subject to DP. For distributed online learning from data streams (e.g., epidemiological surveillance or rolling recruitment), privacy noise vanishes asymptotically and online learning admits similar finite-sample guarantees. On simulated multicenter survival analyses using the AIDS Clinical Trials Group and an advanced-cancer cohort, and a simulated genetic association study over New York City hospitals, our method approaches the non-private baseline at a small privacy budget between 1 and 10, runs 10x to 1000x faster than homomorphic-encryption methods, and incurs up to 100x lower error than first-order private optimization methods. Finally, the level of aggregation is a primary design choice: federating at the organizational rather than the hospital level strengthens privacy, raises statistical power, and lowers communication and administrative burden, so data should be pooled within organizations before setting up federated analytics.

cs.LG↗

Computationally Efficient Estimation of Localized Treatment Effects for Multi-Level, Multi-Component Interventions to Address the Opioid Crisis

The opioid epidemic remains a major public health challenge in the United States, requiring a multi-pronged intervention approach to mitigate harms to communities. Given the heterogeneity of the epidemic, it is crucial for policymakers to understand localized treatment effects of different intervention components and utilize limited resources efficiently. While locally calibrated simulation models can project epidemic outcomes for any given intervention policy, collecting simulation results for all intervention combinations to estimate localized treatment effects for each community is impractical because the number of combinations grows exponentially with the number of interventions and the levels at which they are applied. To tackle this, we develop a two-stage metamodel framework with a two-step sequential design for efficient sampling. The metamodel consists of a response function linking health outcomes to each intervention component's treatment effect, and a Gaussian process regression (GPR) to learn spatial and socio-economic structures of the treatment effects based on locally-contextualized covariates. With two-step sequential sampling, we leverage spatial correlations and posterior uncertainty to sequentially sample the most informative counties and treatment conditions. We apply this framework to estimate the treatment effects of buprenorphine dispensing and naloxone distribution on overdose mortality rates using a calibrated agent-based opioid epidemic model in Pennsylvania counties. Our approach achieves less than 5% average relative error using fewer than 2% of the runs required for an exhaustive simulation. Our two-stage framework provides a computationally efficient approach to support policymakers, enabling an efficient evaluation of alternative resource-allocation strategies to mitigate the opioid epidemic in local communities.

stat.AP↗

Mechanism Design for Privacy-Preserving Information Sharing in Oligopoly Competition

Information sharing among competing suppliers can improve decisions under demand uncertainty, but it may also intensify strategic interaction by aligning firms' beliefs. We study a Cournot oligopoly in which a platform designs an information-sharing mechanism using participation-contingent access, external platform information, and privacy-preserving noise. The central privacy-design challenge is that noise has two opposing effects: it limits how much a firm's report improves rivals' information, but it also reduces the value of the posterior signal released by the platform. In symmetric duopoly, privacy protection alone cannot implement sharing without an external platform signal. More generally, privacy can induce firms that would otherwise not share to participate only when combined with external platform information, which preserves an informational benefit independent of competitors' reports. The $n$-firm case adds a distinct force: under reciprocal access, non-participants lose access to the pooled signal generated by others, so a baseline sharing region may exist even without privacy protection or platform signals. We characterize this sharing-feasible region and show how external information and privacy noise expand it beyond the reciprocal-access baseline. We further provide conditions under which full sharing, rather than partial participation, is the unique participation equilibrium. Finally, we show that privacy noise is valuable as an implementation tool but costly for welfare, so the platform chooses the least distortionary privacy level that implements full sharing, and privacy-induced sharing improves total surplus only when participation gains outweigh informational losses.

econ.TH↗

Seeding with Differentially Private Network Information

In public health interventions such as distributing preexposure prophylaxis (PrEP) for HIV prevention, decision makers often use seeding algorithms to identify key individuals who can amplify intervention impact. However, building a complete sexual activity network is typically infeasible due to privacy concerns. Instead, contact tracing can provide influence samples, observed sequences of sexual contacts, without full network reconstruction. This raises two challenges: protecting individual privacy in these samples and adapting seeding algorithms to incomplete data. We study differential privacy guarantees for influence maximization when the input consists of randomly collected cascades. Building on recent advances in costly seeding, we propose privacy-preserving algorithms that introduce randomization in data or outputs and bound the privacy loss of each node. Theoretical analysis and simulations on synthetic and real-world sexual contact data show that performance degrades gracefully as privacy budgets tighten, with central privacy regimes achieving better trade-offs than local ones.

cs.SI↗

Conformal-DP: A Density-Aware Mechanism for Differential Privacy over Riemannian Manifolds via Conformal Transformation

Differential Privacy (DP) is being increasingly adopted for non-Euclidean data that lie on complex, high-dimensional manifolds. Existing DP mechanisms for manifold data consider geometric properties when calibrating privacy perturbations, but they largely fail to capture variations in data density within datasets, leading to biased perturbations and suboptimal privacy-utility trade-offs due to heterogeneous data distributions. In this paper, we propose a novel density-aware differential privacy mechanism on Riemannian manifolds, referred to as Conformal-DP, that leverages conformal transformations to calibrate perturbations based on local densities and to induce a density-balanced geometry. We prove that our mechanism satisfies $ε$-differential privacy on any complete Riemannian manifold under mild regularity assumptions. In addition, we derive a closed-form expected geodesic error bound that depends only on the underlying data density ratio and is independent of global curvature. Our empirical results on synthetic and real-world datasets demonstrate that the proposed Conformal-DP mechanism substantially improves the privacy-utility trade-off in heterogeneous data distribution settings, with worst-case performance comparable to state-of-the-art manifold DP mechanisms that assume uniformly distributed data.

cs.CR↗

Structural Measures of Resilience for Supply Chains

Modern production systems are increasingly defined by dense networks of multi-tier sourcing dependencies, where localized upstream disruptions can cascade into system-wide collapses. While supply chain resilience has garnered significant managerial attention, we still lack theoretically-grounded, reliable, analytical metrics that can distinguish inherently resilient architectures from fragile ones. This paper addresses this gap by developing a structural resilience framework and a novel metric, defined as the maximum supplier failure rate that a network can sustain while maintaining an aggregate production level. Using node percolation theory and branching processes, we identify four critical structural determinants of resilience: the number of raw materials, the number of finished goods, sourcing requirements, and sourcing influence. Our analysis reveals two distinct regimes: "top hat" architectures, which are characterized by excessive raw materials and high centralization, making them inherently fragile; and "rolling pin" structures, which maintain controlled input/output widths and sparsity, allowing them to absorb non-trivial shocks. To operationalize these insights, we formulate resilience computation as a scalable linear program that approximates cascading failure sizes in large-scale networks with cycles, heterogeneous suppliers, and structural decoupling. Furthermore, we extend our framework to account for exogenous failure correlations, such as those arising from geographic or geopolitical factors that can undermine traditional supplier and input diversification strategies. We validate our theoretical results using multi-echelon supply chain data. These tools can inform network design, supplier diversification, and inventory planning to proactively reduce systemic risk.

cs.SI↗

Differential Privacy for Network Connectedness Indices

Researchers increasingly use data on social and economic networks to study a range of social science questions, but releasing statistics derived from networks can raise significant privacy concerns. We show how to release network connectedness indices that quantify assortative mixing across node attributes under edge-adjacent differential privacy. Standard privacy techniques perform poorly in this setting both because connectedness indices have high global sensitivity and because a single node's attribute can potentially be an input to connectedness in thousands of cells, leading to poor composition. Our method, which is straightforward to apply, first adds noise to node attributes, then analytically debiases downstream statistics, and finally applies a second layer of noise to protect the presence or absence of individual edges. We prove consistency and asymptotic normality of our estimators for both discrete and continuous labels and show our method works well in simulations and on real networks with as few as 200 nodes collected by social scientists.

stat.AP↗

Privacy at Scale in Networked Healthcare

Digitized, networked healthcare promises earlier detection, precision therapeutics, and continuous care; yet, it also expands the surface for privacy loss and compliance risk. We argue for a shift from siloed, application-specific protections to privacy-by-design at scale, centered on decision-theoretic differential privacy (DP) across the full healthcare data lifecycle; network-aware privacy accounting for interdependence in people, sensors, and organizations; and compliance-as-code tooling that lets health systems share evidence while demonstrating regulatory due care. We synthesize the privacy-enhancing technology (PET) landscape in health (federated analytics, DP, cryptographic computation), identify practice gaps, and outline a deployable agenda involving privacy-budget ledgers, a control plane to coordinate PET components across sites, shared testbeds, and PET literacy, to make lawful, trustworthy sharing the default. We illustrate with use cases (multi-site trials, genomics, disease surveillance, mHealth) and highlight distributed inference as a workhorse for multi-institution learning under explicit privacy budgets.

cs.CR↗

Social learning moderates the tradeoffs between efficiency, stability, and equity in group foraging

Collective foragers, from animals to robotic swarms, must balance exploration and exploitation to locate sparse resources efficiently. While social learning is known to facilitate this balance, how the range of information sharing shapes group-level outcomes remains unclear. Here, we develop a minimal collective foraging model in which individuals combine independent exploration, local exploitation, and socially guided movement. We show that foraging efficiency is maximized at an intermediate social learning range, where groups exploit discovered resources without suppressing independent discovery. This optimal regime also minimizes temporal burstiness in resource intake, reducing starvation risk. Increasing social learning range further improves equity among individuals but degrades efficiency through redundant exploitation. Introducing risky (negative) targets shifts the optimal range upward; in contrast, when penalties are ignored, randomly distributed negative cues can further enhance efficiency by constraining unproductive exploration. Together, these results reveal how local information rules regulate a fundamental trade-off between efficiency, stability, and equity, providing design principles for biological foraging systems and engineered collectives.

physics.soc-ph↗

Structural Dynamics of Harmful Content Dissemination on WhatsApp

WhatsApp, a platform with more than two billion global users, plays a crucial role in digital communication, but also serves as a vector for harmful content such as misinformation, hate speech, and political propaganda. This study examines the dynamics of harmful message dissemination in WhatsApp groups, with a focus on their structural characteristics. Using a comprehensive data set of more than 5.1 million messages, including text, images, and videos, collected from approximately 6,000 groups in India, we reconstruct message propagation cascades to analyze dissemination patterns. Our findings reveal that harmful messages consistently achieve greater depth and breadth of dissemination compared to messages without harmful annotations, with videos and images emerging as the primary modes of dissemination. These results suggest a distinctive pattern of dissemination of harmful content. However, our analysis indicates that modality alone cannot fully account for the structural differences in propagation.The findings highlight the critical role of structural characteristics in the spread of these harmful messages, suggesting that strategies targeting structural characteristics of re-sharing could be crucial in managing the dissemination of such content on private messaging platforms.

cs.SI↗

Extended Persistent Homology Distinguishes Simple and Complex Contagions with High Accuracy

The social contagion literature makes a distinction between simple (independent cascade or bond percolation processes that pass infections through edges) and complex contagions (bootstrap percolation or threshold processes that require local reinforcement to spread). However, distinguishing simple and complex contagions using observational data poses a significant challenge in practice. Estimating population-level activation functions from observed contagion dynamics is hindered by confounding factors that influence adoptions (other than neighborhood interactions), as well as heterogeneity in individual behaviors and modeling variations that make it difficult to design appropriate null models for inferring contagion types. Here, we show that a new tool from topological data analysis (TDA), called extended persistent homology (EPH), when applied to contagion processes over networks, can effectively detect simple and complex contagion processes, as well as predict their parameters. We train classification and regression models using EPH-based topological summaries computed on simulated simple and complex contagion dynamics on three real-world network datasets and obtain high predictive performance over a wide range of contagion parameters and under a variety of informational constraints, including uncertainty in model parameters, noise, and partial observability of contagion dynamics. EPH captures the role of cycles of varying lengths in the observed contagion dynamics and offers a useful metric to classify contagion models and predict their parameters. Analyzing geometrical features of network contagion using TDA tools such as EPH can find applications in other network problems such as seeding, vaccination, and quarantine optimization, as well as network inference and reconstruction problems.

cs.SI↗

Democratic Resilience and Sociotechnical Shocks

We focus on the potential fragility of democratic elections given modern information-communication technologies (ICT) in the Web 2.0 era. Our work provides an explanation for the cascading attrition of public officials recently in the United States and offers potential policy interventions from a dynamic system's perspective. We propose that micro-level heterogeneity across individuals within crucial institutions leads to vulnerabilities of election support systems at the macro scale. Our analysis provides comparative statistics to measure the fragility of systems against targeted harassment, disinformation campaigns, and other adversarial manipulations that are now cheaper to scale and deploy. Our analysis also informs policy interventions that seek to retain public officials and increase voter turnout. We show how limited resources (for example, salary incentives to public officials and targeted interventions to increase voter turnout) can be allocated at the population level to improve these outcomes and maximally enhance democratic resilience. On the one hand, structural and individual heterogeneity cause systemic fragility that adversarial actors can exploit, but also provide opportunities for effective interventions that offer significant global improvements from limited and localized actions.

cs.SI↗

Measuring Network Dynamics of Opioid Overdose Deaths in the United States

The US opioid overdose epidemic has been a major public health concern in recent decades. There has been increasing recognition that its etiology is rooted in part in the social contexts that mediate substance use and access; however, reliable statistical measures of social influence are lacking in the literature. We use Facebook's social connectedness index (SCI) as a proxy for real-life social networks across diverse spatial regions that help quantify social connectivity across different spatial units. This is a measure of the relative probability of connections between localities that offers a unique lens to understand the effects of social networks on health outcomes. We use SCI to develop a variable, called "deaths in social proximity", to measure the influence of social networks on opioid overdose deaths (OODs) in US counties. Our results show a statistically significant effect size for deaths in social proximity on OODs in counties in the United States, controlling for spatial proximity, as well as demographic and clinical covariates. The effect size of standardized deaths in social proximity in our cluster-robust linear regression model indicates that a one-standard-deviation increase, equal to 11.70 more deaths per 100,000 population in the social proximity of ego counties in the contiguous United States, is associated with thirteen more deaths per 100,000 population in ego counties. To further validate our findings, we performed a series of robustness checks using a network autocorrelation model to account for social network effects, a spatial autocorrelation model to capture spatial dependencies, and a two-way fixed-effect model to control for unobserved spatial and time-invariant characteristics. These checks consistently provide statistically robust evidence of positive social influence on OODs in US counties.

cs.SI↗

Differentially Private Distributed Estimation and Learning

We study distributed estimation and learning problems in a networked environment where agents exchange information to estimate unknown statistical properties of random variables from their privately observed samples. The agents can collectively estimate the unknown quantities by exchanging information about their private observations, but they also face privacy risks. Our novel algorithms extend the existing distributed estimation literature and enable the participating agents to estimate a complete sufficient statistic from private signals acquired offline or online over time and to preserve the privacy of their signals and network neighborhoods. This is achieved through linear aggregation schemes with adjusted randomization schemes that add noise to the exchanged estimates subject to differential privacy (DP) constraints, both in an offline and online manner. We provide convergence rate analysis and tight finite-time convergence bounds. We show that the noise that minimizes the convergence time to the best estimates is the Laplace noise, with parameters corresponding to each agent's sensitivity to their signal and network characteristics. Our algorithms are amenable to dynamic topologies and balancing privacy and accuracy trade-offs. Finally, to supplement and validate our theoretical results, we run experiments on real-world data from the US Power Grid Network and electric consumption data from German Households to estimate the average power consumption of power stations and households under all privacy regimes and show that our method outperforms existing first-order, privacy-aware, distributed optimization methods.

cs.LG↗

Optimized Model Selection for Estimating Treatment Effects from Costly Simulations of the US Opioid Epidemic

Agent-based simulation with a synthetic population can help us compare different treatment conditions while keeping everything else constant within the same population (i.e., as digital twins). Such population-scale simulations require large computational power (i.e., CPU resources) to get accurate estimates for treatment effects. We can use meta models of the simulation results to circumvent the need to simulate every treatment condition. Selecting the best estimating model at a given sample size (number of simulation runs) is a crucial problem. Depending on the sample size, the ability of the method to estimate accurately can change significantly. In this paper, we discuss different methods to explore what model works best at a specific sample size. In addition to the empirical results, we provide a mathematical analysis of the MSE equation and how its components decide which model to select and why a specific method behaves that way in a range of sample sizes. The analysis showed why the direction estimation method is better than model-based methods in larger sample sizes and how the between-group variation and the within-group variation affect the MSE equation.

stat.ME↗