SearcharxivSearch

arXiv subjects

Rishi Sharma

Publications and source records attributed to Rishi Sharma.

At least 19 recordsLinked to original sources

Communication-Efficient Secure Aggregation in Decentralized Learning

Decentralized learning (DL) enables participants to collaboratively train models without a central server, yet it faces significant scalability challenges that demand sparsification to reduce the prohibitive communication costs of peer-to-peer exchange. While secure aggregation effectively mitigates privacy risks in standard settings, it has remained fundamentally incompatible with sparsification in decentralized networks due to the mismatch of indices across local updates, forcing a trade-off between communication efficiency and privacy. This paper introduces CESAR, a novel protocol that resolves this incompatibility by integrating secure aggregation and sparsification to provide provable defense against honest-but-curious and colluding adversaries. By coordinating masks over parameter intersections, CESAR supports node dropouts and robust privacy without central aggregation. Empirical evaluations on models up to 124 million parameters demonstrate that CESAR matches the accuracy of non-private baselines while cutting total data exchange by 66.7 % compared to a standard full-parameter decentralized protocol (D-PSGD). With TopK sparsification on IID data, CESAR even exceeds by 0.3 % the accuracy achieved by D-PSGD with sparsification. Collectively, these results establish CESAR as the first decentralized protocol to achieve both privacy and communication efficiency through secure aggregation in DL.

cs.LG

Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems

Collaborative agentic AI is projected to transform entire industries by enabling AI-powered agents to autonomously perceive, plan, and act within digital environments. Yet, current solutions in this field are all built in isolation, and we are rapidly heading toward a landscape of fragmented, incompatible ecosystems. In this position paper, we argue that interoperability, achieved by the adoption of minimal standards, is essential to ensure open, secure, web-scale, and widely-adopted agentic ecosystems. To this end, we devise a minimal architectural foundation for collaborative agentic AI, named Web of Agents, which is composed of four building blocks: agent-to-agent messaging, interaction interoperability, state management, and agent discovery. Web of Agents adopts existing standards and reuses existing infrastructure where possible. With Web of Agents, we take the first but critical step toward interoperable agentic systems and offer a pragmatic path forward before ecosystem fragmentation becomes the norm.

cs.NI

Optimizing Agent Planning for Security and Autonomy

Indirect prompt injection attacks threaten AI agents that execute consequential actions, motivating deterministic system-level defenses. Such defenses can provably block unsafe actions by enforcing confidentiality and integrity policies, but currently appear costly: they reduce task completion rates and increase token usage compared to probabilistic defenses. We argue that existing evaluations miss a key benefit of system-level defenses: reduced reliance on human oversight. We introduce autonomy metrics to quantify this benefit: the fraction of consequential actions an agent can execute without human-in-the-loop (HITL) approval while preserving security. To increase autonomy, we design a security-aware agent that (i) introduces richer HITL interactions, and (ii) explicitly plans for both task progress and policy compliance. We implement this agent design atop an existing information-flow control defense against prompt injection and evaluate it on the AgentDojo and WASP benchmarks. Experiments show that this approach yields higher autonomy without sacrificing utility.

cs.CR

Mosaic Learning: A Framework for Decentralized Learning with Model Fragmentation

Decentralized learning (DL) enables collaborative machine learning (ML) without a central server, making it suitable for settings where training data cannot be centrally hosted. We introduce Mosaic Learning, a DL framework that decomposes models into fragments and disseminates them independently across the network. Fragmentation reduces redundant communication across correlated parameters and enables more diverse information propagation without increasing communication cost. We theoretically show that Mosaic Learning (i) shows state-of-the-art worst-case convergence rate, and (ii) leverages parameter correlation in an ML model, improving contraction by reducing the highest eigenvalue of a simplified system. We empirically evaluate Mosaic Learning on four learning tasks and observe up to 12 percentage points higher node-level test accuracy compared to epidemic learning (EL), a state-of-the-art baseline. In summary, Mosaic Learning improves DL performance without sacrificing its utility or efficiency, and positions itself as a new DL standard.

cs.LG

Optimizing Agentic Workflows using Meta-tools

Agentic AI enables LLM to dynamically reason, plan, and interact with tools to solve complex tasks. However, agentic workflows often require many iterative reasoning steps and tool invocations, leading to significant operational expense, end-to-end latency and failures due to hallucinations. This work introduces Agent Workflow Optimization (AWO), a framework that identifies and optimizes redundant tool execution patterns to improve the efficiency and robustness of agentic workflows. AWO analyzes existing workflow traces to discover recurring sequences of tool calls and transforms them into meta-tools, which are deterministic, composite tools that bundle multiple agent actions into a single invocation. Meta-tools bypass unnecessary intermediate LLM reasoning steps and reduce operational cost while also shortening execution paths, leading to fewer failures. Experiments on two agentic AI benchmarks show that AWO reduces the number of LLM calls up to 11.9% while also increasing the task success rate by up to 4.2 percent points.

cs.AI

An effective field theory for thermal QCD with 2+1 flavours

We write a long-distance effective field theory (EFT) for QCD at finite temperature just below the crossover temperature $T_c$. The low energy constants (LECs) of this EFT are obtained from lattice measurements of the screening mass of pions at two temperatures for $N_f=2+1$ using lattice results obtained at physical values of pion and Kaon masses, and $N_f=2$ where the lattice simulations were performed with a heavier pion mass. The EFT gives good predictions for other static pion properties for $N_f=2$, where lattice results are available. We show the corresponding predictions for $N_f=2+1$, where they are not yet measured. We demonstrate that EFT gives excellent predictions for the phase diagram in $N_f=2+1$. The predictions for the pressure are investigated, and predictions are also given for a Wick-rotated real-time quantity called the kinetic mass.

hep-lat

Efficient Pyramidal Analysis of Gigapixel Images on a Decentralized Modest Computer Cluster

Analyzing gigapixel images is recognized as computationally demanding. In this paper, we introduce PyramidAI, a technique for analyzing gigapixel images with reduced computational cost. The proposed approach adopts a gradual analysis of the image, beginning with lower resolutions and progressively concentrating on regions of interest for detailed examination at higher resolutions. We investigated two strategies for tuning the accuracy-computation performance trade-off when implementing the adaptive resolution selection, validated against the Camelyon16 dataset of biomedical images. Our results demonstrate that PyramidAI substantially decreases the amount of processed data required for analysis by up to 2.65x, while preserving the accuracy in identifying relevant sections on a single computer. To ensure democratization of gigapixel image analysis, we evaluated the potential to use mainstream computers to perform the computation by exploiting the parallelism potential of the approach. Using a simulator, we estimated the best data distribution and load balancing algorithm according to the number of workers. The selected algorithms were implemented and highlighted the same conclusions in a real-world setting. Analysis time is reduced from more than an hour to a few minutes using 12 modest workers, offering a practical solution for efficient large-scale image analysis.

cs.DC

The role of the pion mass on the QCD phase diagram in the $T-eB$ plane

We investigated the role of the pion mass on the QCD phase diagram in the $T-eB$ plane using an effective model treatment. Such treatments are able to capture the main features predicted by first-principles calculations. We also employed the model to estimate the pion mass beyond which the inverse magnetic catalysis (IMC) effect disappears. The value is found to be independent of the strength of the magnetic field.

hep-ph

Medium modifications to jet angularities using SCET with Glauber gluons

We perform a comprehensive analysis of medium modifications on ungroomed jet angularities, $τ_a$, within the framework of Soft-Collinear Effective Theory with Glauber gluons (SCET$_{\rm G}$). Angularities are a one-parameter family of jet substructure observables with angularity exponent $a < 2$ for infrared safety. Variation of the angularity exponent allows one to modify the relative weighting of the collinear-to-soft radiations in the jet, thereby giving access to different moments of the jet transverse momentum spectrum. In this article, we focus on $a<1$ and provide detailed results for $a=-1, 0$, and $0.5$. Within SCET$_{\rm G}$, the interactions between jet and medium constituents are mediated by off-shell Glauber gluons generated from the color sources in the medium. While medium modifications are incorporated into the jet function via the use of medium-induced splitting functions, the soft function remains unmodified for $a<1$. For all values of $a$, we find that compared to jets in vacuum, the medium-modified distributions are shifted towards smaller values of jet angularity and have a steeper fall. This redistribution of the ungroomed angularity spectrum is more apparent for a jet with a larger cone size and for higher values of $a$. We also present results for the medium sensitivity towards $p_T$ of the jet and for a jet initiated in a less central event ($10-30\%$ centrality). Finally, we provide the ratios of nucleus-nucleus and proton-proton differential angularity distributions for different angularity exponents, and for two values of the jet radius parameter.

hep-ph

HarMoEny: Efficient Multi-GPU Inference of MoE Models

Mixture-of-Experts (MoE) models offer computational efficiency during inference by activating only a subset of specialized experts for a given input. This enables efficient model scaling on multi-GPU systems that use expert parallelism without compromising performance. However, load imbalance among experts and GPUs introduces waiting times, which can significantly increase inference latency. To address this challenge, we propose HarMoEny, a novel solution to address MoE load imbalance through two simple techniques: (i) dynamic token redistribution to underutilized GPUs and (ii) asynchronous prefetching of experts from the system to GPU memory. These techniques achieve a near-perfect load balance among experts and GPUs and mitigate delays caused by overloaded GPUs. We implement HarMoEny and compare its latency and throughput with four MoE baselines using real-world and synthetic datasets. Under heavy load imbalance, HarMoEny increases throughput by 37%-70% and reduces time-to-first-token by 34%-41%, compared to the next-best baseline. Moreover, our ablation study demonstrates that HarMoEny's scheduling policy reduces the GPU idling time by up to 84% compared to the baseline policies.

cs.DC

Boosting Resource-Constrained Federated Learning Systems with Guessed Updates

Federated learning (FL) enables a set of client devices to collaboratively train a model without sharing raw data. This process, though, operates under the constrained computation and communication resources of edge devices. These constraints combined with systems heterogeneity force some participating clients to perform fewer local updates than expected by the server, thus slowing down convergence. Exhaustive tuning of hyperparameters in FL, furthermore, can be resource-intensive, without which the convergence is adversely affected. In this work, we propose GEL, the guess and learn algorithm. GEL enables constrained edge devices to perform additional learning through guessed updates on top of gradient-based steps. These guesses are gradientless, i.e., participating clients leverage them for free. Our generic guessing algorithm (i) can be flexibly combined with several state-of-the-art algorithms including FEDPROX, FEDNOVA, FEDYOGI or SCALEFL; and (ii) achieves significantly improved performance when the learning rates are not best tuned. We conduct extensive experiments and show that GEL can boost empirical convergence by up to 40% in resource constrained networks while relieving the need for exhaustive learning rate tuning.

cs.LG

A White Paper on The Multi-Messenger Science Landscape in India

The multi-messenger science using different observational windows to the Universe such as Gravitational Waves (GWs), Electromagnetic Waves (EMs), Cosmic Rays (CRs), and Neutrinos offer an opportunity to study from the scale of a neutron star to cosmological scales over a large cosmic time. At the smallest scales, we can explore the structure of the neutron star and the different energetics involved in the transition of a pre-merger neutron star to a post-merger neutron star. This will open up a window to study the properties of matter in extreme conditions and a guaranteed discovery space. On the other hand, at the largest cosmological scales, multi-messenger observations allow us to study the long-standing problems in physical cosmology related to the Hubble constant, dark matter, and dark energy by mapping the expansion history of the Universe using GW sources. Moreover, the multi-messenger studies of astrophysical systems such as white dwarfs, neutron stars, and black holes of different masses, all the way up to a high redshift Universe, will bring insightful understanding into the physical processes associated with them that are inaccessible otherwise. This white paper discusses the key cases in the domain of multi-messenger astronomy and the role of observatories in India which can explore uncharted territories and open discovery spaces in different branches of physics ranging from nuclear physics to astrophysics.

astro-ph.HE

Non-Markovian dynamics of bottomonia in the QGP

The evolution of quarkonia in the QGP medium can be described through the formalism of Open Quantum Systems (OQS). In previous works with OQS, the quarkonium evolution was studied by working in either quantum Brownian or optical regime. In this paper, we set up a general non-Markovian master equation to describe the evolution of the quarkonia in the medium. Due to the non-Markovian nature, it cannot be cast in Lindblad form which makes it challenging to solve. We numerically solve the master equation for $Υ(1S)$ without considering stochastic jumps for Bjorken and viscous hydrodynamic background at energies relevant to LHC and RHIC. We quantify the effect of the hierarchy between the system time scale, $τ_{\textrm{S}}\sim 1/E_b$, and the environment time scale, $τ_{\textrm{E}}$, on quarkonium evolution and show that it significantly affects the nuclear modification factor. A comparison of our results with the existing experimental data from LHC and RHIC is presented.

hep-ph

Low-Cost Privacy-Preserving Decentralized Learning

Decentralized learning (DL) is an emerging paradigm of collaborative machine learning that enables nodes in a network to train models collectively without sharing their raw data or relying on a central server. This paper introduces Zip-DL, a privacy-aware DL algorithm that leverages correlated noise to achieve robust privacy against local adversaries while ensuring efficient convergence at low communication costs. By progressively neutralizing the noise added during distributed averaging, Zip-DL combines strong privacy guarantees with high model accuracy. Its design requires only one communication round per gradient descent iteration, significantly reducing communication overhead compared to competitors. We establish theoretical bounds on both convergence speed and privacy guarantees. Moreover, extensive experiments demonstrating Zip-DL's practical applicability make it outperform state-of-the-art methods in the accuracy vs. vulnerability trade-off. Specifically, Zip-DL (i) reduces membership-inference attack success rates by up to 35% compared to baseline DL, (ii) decreases attack efficacy by up to 13% compared to competitors offering similar utility, and (iii) achieves up to 59% higher accuracy to completely nullify a basic attack scenario, compared to a state-of-the-art privacy-preserving approach under the same threat model. These results position Zip-DL as a practical and efficient solution for privacy-preserving decentralized learning in real-world applications.

cs.LG

Practical Federated Learning without a Server

Federated Learning (FL) enables end-user devices to collaboratively train ML models without sharing raw data, thereby preserving data privacy. In FL, a central parameter server coordinates the learning process by iteratively aggregating the trained models received from clients. Yet, deploying a central server is not always feasible due to hardware unavailability, infrastructure constraints, or operational costs. We present Plexus, a fully decentralized FL system for large networks that operates without the drawbacks originating from having a central server. Plexus distributes the responsibilities of model aggregation and sampling among participating nodes while avoiding network-wide coordination. We evaluate Plexus using realistic traces for compute speed, pairwise latency and network capacity. Our experiments on three common learning tasks and with up to 1000 nodes empirically show that Plexus reduces time-to-accuracy by 1.4-1.6x, communication volume by 15.8-292x and training resources needed for convergence by 30.5-77.9x compared to conventional decentralized learning algorithms.

cs.DC

Boosting Asynchronous Decentralized Learning with Model Fragmentation

Decentralized learning (DL) is an emerging technique that allows nodes on the web to collaboratively train machine learning models without sharing raw data. Dealing with stragglers, i.e., nodes with slower compute or communication than others, is a key challenge in DL. We present DivShare, a novel asynchronous DL algorithm that achieves fast model convergence in the presence of communication stragglers. DivShare achieves this by having nodes fragment their models into parameter subsets and send, in parallel to computation, each subset to a random sample of other nodes instead of sequentially exchanging full models. The transfer of smaller fragments allows more efficient usage of the collective bandwidth and enables nodes with slow network links to quickly contribute with at least some of their model parameters. By theoretically proving the convergence of DivShare, we provide, to the best of our knowledge, the first formal proof of convergence for a DL algorithm that accounts for the effects of asynchronous communication with delays. We experimentally evaluate DivShare against two state-of-the-art DL baselines, AD-PSGD and Swift, and with two standard datasets, CIFAR-10 and MovieLens. We find that DivShare with communication stragglers lowers time-to-accuracy by up to 3.9x compared to AD-PSGD on the CIFAR-10 dataset. Compared to baselines, DivShare also achieves up to 19.4% better accuracy and 9.5% lower test loss on the CIFAR-10 and MovieLens datasets, respectively.

cs.DC

Fair Decentralized Learning

Decentralized learning (DL) is an emerging approach that enables nodes to collaboratively train a machine learning model without sharing raw data. In many application domains, such as healthcare, this approach faces challenges due to the high level of heterogeneity in the training data's feature space. Such feature heterogeneity lowers model utility and negatively impacts fairness, particularly for nodes with under-represented training data. In this paper, we introduce \textsc{Facade}, a clustering-based DL algorithm specifically designed for fair model training when the training data exhibits several distinct features. The challenge of \textsc{Facade} is to assign nodes to clusters, one for each feature, based on the similarity in the features of their local data, without requiring individual nodes to know apriori which cluster they belong to. \textsc{Facade} (1) dynamically assigns nodes to their appropriate clusters over time, and (2) enables nodes to collaboratively train a specialized model for each cluster in a fully decentralized manner. We theoretically prove the convergence of \textsc{Facade}, implement our algorithm, and compare it against three state-of-the-art baselines. Our experimental results on three datasets demonstrate the superiority of our approach in terms of model accuracy and fairness compared to all three competitors. Compared to the best-performing baseline, \textsc{Facade} on the CIFAR-10 dataset also reduces communication costs by 32.3\% to reach a target accuracy when cluster sizes are imbalanced.

cs.LG

Dynamics of Hot QCD Matter 2024 -- Hard Probes

The hot and dense QCD matter, known as the Quark-Gluon Plasma (QGP), is explored through heavy-ion collision experiments at the LHC and RHIC. Jets and heavy flavors, produced from the initial hard scattering, are used as hard probes to study the properties of the QGP. Recent experimental observations on jet quenching and heavy-flavor suppression have strengthened our understanding, allowing for fine-tuning of theoretical models in hard probes. The second conference, HOT QCD Matter 2024, was organized to bring the community together for discussions on key topics in the field. This article comprises 15 sections, each addressing various aspects of hard probes in relativistic heavy-ion collisions, offering a snapshot of current experimental observations and theoretical advancements. The article begins with a discussion on memory effects in the quantum evolution of quarkonia in the quark-gluon plasma, followed by an experimental review, new insights on jet quenching at RHIC and LHC, and concludes with a machine learning approach to heavy flavor production at the Large Hadron Collider.

nucl-ex